- 19comments
- 11comments
- 128comments
- 1comments
- 44comments
- 38comments
- 166comments
- 54comments
- 230comments
- 30comments
- —discuss
- 151comments
- 25comments
- 168comments
- 24comments
- 295comments
- 181comments
- 88comments
- 39comments
- 92comments
- —discuss
- 3comments
- 3comments
- 150comments
- 45comments
- 197comments
- 9comments
- 313comments
- 17comments
- 68comments
Article: "How Good Are Frontier Models at Physics? Expert Re-Grading Reveals Broken Evaluations and Near-Saturation of Leading Benchmarks"
John Sous from Yale posted a fairly solid study on how nearly all the physics benchmarks are broken, as they consistently mark correct answers as incorrect.
When hand grading instead, they found out that the models have actually already saturated the benchmarks which is a little bit scary.
This is interesting and actually very important for robotics.
I've been waiting for this, but all companies seem to not care much now.
There is a way out of this by supplying right context (needs a bit of expertise in physics)
1 more year and frontier will become crazy good at this as well.