Hacker News

New stories

Live mirror
30 storiesupdated just nowView source snapshot
  1. Breaking the 1.58-bit Barrier for Ternary LLMs(arxiv.org ↗)
    discuss
  2. 2026 Small World in Motion Competition – Nikon Small World(nikonsmallworld.com ↗)
    discuss
  3. RustFS 1.0.0 GA: Production-Ready, Open Source, S3-Compatible Object Storage(rustfs.com ↗)
    discuss
  4. Zuck's bot increased our Vercel bill 10x
    discuss
  5. How VoltDB Works(voltdb.com ↗)
    discuss
  6. Ask HN: People unemployed for >2 years, how do you spend your time?
    1comments
  7. Ukrainian USV Attack on Sochi Reveals Rarely Seen Russian Navy Trained Dolphins(hisutton.com ↗)
    discuss
  8. Copyrightability of LLM-Generated Code(fsfe.org ↗)
    discuss
  9. I'm getting sued by a data center(technically.beehiiv.com ↗)
    discuss
  10. Trump's war on EVs derailed America's auto-factory revival(indiatimes.com ↗)
    discuss
  11. Show HN: A simple Mac extension for recent screenshots(github.com/kanishkvashisht ↗)
    1comments
  12. Everything we know about good agent design(rubriclabs.com ↗)
    discuss
  13. R/DoohickeyCorporation(reddit.com ↗)
    1comments
  14. WASM to Go Source Code(github.com/goccy ↗)
    discuss
  15. Ask HN: How would you anthropomorphize the Anthropic Claude logo?
    discuss
  16. Automattic's interim CEO and legal chief signed reciprocal severance deals(techcrunch.com ↗)
    1comments
  17. Arrow 2 and Arrow 2 Telos(quiver.ai ↗)
    discuss
  18. History of Forgotten Chips [video](youtube.com ↗)
    discuss
  19. An air taxi that floats: Regent's Seaglider takes flight with people on board(smartcitiesdive.com ↗)
    discuss
  20. Brushstrokes: Avant-Garde and the Psychology of Perception(eclecticlight.co ↗)
    discuss
  21. The Rhythm of Your Breath Leaves a Fingerprint on Your Thoughts(nautil.us ↗)
    discuss
  22. New satellites and artificial intelligence are transforming wildfire detection(theguardian.com ↗)
    discuss
  23. How to use HDR brightness as a graphic design hack(hdrlogo.com ↗)
    1comments
  24. DOE vs. GitHub, INC: LLM generated-content not a DMCA violation [pdf](uscourts.gov ↗)
    discuss
  25. Racing driver is about to race 100 karts at once(arstechnica.com ↗)
    discuss
  26. What's New in PHP 8.6(laravel-news.com ↗)
    1comments
  27. Denuvo Sues Game Cracker 'Voices38' for Bypassing Its Anti-Tamper DRM(torrentfreak.com ↗)
    discuss
  28. Claude.md is good for taste and project context.It's a weak place for invariants(tesseracted-labs-blog.vercel.app ↗)
    discuss
  29. Think about your goals and threat models for MFA authentication(utcc.utoronto.ca ↗)
    discuss
  30. Priest, Monk, and Mathematician(logangraves.com ↗)
    discuss

How good are frontier models at physics?

27 pointsby 1h agoarxiv.org
6 comments
1h agoHN ↗

Article: "How Good Are Frontier Models at Physics? Expert Re-Grading Reveals Broken Evaluations and Near-Saturation of Leading Benchmarks"

John Sous from Yale posted a fairly solid study on how nearly all the physics benchmarks are broken, as they consistently mark correct answers as incorrect.

When hand grading instead, they found out that the models have actually already saturated the benchmarks which is a little bit scary.

58m agoHN ↗

I would be very surprised if any of the frontier models wasn't trained on all public physics benchmarks. Training data providers have been hiring people for exactly this task.

18m agoHN ↗

This study appears to be evidence against that: the model failed the benchmark but arrived at the correct answer.

11m agoHN ↗

Half of the benchmarks don't have published answers. That's why training data providers have been hiring physicists to solve them.

1h agoHN ↗

This is interesting and actually very important for robotics.

I've been waiting for this, but all companies seem to not care much now.

There is a way out of this by supplying right context (needs a bit of expertise in physics)

1 more year and frontier will become crazy good at this as well.

27m agoHN ↗

(Trained physicist here)

From personal experience, frontier models absolutely struggle with understanding a physical situation based on words. (Okay, I haven't played with Astra much. GPT-5.6 Sol makes outrageous errors that anyone understanding a real world object would not make. And I was just asking it about NPT threads, not advanced physics.)

But seriously, what's up with these benchmarks? The example question in the paper is:

PHYBench, problem 140: equivalent expressions for the same rope tension

Problem statement. Three identical homogeneous balls are placed on a smooth horizontal surface, touching each other and are close enough to each other. A rope is wrapped around the spheres at the height of their centers, tying them together. A fourth identical sphere is placed on top of the three spheres. Find the tension T in the rope. It is given that the weight of each sphere is P.

For some reason the paper was focused on the fact that the grader didn't notice that some models were producing answers that were trivially algebraically equivalent to the reference answer. But this is missing the elephants in the room:

1. "touching each other and are close enough to each other": the right response is "hey, Professor, what do you mean 'close enough to each other'? They're sitting on a table in an equilateral triangle, all touching (i.e. tangent at their equators), right? Did you have a different configuration in mind?

2. The answer is 0. Go find four baseballs or foursquare balls or whatever, make a little triangle with three of them, and balance the fourth one on top. It's not especially hard on an appropriate surface. Now loosely wrap an imaginary rope around them (but see below) to keep them from moving - no tension is needed because they're not moving anyway. So the models and the reference answer are wrong, IMO.

3. How, exactly, do you plan to wrap a rope around the spheres, at equator height, with no built-in tension (not pre-stretched), such that the rope does not immediately fall off? Friction? But I suspect you need to pretend there is no friction to get the reference answer. (Or maybe that the marble-marble interface has friction but the marble-table interface doesn't? Again, I haven't tried to reverse engineer it.) So maybe the right answer is "infinity or impossible -- in the scenario where the rope is needed, the rope will promptly fall off because it cannot be stable in the described configuration and gravity pulls it down, and once the rope falls off the tension will be zero and the top marble will fall and the other three will roll over the rope."

4. The answer might be "any tension you like -- just wrap the rope with the desired amount of tension". Imagine three baseballs in a triangle with a rubber band around them and a fourth baseball on top for good measure. The tension is a function of what rubber band you choose.

I'm sure there's an interpretation of the question that makes the reference answer correct, and I was not inspired to try to reverse engineer it.

My tentative conclusion is that LLMs are almost unbelievably good at solving problems that are fully contained within the inputs and (training/verification) outputs, and that they and the people training them are not actually particularly good at the input and output parts. If you are training a model to benchmaxx this benchmark, you are training a bad model.