Hacker News

New stories

Live mirror
30 storiesupdated just nowView source snapshot
  1. Understanding the Dual Polytope for Hull Simplification (cairnc.github.io)
    1comments
  2. Claude Is "At Capacity"
    —discuss
  3. Claude partial outage (claude.com)
    —discuss
  4. Why Doesn't Anyone Want to Fix One of America's Scariest Roads? (newyorker.com)
    —discuss
  5. Show HN: Ctxfw – AST context firewall that cuts agent prompt tokens by 67% (github.com/heuristicolab)
    —discuss
  6. HardenedBSD August / September 2026 Status Report (hardenedbsd.org)
    —discuss
  7. What's the Future for Pure Math Research in the Age of AI? (stephenwolfram.com)
    —discuss
  8. Why Secret Collaboration Is an AI Agent Security Risk (ieee.org)
    —discuss
  9. 2026 in LLMs (So Far) (simonwillison.net)
    —discuss
  10. Forget About Your IPO, Schedule Your IOP (groupicorn.com)
    —discuss
  11. Cybernetics and Dynamics (slimemoldtimemold.com)
    —discuss
  12. ESP32 can stream I/Q up to 80MSPS 10bit from 2GHz/5GHz by tapping Wifi/BT stage (github.com/espargos)
    1comments
  13. Shopify adds WebMCP checkout for browser-based AI agents (shopify.dev)
    —discuss
  14. Unsurprisingly, Meta's new Muse AI agent blatantly ignores users permissions (appleinsider.com)
    —discuss
  15. Show HN: We built a way for web apps to message users' phones without SMS (thenotifier.app)
    1comments
  16. Inside the Biggest Feud in Artificial Intelligence (theatlantic.com)
    1comments
  17. Google ending ChromeOS support two years early (theregister.com)
    —discuss
  18. Highlights from Git 2.56 (github.blog)
    —discuss
  19. Samhain (wikipedia.org)
    —discuss
  20. AI coding models state their assumptions only 46% of the time (bito.ai)
    —discuss
  21. Anthropic is "destructively scanning" books. Should you care? (siliconprairies.substack.com)
    1comments
  22. ListingLift – AI Listings for Etsy, eBay, and Shopify (listinglift.online)
    —discuss
  23. Lost Dinosaur Cities on Mars (mceglowski.substack.com)
    —discuss
  24. Session Replay: See What Your Agent Did, Step by Step (medium.com/mirarshadtalpur)
    1comments
  25. AI Delirium (piece-of-cake.no)
    —discuss
  26. State of the (Tagged) Union Address by Andrew Kelley [video] (youtube.com)
    —discuss
  27. Google Messages is closing in on RCS video call support (androidauthority.com)
    —discuss
  28. Viewpoint: Intuitive Equals Familiar (acm.org)
    —discuss
  29. A field guide to functional mushrooms (ft.com)
    1comments
  30. Astra, Opus 5.5 Demonstrate Jagged Performance on Web to Robotics Tasks (fig.inc)
    2comments

Astra, Opus 5.5 Demonstrate Jagged Performance on Web to Robotics Tasks

7 pointsby 23m agofig.inc
2 comments
19m agoHN ↗

Author here. We ran eight frontier models on web tasks across offline physical domain tasks from Bench2Drive, VLABench, IndEgo and Assembly101.

We expected these models to be jagged, but the shape of it surprised us. In all 15 model pairs, the lower-scoring model solves at least 3 tasks the higher-scoring one fails. We find that an update inside one model family moved the mean by -1.1pp, while flipping 36 of 177 episodes in both directions. The same was found on the four physical domains as well. Whether models would succeed or fail on a task is hard to predict before hand, since human labeled benchmark difficulty levels don’t necessarily mean the same to frontier models.

We also release the per-item results and model traces for exploration: https://huggingface.co/datasets/figai/RIDGE.

9m agoHN ↗

thank you for sharing. Do you have plans to test this on more robust computer control or robotics tasks? It would be interesting to see how the performance scales