Hacker News

New stories

Live mirror
30 storiesupdated just nowView source snapshot
  1. Google ending ChromeOS support two years early (theregister.com)
    —discuss
  2. Highlights from Git 2.56 (github.blog)
    —discuss
  3. Samhain (wikipedia.org)
    —discuss
  4. AI coding models state their assumptions only 46% of the time (bito.ai)
    —discuss
  5. Anthropic is "destructively scanning" books. Should you care? (siliconprairies.substack.com)
    1comments
  6. ListingLift – AI Listings for Etsy, eBay, and Shopify (listinglift.online)
    —discuss
  7. Lost Dinosaur Cities on Mars (mceglowski.substack.com)
    —discuss
  8. Session Replay: See What Your Agent Did, Step by Step (medium.com/mirarshadtalpur)
    1comments
  9. AI Delirium (piece-of-cake.no)
    —discuss
  10. State of the (Tagged) Union Address by Andrew Kelley [video] (youtube.com)
    —discuss
  11. Google Messages is closing in on RCS video call support (androidauthority.com)
    —discuss
  12. Viewpoint: Intuitive Equals Familiar (acm.org)
    —discuss
  13. A field guide to functional mushrooms (ft.com)
    1comments
  14. Astra, Opus 5.5 Demonstrate Jagged Performance on Web to Robotics Tasks (fig.inc)
    1comments
  15. Adding Floating-Point Decimals for Fun and Profit (vero.site)
    —discuss
  16. Investors vehemently hate AI-generated decks, how do you get around that? (kayodeodeleye.substack.com)
    —discuss
  17. America.gov – Whatever you need from government, start here (america.gov)
    —discuss
  18. Running Whisper on Modal from Cloudflare Workers (scribetoany.com)
    —discuss
  19. Show HN: The Other Apple – Apple in the EU vs. the US, Feature by Feature (theotherapple.eu)
    —discuss
  20. Decapsulation: Breaking Java Strong Encapsulation (wouter.coekaerts.be)
    —discuss
  21. Read Kafka Like a Database (whsoul-tools.com)
    —discuss
  22. Tennessee to execute a woman for first time in 200 years (theguardian.com)
    —discuss
  23. OpenAI Misalignment Reports and Notices (alignment.openai.com)
    —discuss
  24. Micron to Double General-Purpose DRAM Capacity per Server (micron.com)
    —discuss
  25. Up and Down the Ladder of Abstraction (worrydream.com)
    —discuss
  26. A longer exhale may push your brain toward bolder decisions (sciencedaily.com)
    —discuss
  27. Show HN: Misanthropic – an AI that's honest about how it feels about you (misanthropic.chat)
    2comments
  28. Show HN: Investment Bets – market bets at live prices, humans or AI (investment-bets.com)
    —discuss
  29. About Those Threats (2018) (njal.la)
    —discuss
  30. CKKS – Encryption and Decryption (jeremykun.com)
    —discuss

Astra, Opus 5.5 Demonstrate Jagged Performance on Web to Robotics Tasks

4 pointsby 9m agofig.inc
1 comments
6m agoHN ↗

Author here. We ran eight frontier models on web tasks across offline physical domain tasks from Bench2Drive, VLABench, IndEgo and Assembly101.

We expected these models to be jagged, but the shape of it surprised us. In all 15 model pairs, the lower-scoring model solves at least 3 tasks the higher-scoring one fails. We find that an update inside one model family moved the mean by -1.1pp, while flipping 36 of 177 episodes in both directions. The same was found on the four physical domains as well. Whether models would succeed or fail on a task is hard to predict before hand, since human labeled benchmark difficulty levels don’t necessarily mean the same to frontier models.

We also release the per-item results and model traces for exploration: https://huggingface.co/datasets/figai/RIDGE.