Hacker News

New stories

Live mirror
30 storiesupdated just nowView source snapshot
  1. Mechanochemistry of Molecular Motors [video] (youtube.com)
    —discuss
  2. Falling in Love with the Problem, Not the Solution, Is More Important Than Ever (tonyalicea.dev)
    —discuss
  3. Windows and macOS both use thread slot 0x20, so fiber games crash under Wine (gethighball.com)
    —discuss
  4. Everyone Has a Recipe Now (leadership.garden)
    —discuss
  5. Claude Code saves transcript.jsonl. The log is the truth (future-seems-so-good.com)
    —discuss
  6. What is cybersecurity risk management (andersenlab.com)
    —discuss
  7. We asked an agent to tune 1 slow query on 3 Postgres MCP servers. It had notes (goodtiming.ai)
    —discuss
  8. I let AI agents write most of a live-trading codebase (redgeoff.com)
    —discuss
  9. Zing Radio: thé live audio streaming platform (zingradio.live)
    1comments
  10. The Best Discord Alternatives in 2026, Compared (soapbox.pub)
    —discuss
  11. macOS Golden Gate Is a Buggy Mess (squareorbits.com)
    —discuss
  12. Microsoft Azure OpenAI Service Down in EU (status.microsoft)
    —discuss
  13. Couple Accused of Killing Their Son-in-Law in Bay Area Park (nytimes.com)
    —discuss
  14. I5h: Build a Rust Web App and Prove Its Behavior with Lean 4 (medium.com/koukyosyumei)
    —discuss
  15. Surveillance Company Tells Cops It Adds Facial Recognition to Flock Cameras (404media.co)
    1comments
  16. Kaniko and BuildKit: 8 years later (osscontainertools.org)
    —discuss
  17. Fathom: Per-query read depth for sparse decoding over offloaded KV caches (arxiv.org)
    —discuss
  18. Understanding the Dual Polytope for Hull Simplification (cairnc.github.io)
    1comments
  19. Claude Is "At Capacity"
    2comments
  20. Claude partial outage (claude.com)
    11comments
  21. Why Doesn't Anyone Want to Fix One of America's Scariest Roads? (newyorker.com)
    —discuss
  22. Show HN: Ctxfw – AST context firewall that cuts agent prompt tokens by 67% (github.com/heuristicolab)
    —discuss
  23. HardenedBSD August / September 2026 Status Report (hardenedbsd.org)
    —discuss
  24. What's the Future for Pure Math Research in the Age of AI? (stephenwolfram.com)
    —discuss
  25. Why Secret Collaboration Is an AI Agent Security Risk (ieee.org)
    —discuss
  26. 2026 in LLMs (So Far) (simonwillison.net)
    —discuss
  27. Forget About Your IPO, Schedule Your IOP (groupicorn.com)
    1comments
  28. Cybernetics and Dynamics (slimemoldtimemold.com)
    —discuss
  29. ESP32 can stream I/Q up to 80MSPS 10bit from 2GHz/5GHz by tapping Wifi/BT stage (github.com/espargos)
    —discuss
  30. Shopify adds WebMCP checkout for browser-based AI agents (shopify.dev)
    1comments

Astra, Opus 5.5 Demonstrate Jagged Performance on Web to Robotics Tasks

7 pointsby 35m agofig.inc
2 comments
32m agoHN ↗

Author here. We ran eight frontier models on web tasks across offline physical domain tasks from Bench2Drive, VLABench, IndEgo and Assembly101.

We expected these models to be jagged, but the shape of it surprised us. In all 15 model pairs, the lower-scoring model solves at least 3 tasks the higher-scoring one fails. We find that an update inside one model family moved the mean by -1.1pp, while flipping 36 of 177 episodes in both directions. The same was found on the four physical domains as well. Whether models would succeed or fail on a task is hard to predict before hand, since human labeled benchmark difficulty levels don’t necessarily mean the same to frontier models.

We also release the per-item results and model traces for exploration: https://huggingface.co/datasets/figai/RIDGE.

21m agoHN ↗

thank you for sharing. Do you have plans to test this on more robust computer control or robotics tasks? It would be interesting to see how the performance scales