Hacker News

New stories

Live mirror
30 storiesupdated just nowView source snapshot
  1. Anthropic is planning to launch its IPO in November(wsj.com ↗)
    discuss
  2. Valve has open-sourced Lepton, its tool to bring Android games to Steam(theverge.com ↗)
    discuss
  3. Exclusive-Anthropic sets up biology lab as it ramps AI drug program(yahoo.com ↗)
    discuss
  4. I Hate Workday(nelson.cloud ↗)
    discuss
  5. Y Combinator's PAC is throwing money at Republicans across the country(gazetteer.co ↗)
    discuss
  6. War may be coming. Are we psychologically ready?(bbc.com ↗)
    discuss
  7. Flawed AI Intel Brought US to the Brink of Confronting China, Claims Report(ndtvprofit.com ↗)
    discuss
  8. AI cloud provider Nscale files for IPO on $140.6M revenue, $1.02B net loss(cnbc.com ↗)
    discuss
  9. Show HN: Find Street Parking in NYC(avi.nyc ↗)
    discuss
  10. Software-Based Live Migration for RDMA – Proceedings of the ACM Sigcomm 2025(acm.org ↗)
    discuss
  11. Anthropic sets up a bio research lab for physical experiments(engadget.com ↗)
    2comments
  12. Show HN: Prohibition of Nuclear Launch Automation
    discuss
  13. A CPU Backdoor (2025)(phrack.org ↗)
    discuss
  14. Show HN: TypeSeer – on-device autocomplete for every text field on macOS(typeseer.com ↗)
    1comments
  15. Show HN: Rediagram – put the diagrams you have on your brand(rediagram.app ↗)
    discuss
  16. We Must Create the Shit Machine(mcsweeneys.net ↗)
    1comments
  17. Intel Appears to End Its Bug Bounty Program(phoronix.com ↗)
    discuss
  18. DeepSeek v4.1 Flash avg 102 tps on 4x RTX6000 pro max-q, 2.1x up from v4-flash(level1techs.com ↗)
    2comments
  19. We were right (about passkeys) all along(mailpace.com ↗)
    discuss
  20. Stack Overflow relaunched Developer Story (who certifies that a human wrote it?)(stackoverflow.blog ↗)
    discuss
  21. Friday Facts #446 – An ARM and a Frame(factorio.com ↗)
    discuss
  22. The Pain Axis: LLMs Represent Self-Directed Harm and Act to Relieve It(arxiv.org ↗)
    discuss
  23. I Have Been a DelGuard
    discuss
  24. The Mic Is On. So Is Live Auto-Tune.(nytimes.com ↗)
    1comments
  25. Institutional Parasitism in Open Technology Communities(wasabisys.com ↗)
    discuss
  26. UK could force phone companies to add 'anti-theft protections'(bbc.com ↗)
    discuss
  27. Using jev to improve product experiences is pretty crazy(elvex.com ↗)
    3comments
  28. What I learned from using FreeBSD as a main OS for a summer(divanv.com ↗)
    discuss
  29. A Lack of Honesty Is the Ultimate Killer(phillipspobrien.substack.com ↗)
    discuss
  30. Ask HN: What do you think of Noul, a new decision primitive
    discuss

Fable 5.1 vs. Astra for coding: Fable 2X more expensive per task but solves more

1 pointsby 6h agoaistack.imec-int.com
4 comments
6h agoHN ↗

One of the authors here - We had both run through a curated set of 64 long horizon coding tasks to evaluate cost/intelligence. Astra surprised us in this default setting. How is the rest of you faring (especially interested in those who use both)

6h agoHN ↗

Fable solved 17 more tasks than Astra, but also billed €68 more for them.

So Astra billed for tasks it couldn’t solve?

5h agoHN ↗

Correct. But that's by design here.

Context : We continuously do runs of a subset of 64 curated SWE-Bench Pro long horizon tasks on various models. Some via their respective API's. Some on reference hardware setups. We do this to approximate 'real world use' of these models and get more insights on how these use the underlying compute (we advise HW builders)

In our runs we create multiple separate sandboxed agents that use a given model/harness combo (in this case just the default harnesses for both) and feed them the benchmark tasks, which we compare against the golden resolution. (There's also a time out, just to make sure we don't blow through our entire budget by accident). You pay for the tokens no matter if the task gets solved or not, so both bills cover all 64 attempts, also the failed ones (30 for Astra, 13 for Fable). The €68 is just the difference between the two full runs, not the price of those 17 tasks. That's basically why we look at cost per solved task, €1.6 vs €2.4. But we see how the wording of that sentence wasn't optimal. We'll fix that bullet (thx)

Aside from it giving us a good view on evolving token needs of different model generations (and being able to compare with open weights models), the variations in "intelligence per dollar" per generation is also pretty interesting. Hence posts like these, just to share the data, which is hopefully useful to others too. And : always interested in seeing different results.

5h agoHN ↗

I’d rather pay more and get things solved.

Penny wise and pound foolish comes to mind.