Hacker News

New stories

Live mirror
30 storiesupdated just nowView source snapshot
  1. Automattic names interim CFO after exec departures(techcrunch.com ↗)
    discuss
  2. What is a System One model and why we need it?(stackness.dev ↗)
    discuss
  3. Show HN: ReacherX – Open-source platform to find and reach the right people(github.com/vecterai ↗)
    discuss
  4. Sam Altman to brief UN Security Council next week(reuters.com ↗)
    discuss
  5. Silex: Laser Enrichment Between Promise and Proliferation Risk(csis.org ↗)
    discuss
  6. Quantum computers will not be that different(arxiv.org ↗)
    discuss
  7. Single player games require age verification if they use Steam under EU KIDS Act(rockpapershotgun.com ↗)
    1comments
  8. A solo founder runs a five-continent tender platform on AlloyDB and MCP(cloud.google.com ↗)
    discuss
  9. Trump Announces Ban of CNN, Politico and MS Now from White House(time.com ↗)
    discuss
  10. Using Cyber Decoys to Strengthen Detection and Response(cisa.gov ↗)
    1comments
  11. Show HN: 3D World Explorer and Location Guesser Game(tadget.net ↗)
    discuss
  12. Saying Goodbye to Firebug (2017)(hacks.mozilla.org ↗)
    1comments
  13. Jev's Architecture Unmasked(archerhume.com ↗)
    discuss
  14. Senior Engineers Are the Next DRAM Shortage(herlein.com ↗)
    1comments
  15. SR-71's "R2-D2" Could Be Key to Winning Future Fights in GPS Denied Environments(twz.com ↗)
    1comments
  16. Terence Tao: SAIR's Open Math Model Initiative [video](youtube.com ↗)
    discuss
  17. California governor signs order to explore AI kill switch(techxplore.com ↗)
    1comments
  18. Communication Doesn't Have a Compiler(wreckitrob.dev ↗)
    discuss
  19. Anthropic Moves Ahead with IPO Plans Amid A.I. Safety Debate(nytimes.com ↗)
    1comments
  20. When the FM Band Goes Transatlantic(radioworld.com ↗)
    discuss
  21. Punctum Books Catalog(punctumbooks.com ↗)
    discuss
  22. AI Error Nearly Triggered U.S. Intercept of Chinese Ship(gcaptain.com ↗)
    1comments
  23. Labeled matches: why is this not in every regex engine?(iev.ee ↗)
    discuss
  24. I Cancelled My Claude Subscription(williamangel.net ↗)
    2comments
  25. Graduating in AI Era Is Like Large Recession for Starting Pay(census.gov ↗)
    1comments
  26. Human brain is two separate organs(stanford.edu ↗)
    1comments
  27. IAM Needs an Architectural Split(gluufederation.medium.com ↗)
    discuss
  28. I used Jev to control a swarm of 15 simulated drones in real time(github.com/khordoo ↗)
    discuss
  29. Napster Is Now Making AI-Powered 'Digital Twins' of Teachers(gizmodo.com ↗)
    discuss
  30. Data Centers Are Breaking the Power Grid [video](youtube.com ↗)
    1comments

Fable 5.1 vs. Astra for coding: Fable 2X more expensive per task but solves more

1 pointsby 5h agoaistack.imec-int.com
4 comments
5h agoHN ↗

One of the authors here - We had both run through a curated set of 64 long horizon coding tasks to evaluate cost/intelligence. Astra surprised us in this default setting. How is the rest of you faring (especially interested in those who use both)

5h agoHN ↗

Fable solved 17 more tasks than Astra, but also billed €68 more for them.

So Astra billed for tasks it couldn’t solve?

4h agoHN ↗

Correct. But that's by design here.

Context : We continuously do runs of a subset of 64 curated SWE-Bench Pro long horizon tasks on various models. Some via their respective API's. Some on reference hardware setups. We do this to approximate 'real world use' of these models and get more insights on how these use the underlying compute (we advise HW builders)

In our runs we create multiple separate sandboxed agents that use a given model/harness combo (in this case just the default harnesses for both) and feed them the benchmark tasks, which we compare against the golden resolution. (There's also a time out, just to make sure we don't blow through our entire budget by accident). You pay for the tokens no matter if the task gets solved or not, so both bills cover all 64 attempts, also the failed ones (30 for Astra, 13 for Fable). The €68 is just the difference between the two full runs, not the price of those 17 tasks. That's basically why we look at cost per solved task, €1.6 vs €2.4. But we see how the wording of that sentence wasn't optimal. We'll fix that bullet (thx)

Aside from it giving us a good view on evolving token needs of different model generations (and being able to compare with open weights models), the variations in "intelligence per dollar" per generation is also pretty interesting. Hence posts like these, just to share the data, which is hopefully useful to others too. And : always interested in seeing different results.

4h agoHN ↗

I’d rather pay more and get things solved.

Penny wise and pound foolish comes to mind.