Hacker News

New stories

Live mirror
30 storiesupdated just nowView source snapshot
  1. Grok Voice Transcribe 2.0(x.ai ↗)
    1comments
  2. Red and Blue America Have Found Something to Agree On: Flock Cameras Must Go(wsj.com ↗)
    discuss
  3. A non-established environment for Reinforcement Learning of PC-98 Touhou games(github.com/touhourl ↗)
    discuss
  4. Epipe on write might mean you're doing it wrong(rachelbythebay.com ↗)
    discuss
  5. GPT-6 Astra: 3D, Embodied AI, and Beyond(wentao.live ↗)
    discuss
  6. TikTok users' cameras hacked with open weight model(washingtonpost.com ↗)
    discuss
  7. Copaganda, Punishment, and Policing in the United States(teenvogue.com ↗)
    discuss
  8. Trump Calls AI Fears a Hoax. Inside the White House, the Debate Is More Complex(nytimes.com ↗)
    discuss
  9. Show HN: Signal – Audience Intelligence for Film, TV and Game Releases(skipthecritics.com ↗)
    discuss
  10. Gavin Newsom Wants AI Kill Switch to Boost CA AI Safety(bloomberg.com ↗)
    discuss
  11. At Dreamforce business leaders say last year's models are enough(cnbc.com ↗)
    discuss
  12. Show HN: Jeff – A read-only CLI for semantic code review using Jev(github.com/alurith ↗)
    discuss
  13. Conway's Law and Programming Languages(weblog.lol ↗)
    discuss
  14. Jev as a Primitive Feature of Ruby(github.com/innocentdiaz ↗)
    discuss
  15. The Stoic Library(thestoiclibrary.com ↗)
    discuss
  16. EA Safety(venkateshrao.com ↗)
    1comments
  17. Character.ai's CEO Karandeep Anand Is Joining Disney(character.ai ↗)
    discuss
  18. The case for a robot tax to redistribute wealth(restofworld.org ↗)
    discuss
  19. Asus Ascent QN10, first mini PC powered by Qualcomm 18-core ARMv9 CPU(cnx-software.com ↗)
    discuss
  20. How China is preparing for the risk of AI escaping human control(reuters.com ↗)
    discuss
  21. Singapore is paying people to read books – can it fix the reading crisis?(theguardian.com ↗)
    discuss
  22. ReallyFree(reallyfree.app ↗)
    discuss
  23. Show HN: Knucklebones(apps.apple.com ↗)
    discuss
  24. Internet Phone Book(internetphonebook.net ↗)
    discuss
  25. Jev might be Qwen finetune(nitter.click ↗)
    discuss
  26. Benchmarking Wild vs. Mold(davidlattimore.github.io ↗)
    discuss
  27. Product Roadmaps: How the Best Product Teams Plan for Uncertainty(producttalk.org ↗)
    discuss
  28. Humanoid Robot Regulations have begun(youtube.com ↗)
    discuss
  29. Turns out that "HDMI 2.1" ports don't need to support HDMI 2.1 features(arstechnica.com ↗)
    discuss
  30. Disney Hires Its First CTO: Karandeep Anand, Former CEO of AI Chatbot Startup(variety.com ↗)
    discuss

Fable 5.1 vs. Astra for coding: Fable 2X more expensive per task but solves more

1 pointsby 2h agoaistack.imec-int.com
4 comments
2h agoHN ↗

One of the authors here - We had both run through a curated set of 64 long horizon coding tasks to evaluate cost/intelligence. Astra surprised us in this default setting. How is the rest of you faring (especially interested in those who use both)

2h agoHN ↗

Fable solved 17 more tasks than Astra, but also billed €68 more for them.

So Astra billed for tasks it couldn’t solve?

1h agoHN ↗

Correct. But that's by design here.

Context : We continuously do runs of a subset of 64 curated SWE-Bench Pro long horizon tasks on various models. Some via their respective API's. Some on reference hardware setups. We do this to approximate 'real world use' of these models and get more insights on how these use the underlying compute (we advise HW builders)

In our runs we create multiple separate sandboxed agents that use a given model/harness combo (in this case just the default harnesses for both) and feed them the benchmark tasks, which we compare against the golden resolution. (There's also a time out, just to make sure we don't blow through our entire budget by accident). You pay for the tokens no matter if the task gets solved or not, so both bills cover all 64 attempts, also the failed ones (30 for Astra, 13 for Fable). The €68 is just the difference between the two full runs, not the price of those 17 tasks. That's basically why we look at cost per solved task, €1.6 vs €2.4. But we see how the wording of that sentence wasn't optimal. We'll fix that bullet (thx)

Aside from it giving us a good view on evolving token needs of different model generations (and being able to compare with open weights models), the variations in "intelligence per dollar" per generation is also pretty interesting. Hence posts like these, just to share the data, which is hopefully useful to others too. And : always interested in seeing different results.

1h agoHN ↗

I’d rather pay more and get things solved.

Penny wise and pound foolish comes to mind.