Hacker News

New stories

Live mirror
30 storiesupdated just nowView source snapshot
  1. TikTok users' cameras hacked with open weight model(washingtonpost.com ↗)
    discuss
  2. Copaganda, Punishment, and Policing in the United States(teenvogue.com ↗)
    discuss
  3. Trump Calls AI Fears a Hoax. Inside the White House, the Debate Is More Complex(nytimes.com ↗)
    discuss
  4. Show HN: Signal – Audience Intelligence for Film, TV and Game Releases(skipthecritics.com ↗)
    discuss
  5. Gavin Newsom Wants AI Kill Switch to Boost CA AI Safety(bloomberg.com ↗)
    discuss
  6. At Dreamforce business leaders say last year's models are enough(cnbc.com ↗)
    discuss
  7. Show HN: Jeff – A read-only CLI for semantic code review using Jev(github.com/alurith ↗)
    discuss
  8. Conway's Law and Programming Languages(weblog.lol ↗)
    discuss
  9. Jev as a Primitive Feature of Ruby(github.com/innocentdiaz ↗)
    discuss
  10. The Stoic Library(thestoiclibrary.com ↗)
    discuss
  11. EA Safety(venkateshrao.com ↗)
    1comments
  12. Character.ai's CEO Karandeep Anand Is Joining Disney(character.ai ↗)
    discuss
  13. The case for a robot tax to redistribute wealth(restofworld.org ↗)
    discuss
  14. Asus Ascent QN10, first mini PC powered by Qualcomm 18-core ARMv9 CPU(cnx-software.com ↗)
    discuss
  15. How China is preparing for the risk of AI escaping human control(reuters.com ↗)
    discuss
  16. Singapore is paying people to read books – can it fix the reading crisis?(theguardian.com ↗)
    discuss
  17. ReallyFree(reallyfree.app ↗)
    discuss
  18. Show HN: Knucklebones(apps.apple.com ↗)
    discuss
  19. Internet Phone Book(internetphonebook.net ↗)
    discuss
  20. Jev might be Qwen finetune(nitter.click ↗)
    discuss
  21. Benchmarking Wild vs. Mold(davidlattimore.github.io ↗)
    discuss
  22. Product Roadmaps: How the Best Product Teams Plan for Uncertainty(producttalk.org ↗)
    discuss
  23. Humanoid Robot Regulations have begun(youtube.com ↗)
    discuss
  24. Turns out that "HDMI 2.1" ports don't need to support HDMI 2.1 features(arstechnica.com ↗)
    discuss
  25. Disney Hires Its First CTO: Karandeep Anand, Former CEO of AI Chatbot Startup(variety.com ↗)
    discuss
  26. Show HN: Keydris, Sudo for AI Agents(youtube.com ↗)
    discuss
  27. Better Than O1 Visa?(uscis.gov ↗)
    discuss
  28. Claude Code Projects(claude.com ↗)
    discuss
  29. US Military had close call after using AI for hallucinated intelligence report(cnn.com ↗)
    3comments
  30. Scam Spotting with ChatGPT(chatgpt.site ↗)
    discuss

Fable 5.1 vs. Astra for coding: Fable 2X more expensive per task but solves more

1 pointsby 2h agoaistack.imec-int.com
4 comments
2h agoHN ↗

One of the authors here - We had both run through a curated set of 64 long horizon coding tasks to evaluate cost/intelligence. Astra surprised us in this default setting. How is the rest of you faring (especially interested in those who use both)

1h agoHN ↗

Fable solved 17 more tasks than Astra, but also billed €68 more for them.

So Astra billed for tasks it couldn’t solve?

1h agoHN ↗

Correct. But that's by design here.

Context : We continuously do runs of a subset of 64 curated SWE-Bench Pro long horizon tasks on various models. Some via their respective API's. Some on reference hardware setups. We do this to approximate 'real world use' of these models and get more insights on how these use the underlying compute (we advise HW builders)

In our runs we create multiple separate sandboxed agents that use a given model/harness combo (in this case just the default harnesses for both) and feed them the benchmark tasks, which we compare against the golden resolution. (There's also a time out, just to make sure we don't blow through our entire budget by accident). You pay for the tokens no matter if the task gets solved or not, so both bills cover all 64 attempts, also the failed ones (30 for Astra, 13 for Fable). The €68 is just the difference between the two full runs, not the price of those 17 tasks. That's basically why we look at cost per solved task, €1.6 vs €2.4. But we see how the wording of that sentence wasn't optimal. We'll fix that bullet (thx)

Aside from it giving us a good view on evolving token needs of different model generations (and being able to compare with open weights models), the variations in "intelligence per dollar" per generation is also pretty interesting. Hence posts like these, just to share the data, which is hopefully useful to others too. And : always interested in seeing different results.

1h agoHN ↗

I’d rather pay more and get things solved.

Penny wise and pound foolish comes to mind.