Hacker News

New stories

Live mirror
30 storiesupdated just nowView source snapshot
  1. What Jev Means for the Future of Evals(armank.com ↗)
    discuss
  2. Hands-On with Googlebooks(tomshardware.com ↗)
    discuss
  3. North Korea Stopped Nuclear Testing in 2017, Triggered 1k Earthquakes Since(sciencealert.com ↗)
    discuss
  4. The Sun Doesn't Shine on Me (2006)(fullhoffman.com ↗)
    discuss
  5. BBC removes 11 Mitchell and Webb comedy sketches from iPlayer(bbc.co.uk ↗)
    discuss
  6. You're a Meat Proxy(twitter.com/i ↗)
    discuss
  7. Antigravity Spaces – Switch Antigravity IDE Projects Easily(github.com/apolloswave ↗)
    discuss
  8. Iron willpower is an illusion – composure is the real engine of self-control(drdeborst.substack.com ↗)
    discuss
  9. Show HN: Jobs at Recently Funded Startups(vcbacked.co ↗)
    discuss
  10. Chess Design Timeline(chess-timeline.vercel.app ↗)
    1comments
  11. Database Normalization(wikipedia.org ↗)
    discuss
  12. We Asked a Urologist Whether Icing Your Testicles Boosts Testosterone(dartmouth.edu ↗)
    discuss
  13. Guess the Country from Demographic Data(demoguessr.com ↗)
    1comments
  14. Show HN: TickerWhale – nightly 0-100 scores and fair values for 1,300 stocks(tickerwhale.com ↗)
    discuss
  15. Googlebook – Meet the Lineup and pre-order(googlebook.google ↗)
    1comments
  16. Grok 4.7 Intelligence, Performance and Price Analysis(artificialanalysis.ai ↗)
    discuss
  17. Show HN: Agent Chaperone – Screen AI agent tool calls and results with Jev(github.com/agent-chaperone ↗)
    discuss
  18. Show HN: flyOS – A fruit fly connectome simulated in real time on an iPhone(becomethefly.com ↗)
    discuss
  19. Saudi Arabia's Ceer launches flagship electric vehicles(agbi.com ↗)
    discuss
  20. Modulate ML Team Announces New Public Entity Transcription Benchmark(modulate.ai ↗)
    discuss
  21. WTF Is Up with Napster's AI Pivot?(tedium.co ↗)
    discuss
  22. Alcor: Simulate cpuid and sgdt/sidt results per-process(github.com/er-azh ↗)
    discuss
  23. Google hit with €403M fine by Irish data watchdog over GDPR violations(bbc.co.uk ↗)
    discuss
  24. Raspberry Pi founder Eben Upton: 'I'm an Omni-geek(ft.com ↗)
    discuss
  25. Grok 4.7 is here with Electrical engineering benchmark which beats fable 5.1 max(twitter.com/hive_echo ↗)
    discuss
  26. Delta A21N at Kahului on Sep 19th 2026, fuel fumes on board(avherald.com ↗)
    discuss
  27. WWLD #1: Domains and gas station hotdogs(chaosguru.substack.com ↗)
    1comments
  28. The agents, they just want to talk(snats.xyz ↗)
    1comments
  29. Amazon Blocks Meta's Muse AI Agent from Its Retail Site(bloomberg.com ↗)
    1comments
  30. Show HN: Foremerge – Catch Intent Conflicts Between Parallel Coding Agents(github.com/naw103 ↗)
    discuss

Fable 5 – Median thinking declined in August

30 pointsby 43m agotwitter.com
11 comments
20m agoHN ↗

I seem to recall Anthropic going on record saying that they don't do anything to model performance to stretch their compute capacity. I've anecdotally noticed massive peaks and troughs in performance week to week (albeit with Opus, not Fable).

I wonder what their official explanation for this behavior is.

9m agoHN ↗

Last time they were called out, it was a regression in Claude code itself.

At least that's their explanation. Either way, it wasn't a good look for "vibecoding" but it got brushed over.

17m agoHN ↗

How do you measure thinking tokens? They don't send those back to the client.

11m agoHN ↗

They tell you how many tokens are used, however, right? Otherwise you couldn't see your own token consumption.

9m agoHN ↗

Good point. I suppose watching the number go up is useful information in itself.

I have been using CC with DeepSeek 4.1 Flash lately, and it's nice to see how the sausage is being made (even if it's partly illusory, as CoT always is.)

12m agoHN ↗

Obviously. The standard pattern is that model X is basically AGI and wins all benchmarks, followed the next day by Y and Z, which both win all benchmarks, too.

Then weeks later people find out that they have been duped and complain that the models have been quantized or employ worse inference.

Buy decent coffee instead of your $200 subscription and sidestep all the scams.

4m agoHN ↗

You forgot a stage or two:

1: "Our model will bring about the end of all things. Flee, flee for your lives"

2: "Our model is basically AGI"

3: "Our model will be available in limited release next week"

4: "Everybody who subscribes at the $200 level gets access now"

5: "Everybody who subscribes at the $20 level gets access now"

6, at least at Google: "Our model will be shoved down your throat every time you do a search, whether you want it or not"

4m agoHN ↗

Well I happen to enjoy coffee and $200 AI plans. What if Blue Bottle started watering down it's coffee? Is your answer to stop drinking coffee and make myself tea instead?

Information that vendors are watering down or otherwise being misleading in what they are delivering in their product is important to share even if you don't use that product yourself.

7m agoHN ↗

Anecdotally, I have found the same. I spend a lot of time with these frontier models, brainstorming, etc. and the drop in performance from, say, week 1 to week 8 is often massive. Whereas in the beginning, it seemed like a capable research assistant, by the end of week 8 or so it starts acting like a puppy dog eager to make its 'master' happy for a few treats.

5m agoHN ↗

How do create repeatable tests in a non-deterministic system? Every time you send the same prompt you get a different answer.