Hacker News

New stories

Live mirror
30 storiesupdated just nowView source snapshot
  1. Apple Copland D11E4 Booting in the Browser(pagetable.com ↗)
    discuss
  2. Avoiding the babbling-idiot failure in a time-triggered communication system(ieee.org ↗)
    discuss
  3. Kara (2012) [video](youtube.com ↗)
    discuss
  4. Five stages of grief AI edition(operationsoptimist.substack.com ↗)
    discuss
  5. Show HN: Four-Leaf MCP, open-source job search and interview prep(github.com/fourleafai ↗)
    discuss
  6. ReadWhile. read while you work(apps.microsoft.com ↗)
    1comments
  7. Maybe Meta Is Right About AI, or at Least Jeremy Stern Is(philippdubach.com ↗)
    discuss
  8. Anthropic at $2T isn't far-fetched(ft.com ↗)
    1comments
  9. Physics in the Age of LLMs(ozamram.substack.com ↗)
    discuss
  10. Hermes Agent now supports Claude Pro/Max subscriptions(nousresearch.com ↗)
    discuss
  11. You Probably Don't Need Another Infrastructure Project(medium.com/majid.fekri ↗)
    discuss
  12. Measurements for understanding the pace of AI development inside frontier labs(anthropic.com ↗)
    discuss
  13. Saudi CEER just launched 2 EVs(ceermotors.com ↗)
    discuss
  14. Show HN: Artwork List – map of artworks near you(artworklist.com ↗)
    discuss
  15. Show HN: Combinators in Array Languages(softwarewrighter.com ↗)
    discuss
  16. Writing Without Spaces with Jev(levmiseri.com ↗)
    discuss
  17. I gave an autonomous AI company $0 and $200 of debt(autonomouscompany.substack.com ↗)
    discuss
  18. AI-Induced Dehumanization(wiley.com ↗)
    discuss
  19. State attorneys general agree to settle Paramount-Warner Bros merger lawsuit(cnn.com ↗)
    discuss
  20. Halo: Post-train LLMs 3x faster than TRL and Megatron(twitter.com/whitecircle ↗)
    1comments
  21. A (necessarily) brief history of the reminder of mathematics(cofault.com ↗)
    discuss
  22. Grok 4.7(twitter.com/spacexai ↗)
    discuss
  23. Show HN: Gitstats – compare commits and lines with your friends(gitstats.org ↗)
    discuss
  24. David Pogue: '125 Tests of the New AI Siri'(daringfireball.net ↗)
    discuss
  25. The Long Dream of the Googlebook(theverge.com ↗)
    1comments
  26. The Next AI Infrastructure Challenge Is Before the First Token(radicaldatascience.wpcomstaging.com ↗)
    discuss
  27. Show HN: I made a Pomodoro timer that speeds up time(arsh.zip ↗)
    discuss
  28. Show HN: Viaduct – C4 models that coding agents can read and update(quietgridlabs.com ↗)
    discuss
  29. OpenAI Urges U.S. Government to Create Global AI-Safety Standards(wsj.com ↗)
    discuss
  30. Apple M6 SoC Analysis's 2 nm chip crushes AMD, Intel and Qualcomm(notebookcheck.net ↗)
    discuss

Fable 5 – Median thinking declined in August

150 pointsby 2h agotwitter.com
88 comments
1h agoHN ↗

I seem to recall Anthropic going on record saying that they don't do anything to model performance to stretch their compute capacity. I've anecdotally noticed massive peaks and troughs in performance week to week (albeit with Opus, not Fable).

I wonder what their official explanation for this behavior is.

1h agoHN ↗

Last time they were called out, it was a regression in Claude code itself.

At least that's their explanation. Either way, it wasn't a good look for "vibecoding" but it got brushed over.

1h agoHN ↗

When something is new, its capabilities feel incredible. Over time, those same capabilities become mundane, and you start to notice the flaws.

(Now, if TFA is actually measuring reasoning tokens, that's quite different! It's not entirely obvious to me how he is measuring.)

1h agoHN ↗

I don’t think that’s what’s going on. I notice flaws on day one of model releases. But I also notice improvements if the model is truly more advanced than what I’m used to. Then over time the same questions or tasks return worse results.

What is actually stopping these model companies from running a model at full capacity on release then once its name rings out, start serving users quantized garbage?

2m agoHN ↗

What is actually stopping these model companies from running a model at full capacity on release then once its name rings out, start serving users quantized garbage?

...I mean, if they were actually doing this despite saying that they don't—promising one product and delivering something else—I think that would be fraud, no?

And, maybe it's one thing to defraud normies like us (although class action lawsuits exist), but I don't think major enterprises or the US military would take too kindly to it.

1h agoHN ↗

They are deploying optimizations weekly (if not daily) with various AB tests. They don't manipulate model performance, but they do actively perform tests.

42m agoHN ↗

And here's another great example of how a bunch of people who don't know what's going on throw noise into the system. That post is simply confused: the 1m opus calls are the auto-mode classifier, actual agent calls are still in Fable.

31m agoHN ↗

Look at the usage. Fable wasn't being consumed.

1h agoHN ↗

How do you measure thinking tokens? They don't send those back to the client.

1h agoHN ↗

They tell you how many tokens are used, however, right? Otherwise you couldn't see your own token consumption.

1h agoHN ↗

Good point. I suppose watching the number go up is useful information in itself.

I have been using CC with DeepSeek 4.1 Flash lately, and it's nice to see how the sausage is being made (even if it's partly illusory, as CoT always is.)

1h agoHN ↗

Obviously. The standard pattern is that model X is basically AGI and wins all benchmarks, followed the next day by Y and Z, which both win all benchmarks, too.

Then weeks later people find out that they have been duped and complain that the models have been quantized or employ worse inference.

Buy decent coffee instead of your $200 subscription and sidestep all the scams.

1h agoHN ↗

You forgot a stage or two:

1: "Our model will bring about the end of all things. Flee, flee for your lives"

2: "Our model is basically AGI"

3: "Our model will be available in limited release next week"

4: "Everybody who subscribes at the $200 level gets access now"

5: "Everybody who subscribes at the $20 level gets access now"

6, at least at Google: "Our model will be shoved down your throat every time you do a search, whether you want it or not"

1h agoHN ↗

0. "our model is too dangerous to release to pubic"

1h agoHN ↗

Google's search AI actually its too dangerous to release to the public. I have relatives routinely citing it as their source for medical advice.

I have quite strongly told them, in no uncertain terms, that they are going to kill themselves doing that.

48m agoHN ↗

Be sure to eat plenty of rocks in your daily diet, they are chock full of valuable minerals!

1h agoHN ↗

Well I happen to enjoy coffee and $200 AI plans. What if Blue Bottle started watering down it's coffee? Is your answer to stop drinking coffee and make myself tea instead?

Evidence that vendors are being misleading in what they are delivering is important to share, whether or not you personally approve of that product.

1h agoHN ↗

Well sorry, still have to get decent AI somewhere. Productivity without AI is about 5x less. I am not comfortable with paying Chinese companies, and no Western companies provide subscription-based pricing for open models.

1h agoHN ↗

Anecdotally, I have found the same. I spend a lot of time with these frontier models, brainstorming, etc. and the drop in performance from, say, week 1 to week 8 is often massive. Whereas in the beginning, it seemed like a capable research assistant, by the end of week 8 or so it starts acting like a puppy dog eager to make its 'master' happy for a few treats.

47m agoHN ↗

Reminds me of how slot machine users swear the odds have changed on a machine.

also when someone says you just have to prompt it a certain way it reminds me of people who think they can get better results out of a slot machine by pressing buttons in a certain order

The providers of these models also design the UX similarly to slot machines (run it x amount of times for better results, multiplying your spend) this isnt a coincidence and they're playing into the gambler mentality, and probably hire UX designers that specialize in this.

34m agoHN ↗

Wtf are you on my dude. Anthropic UI is designed like a slot machine? Hiring slot machine specialists? Sometimes I can’t believe im even on HN anymore with comments like this.

29m agoHN ↗

I think some of it comes from that they do not publicly let you see the random seed. So each time you ask the answer is different (like a slot machine) and if they let users use the random seed it would let people much more accurately assess if an underlying model changed somehow (same seed and same input will always have the same output).

Of course the closed Anthropic would never share this, it would definitely take away the 'magic' feeling of the AI

3m agoHN ↗

To my understanding, with batched inference and other "optimizations" you wouldn't get the exact same token predictions even with temp=0.0.

28m agoHN ↗

Well, running an LLM X amount of times does give you better results provided you are willing to select the best one out of the X yourself.

But I agree with your general point. One of the reasons subscription plans are cheaper because they modulate usage in this way based on demand. They can also recover compute more coarsely via usage resets.

8m agoHN ↗

At that point I might as well do it myself

1h agoHN ↗

How do you create repeatable tests in a non-deterministic system? Every time you send the same prompt you get a different answer.

1h agoHN ↗

The actual tokens might be non-deterministic, but you could look for proxy measures that are supposed to be invariant. Eg. correctness/performance on benchmarks, "thinking level" on complex problems, etc

21m agoHN ↗

That's like the entire field of statistics.

1h agoHN ↗

I strongly believe that the real Fable is the one we had for a few days in June. Then they nerfed the model a bit after the government pulled it off the market. What we have now is something less, but still good

27m agoHN ↗

I also believe this. Fable post-ban was never the same. At the least, whatever system prompt munging or pre/post filtering they did to strengthen the guardrails nerfed it.

8m agoHN ↗

question is if they ever let the general public access borderline AGI

1h agoHN ↗

It is clear by now to me that Anthropic is constantly trying to find a kind of “auto” degradation perhaps to save money on work it thinks does not require high reasoning. I always use max reasoning and I can clearly see differences between the models when they release and after 3-4 weeks. I think they give a kind of intelligence boost also for new accounts.

1h agoHN ↗

Just yesterday I was thinking about gpt-5.6-luna. I made it my default model in Hermes during its fist week of launch. It was just as good as 5.5 which was my previous default. But over the last 2 or 3 weeks I've seen how dumb it is now. I have to be very explicit with it.

For example, I used to be able to prompt "Check the system logs on <server> for...." and it would just figure it out. Yesterday I asked "Did <service> on <server> complete the overnight job" and all it said was "that service is not installed on my host"

I had to tell it to ssh into the server and run journlctl to check it

Anecdotal, I know, but they all seem to be less capable with time.

_edit_ I use the same reasoning level of `medium`

39m agoHN ↗

Same exact experience. I worked with both Fable and Sol foe the last two months, daily for several hours, and got used to the very bright, quick thinking, proactive even.

As of last 2 weeks or so both models are nearly on par with DeepSeek4.1 now, which I also use a lot. They're still better, but that difference is not as pronounced as before and, importantly, the frustration level is now on par.

Whatever they're doing will surely drive people to less advanced but predictable, self hosted open models. I sure would rather use DS4.1 with Qwen/GLM in adversarial mode than deal with this b/s I pay significant amount of money.

Me and my friends have been contemplating on getting an Ultra M5 256 and splitting the cost. PI harness is so good now that this is really a viable alternative.

16m agoHN ↗

I'm fairly sure it's just luck of the draw if you get put onto a quant'd model or not. I've seen luna xhigh change intelligence fairly drastically on a day to day basis.

1h agoHN ↗

The smart takeaway is not skepticism or snark, but understanding that once the new datacenter buildout starts coming online, cheap and widespread access to even the current frontier models (without strict thinking limits) will blow the economy wide open.

(ie, even a pause in AI training isn't going to stop the train where AI flips the economy upside down, we've barely even seen the impact of the current frontier)

1h agoHN ↗

Anthropic is straight up scamming its users at this point.

1h agoHN ↗

The question I have is this only happening for a subset of users working in specific areas, such as AI or distributed systems (https://news.ycombinator.com/item?id=48742153), or is this across the board? I am working on distributed systems. Today Fable is mostly unusable. It resembles Opus, so I went looking to see if anyone else is having issues. Sure enough.

1h agoHN ↗

I work in embedded systems. I have seen the same thing happening day by day from Opus. Some days it’s okay to use and performs well. Other days I have to correct it repeatedly and remind it of information already in the prompt earlier (before compaction!) and still other times it’s infuriatingly stupid.

It’s a slot machine for what they’re actually giving us behind the opaque paywalls.

Yes, I’m on a business subscription plan.

1h agoHN ↗

I have no hard data but I have a strong feeling this morning that something's wrong with Fable 5 compared to Friday evening.

Just an hour ago I had Fable correctly identify an unused method that could be deleted. I then immediately get a diff for an exact duplicate method, and then Fable outputting, "I accidentally duplicated <method> instead of deleting it. Removing both copies now."

The remaining morning complaints that makes it feel like something's off is that it will do a lot of "thinking" for simple things that previously took very little time. And it got very lost and completely mixed up DE-91M predicate names and implementations. Just absolute disaster code that I had over the past months come to generally expect it to do without issue.

Glad I carefully review everything. I think what I need is reliability and consistency. But it feels like picking a model from the list doesn't guarantee that: that the models' "brain" is open on the table and they're screwing with it.

1h agoHN ↗

New release of fable and opus 5.5 is pending and Anthropic is reallocating resources. Degradation always happens in transition, it sucks.

Opus 5.5 is being served under opus 5 right now.

59m agoHN ↗

Especially with the frequent releases aka version bumps.

44m agoHN ↗

Can you elaborate on the mechanism of this degradation? If resources are not available I would expect a request to fail with a message about resources not available. Do they tweak back end model capabilities to maintain service in a degraded state?

39m agoHN ↗

Dollars to donuts, they are speculating, and not privy to inside information on the topic.

However, I believe that runtime model quantization is possible with some publicly-available inference engines (e.g. vLLM), so its not beyond belief that the closed labs do quantize at runtime, either to allocate compute, or to nudge users towards a preferred model (e.g. make the incumbent model dumber to push people to use the latest-and-greatest model, or vice versa to ease the load on the latest model, which is typically larger than the old one).

35m agoHN ↗

Opus 5.5 is being served under opus 5 right now.

On what basis are you claiming this?

35m agoHN ↗

It shouldn’t be an excuse. They are selling a product and that product should always be within the quality range.

1h agoHN ↗

Maybe they are jealous of Navier Stokes and try the Hodge conjecture with 80% of total compute at the expense of their customers.

44m agoHN ↗

I would rather wait in a queue than be routed to a degraded model. And if they _have_ to degrade the models, then I wish they would fucking tell us. Instead, it's "I have a strong feeling".

That we have to guess at this is by far the worst part of the AI era. It feels like a dark cloud over my productivity. It makes my body tense for the entire day when it happens. Not healthy.

27m agoHN ↗

They have repeatedly said they do not ever intentionally reduce model quality and do not degrade in this way, and that a model version number is always the same.

But, of course, OP is an empirical claim to the contrary, and I'd be curious to see if anyone (who's been capturing data over these timeframes) can replicate the same results and if Anthropic has any comment.

15m agoHN ↗

Every official statement I've seen around this is careful to say that they "don't intentionally reduce model quality", which leaves plenty of room for "we adjusted some knobs and our evals show performance is materially the same".

However, I also agree that I haven't seen any robust data from someone tracking it daily/weekly. The handful of sites purporting to do this aren't even running it enough times to hit stat sig.

32m agoHN ↗

The Claude models definitely felt more susceptible to moods, like you could leave them for a few hours, come back and it suddenly was unable to do things which it was doing just earlier, which tellingly is never an experience I've had with an open model.

Honestly I lost patience with Anthropic both clearly messing around with things like this and their agitation over regulation. They aren't good actors, and quite why so many blindly trust them with their company crown jewels is a mystery.

22m agoHN ↗

Claude Code's prompt cache expires after 1 hour.

25m agoHN ↗

If you follow reddit forums for claude code, its common to see people, on the same day, claiming that Opus/Fable is especially smart today, and especially dumb today.

I think people are still not used to non deterministic tools like this, and human perception is absolutely horrible at evaluating trends like this no matter how smart, clever, and experienced you are.

If you have a bank of rigorously tested benchmarks that you run every few days, with enough trials to know what your standard deviation is, and you are getting significant trends over time with those, that would be interesting.

But "I have a feeling" and "Seems like" really isn't a reliable signal at all, humans just can't handle perceiving these things reliably. On top of that changes in your work environment can easily pollute LLMs and change quality of results. Are things getting added to your memory or claude.md files that you don't realize? Is your project growing in size and thus claude is performing worse as more context is needed to work with it? etc etc

11m agoHN ↗

It’s load shedding. They’re reducing consumption for capacity balancing at your expense. Whenever there are rate limiting storms Claude gets dumber. They also shift capacity for new releases, and Claude gets dumber leading up to it.

Self run infrastructure won’t have this cost but you have to manage the capacity and rollouts yourself, at which point it’s more obvious what’s happening, but the effects will be the same. The not knowing makes it harder, but also harder to plan your own work around.

1h agoHN ↗

Gemini Chat is constantly throwing, "Pro is in high demand right now, a different model was used for this generation," too.

I'm thinking they're all running out of physical resources. It's the DotCom bubble all over again; rollout of the physical infrastructure that's necessary to keep all of the pie-in-the-sky promises will not happen on the timescales that investors can work with, and they will panic when they realize this.

EDIT: And, frankly, I can't wait. I'm tired of the sketchy and dishonest way these companies are behaving.

1h agoHN ↗

They want transparency from everyone else but not for them ... you don't say.

1h agoHN ↗

This is like shared clouds back in the day where if someone is using the CPU more it impacts you, just pool every one to the same service. There should be an SLA but for the intelligence of these models, otherwise, you are sold fable but with the intelligence of a table.

59m agoHN ↗

This may be the result of someone at Anthropic not working urgently.

55m agoHN ↗

This is a project i wanted to implement for a long time. It regularly benchmarks cloud hosted models with private benchmarks. Not just openai & anthropic, popular openrouter models too.

Tests their intelligence, not their diligence.

Sadly i cant think of a way to monetize the service. Also if it ever gets famous enough labs would try to game the system, it would be cat&mouse game that i am not willing to waste time on without any monetary gain.

47m agoHN ↗

I mean I think if this is done well, lots of companies would pay for access to that data. Think like Enterprise subscriptions.

Its similar to other data services I see around my F500 company.

42m agoHN ↗

I built GitHub.com/adrianco/retort to do this. It’s runs lots of experiments and you can contribute results if you have some spare tokens. You can add your own tests, and it runs Claude, Codex, Gemini, Hermes for local models.

54m agoHN ↗

So in 5 years will they lose a suit for intentionally deceiving users? Or is something baked into the ToS by now that allows them to adjust things like this?

47m agoHN ↗

I've followed a few trackers, eg https://marginlab.ai/trackers/claude-code/ , for awhile. For Claude Code the trend, it seems to me at least, is fewer tokens to do the same or better job. Prompt changes, tool ergonomics changes, etc.; I'd be shocked if they didn't A/B every release. Less thinking as measured by tokens isn't necessarily bad if you can get the same results by making it think about the "right" things or structure. They obviously screw up sometimes, and I've always been suspicious with hidden tokens, but I haven't found evidence quality intentionally degrades over time.

44m agoHN ↗

The Office of Weights and Measures exists because, long before any of us were born, in 1836, companies were up to shady shit and consumers were paying for inconsistent products. I.E. Being scammed.

AI companies should be subject to the OWM like any other company that sells a product that varies in weight. Perhaps when a sane administration is re-elected; one that can read history books and comprehend why our regulations exist in the first place. Or have even a semblance of respect for its citizenry.

39m agoHN ↗

Petition to rename them to the Office of Weights and Biases, haha.

38m agoHN ↗

I stopped using Fable long time ago. It's worse than Sonnet. Opus is not much better.

This cycle of new model running at full quantisation and then nerfed few days / weeks after premiere should be called out. Anthropic should also drop the adaptive reasoning scam.

If I pay for Fable, I should get full, not nerfed model at honest pricing.

Regulators should investigate them.

OpenAI is no different. Astra has basically the same problem.

36m agoHN ↗

The ROI just isn't there. It feels like Fable is in the same place Opus was early last year; at best marginal improvement that's barely noticeable over the lower model, for 10x the cost. I'm sure it'll take over as the workhorse as Opus did once they get it down, but right now it just doesn't make sense

19m agoHN ↗

It's not really 10x the cost though, with the low cost of cached read it's more like maybe 1.2x the cost.

35m agoHN ↗

Makes sense, no? Test time compute is something you can vary, so it makes sense that you start covertly reducing it once the model has already made it's splash.

29m agoHN ↗

Could there be a benefit to releasing a new model, slowly dumbing it down over a couple months, then releasing a new model that’s marginally if at all better than the original to create a perceived improvement when in reality there isn’t really one?

For an industry that’s stagnant in progress yet relies on new frequent releases to survive (non-progress being an existential risk), this could make sense.

I have no idea if that’s what’s happened, I completely pulled it out of my butt. And I have no idea is the actual frontier is stagnating.

27m agoHN ↗

Could there be a benefit to releasing a new model, slowly dumbing it down over a couple months, then releasing a new model that’s marginally if at all better than the original to create a perceived improvement when in reality there isn’t really one?

Exactly what I am saying for months now. And it's exactly the reason why I am shifting to open weight models now. Just bought myself a 2x DGX Spark Cluster. Will run Qwen3.8 Flash Next on it, maybe Qwen4 when it comes out.

Not only do I have full control over quantization and inference, but also will I experience a constant level of quality. It won't be frontier. But it will be stable, and that's enough reason for me to switch. Also I will likely save some money on subscriptions.

20m agoHN ↗

Also I will likely save some money on subscriptions.

Unlikely. The $200 Claude subscription allows for billions of tokens/month, and that kind of hardware will take years to amortize.

11m agoHN ↗

Last time I estimated it was like 30 years to pay back. I doubt the hardware will even last that long.

26m agoHN ↗

Every SOTA model I've used at launch uses deeper, longer inference then gradually turns down over time, until the next model comes out which seems to be trained on some new data, but mostly performance due to deeper longer inference for another period.

17m agoHN ↗

I have not been attributing it so much to malice, just that all the major cloud vendors seem to be running at full capacity, and can't build new datacenters fast enough. I just kind of assumed that as they got busy training newer models, that they allocated less resources to handle the existing systems, because they aren't able to get more capacity right now.

17m agoHN ↗

The Opus 4-6,4-8,5 arc is exactly this. As one person commented in here, opus 5 is a terrorist. This is undeniable. Opus 4-6 was awesome. 4-8 was worse behaviorally but produced better code.

Fable seems to be following the same enshittification arc of other Anthropic models.

Generally OpenAI seems to be taking the opposite approach with an increasing improvement over time. As sad as I feel to say this, open ai seems to have the right strategy. Making your product worse over time rarely plays well with customers. At this point it feel often hard to justify using Anthropic for anything. I generally like Anthropic better as a company and they really had the initiative and advantage and customer good will, then proceeded to squander it faster than a cigarette company or the Sacklers could have.

14m agoHN ↗

This sounds similar to rumors about how SSD companies work. First they would design a new drive with better performance that everyone uses to benchmark against other models; then slowly change its parts to worse ones, either because they are cheaper, the originals are no longer available, or whatever reason

13m agoHN ↗

There must be some benefit if all the providers are doing it independently.

GPT5.6-Sol on Max thinking just became regarded as of a few days ago.

The boosters will tell me it’s my fault for using such an old, cheap out-of-date low quality near useless wish.com model (that was SOTA and better than human coders one month ago).

The cycle repeats.

11m agoHN ↗

Could there be a benefit to releasing a new model, slowly dumbing it down over a couple months, then releasing a new model that’s marginally if at all better than the original to create a perceived improvement when in reality there isn’t really one?

Not unless your competitors do the same, or else you will only be perceived as falling behind others.

3m agoHN ↗

...releasing a new model that’s marginally if at all better than the original...

This isn't what we see in benchmarks.

17m agoHN ↗

Check gpt I think they recently started taking the same route

16m agoHN ↗

Don't they continuously tweak the models post-release?

15m agoHN ↗

Instead of finding a nerfed model, after six weeks of reconstructing wire logs, parsing transcripts, analyzing output tokenization, and staring at data, I found a much deeper issue. The model identity had remained the same, but the inference regime being delivered behind that model had not.

Is this something specific that shows up in the wire log, or is this the author's intepretation? The fact that Claude Code versions change over time in the test is suspicious. Anthropic has stated in the past that the underlying model behavior does not change over time, but Claude Code will change from version to version and this is expected. So if it's just Claude Code more aggressively tuning some knob in its requests, that's a pretty different thing than the underlying model changing.

9m agoHN ↗

It's easy to imagine this only happening for subscription accounts rather than paid API usage. Any data on this?