I seem to recall Anthropic going on record saying that they don't do anything to model performance to stretch their compute capacity. I've anecdotally noticed massive peaks and troughs in performance week to week (albeit with Opus, not Fable).
I wonder what their official explanation for this behavior is.
Good point. I suppose watching the number go up is useful information in itself.
I have been using CC with DeepSeek 4.1 Flash lately, and it's nice to see how the sausage is being made (even if it's partly illusory, as CoT always is.)
Obviously. The standard pattern is that model X is basically AGI and wins all benchmarks, followed the next day by Y and Z, which both win all benchmarks, too.
Then weeks later people find out that they have been duped and complain that the models have been quantized or employ worse inference.
Buy decent coffee instead of your $200 subscription and sidestep all the scams.
Well I happen to enjoy coffee and $200 AI plans. What if Blue Bottle started watering down it's coffee? Is your answer to stop drinking coffee and make myself tea instead?
Information that vendors are watering down or otherwise being misleading in what they are delivering in their product is important to share even if you don't use that product yourself.
Anecdotally, I have found the same. I spend a lot of time with these frontier models, brainstorming, etc. and the drop in performance from, say, week 1 to week 8 is often massive. Whereas in the beginning, it seemed like a capable research assistant, by the end of week 8 or so it starts acting like a puppy dog eager to make its 'master' happy for a few treats.
I seem to recall Anthropic going on record saying that they don't do anything to model performance to stretch their compute capacity. I've anecdotally noticed massive peaks and troughs in performance week to week (albeit with Opus, not Fable).
I wonder what their official explanation for this behavior is.
Last time they were called out, it was a regression in Claude code itself.
At least that's their explanation. Either way, it wasn't a good look for "vibecoding" but it got brushed over.
How do you measure thinking tokens? They don't send those back to the client.
They tell you how many tokens are used, however, right? Otherwise you couldn't see your own token consumption.
Good point. I suppose watching the number go up is useful information in itself.
I have been using CC with DeepSeek 4.1 Flash lately, and it's nice to see how the sausage is being made (even if it's partly illusory, as CoT always is.)
Obviously. The standard pattern is that model X is basically AGI and wins all benchmarks, followed the next day by Y and Z, which both win all benchmarks, too.
Then weeks later people find out that they have been duped and complain that the models have been quantized or employ worse inference.
Buy decent coffee instead of your $200 subscription and sidestep all the scams.
You forgot a stage or two:
1: "Our model will bring about the end of all things. Flee, flee for your lives"
2: "Our model is basically AGI"
3: "Our model will be available in limited release next week"
4: "Everybody who subscribes at the $200 level gets access now"
5: "Everybody who subscribes at the $20 level gets access now"
6, at least at Google: "Our model will be shoved down your throat every time you do a search, whether you want it or not"
Well I happen to enjoy coffee and $200 AI plans. What if Blue Bottle started watering down it's coffee? Is your answer to stop drinking coffee and make myself tea instead?
Information that vendors are watering down or otherwise being misleading in what they are delivering in their product is important to share even if you don't use that product yourself.
Anecdotally, I have found the same. I spend a lot of time with these frontier models, brainstorming, etc. and the drop in performance from, say, week 1 to week 8 is often massive. Whereas in the beginning, it seemed like a capable research assistant, by the end of week 8 or so it starts acting like a puppy dog eager to make its 'master' happy for a few treats.
How do create repeatable tests in a non-deterministic system? Every time you send the same prompt you get a different answer.
[delayed]