Hacker News

Best stories

Live mirror
30 storiesupdated just nowView source snapshot
  1. When did Google get so weird? (sancho.bearblog.dev)
    1080comments
  2. Owed a billion dollars in Nvidia stock (colo.to)
    451comments
  3. Sonnet 5.5 (anthropic.com)
    590comments
  4. Updated Google Maps shows destruction of the city of Rafah (twitter.com/aliabunimah)
    780comments
  5. Pirating the Pirates (mubi.com)
    344comments
  6. You are no longer invited to dinner (derekthompson.org)
    525comments
  7. Ember-1 (fireworks.ai)
    248comments
  8. It's Time to Investigate the AI Labs (calnewport.com)
    246comments
  9. Jeff – Jev-compatible 0.8B decision models, trained at home, ~30 ms (github.com/firelex)
    215comments
  10. Coding is not solved (alexewerlof.com)
    522comments
  11. Windows 11½ (definitelynotwindows.com)
    169comments
  12. Kids turned low-traffic NPR Spotify comments into a secret group chat (thisamericanlife.org)
    251comments
  13. AI companies in race to demonstrate their model most threatening to humanity (thecivilian.co.nz)
    392comments
  14. 500k facial scans at UK stations yield no arrests, 1 false positive (theguardian.com)
    249comments
  15. A Privacy Analysis of Web and Mobile Conversational AI Agents [pdf] (jorgegarciaherrero.com)
    123comments
  16. The problem is not AI code, but not knowing about system architecture or intent (ssp.sh)
    238comments
  17. MongoDB CEO resigns to join Meta (reuters.com)
    291comments
  18. macOS Golden Gate Is a Buggy Mess (squareorbits.com)
    258comments
  19. California farmers are struggling to sell grapes as demand for wine drops (kqed.org)
    854comments
  20. GPT 6.1 Sol: Near-Astra intelligence for a fifth of the price (openai.com)
    286comments
  21. Self-Hosting on the Dark Web (alvarezrosa.com)
    115comments
  22. US sanctions force The Netherlands off Microsoft and toward alternative NixOS (tomshardware.com)
    328comments
  23. How Delhi cut electricity loss from 50 to 5 percent (ieee.org)
    189comments
  24. Parley: Federated, decentralised chat that speaks plain IRC (mills.io)
    187comments
  25. Show HN: Lofi Cities – Pixel-art city nights with browser-generated lofi (loficities.com)
    135comments
  26. SpaceX's Starship launching to orbit for first time ever today (space.com)
    360comments
  27. World Labs is Joining AMD (worldlabs.ai)
    111comments
  28. Hijacking the PS5's RTMP stream (yashgarg.dev)
    90comments
  29. Does Reddit have an astroturfing problem? What the data suggests (petervijeh.com)
    360comments
  30. MicroLLM Lab – Try 7 tiny LLM's in the browser (stateofutopia.com)
    111comments

GPT 6.1 Sol: Near-Astra intelligence for a fifth of the price

347 pointsby 1h agoopenai.com
273 comments
1h agoHN ↗

Weren't there headlines just yesterday that they weren't releasing this due to safety concerns?

1h agoHN ↗

GPT-6.1 Astra is what those headlines referred to. This is GPT-6.1 Sol.

1h agoHN ↗

That was 6.1 Astra. And I'm assuming it's being tabled because it still doesn't match Opus 5.5.

This is a decent win though, if it really is better. 6-sol was really no good, at least in my work.

1h agoHN ↗

6 Sol was worse than 5.6 Sol from my own experiences. Far worse.

Will see if this remedies things.

1h agoHN ↗

Yup, same experience here. I used it for one day, spent the next day fixing its lousy code, then went back to 5.6.

1h agoHN ↗

Shots fired, half the price of Opus 5.5.

1h agoHN ↗

Wasn't 6 released like last week? I can't keep up anymore.

1h agoHN ↗

Yes but it was underwhelming, so they seem to have rushed 6.1 Sol out. Also Opus 5.5 may have spooked them too.

1h agoHN ↗

Do you need to? Do you always keep up with all the version bumps on the software you use?

1h agoHN ↗

It was so underwhelming that it didn't even make it to chatgpt chat interface

35m agoHN ↗

Sol 6 is in there? You may be on Enterprise where it didn't roll out by default and comes out in a week or so. (Which is a weird and bad change to their model releases.)

1h agoHN ↗

Ominous for the industry and investors that token price is becoming the main battleground. Could be Anthropic's rationale for IPOing this year.

1h agoHN ↗

Great for the consumer.

I remember when bandwidth was super expensive and now it’s dirt cheap.

1h agoHN ↗

China will do to llms what they did to german cars

59m agoHN ↗

Why make a new account just to post this comment?

It's not even anything controversial..

56m agoHN ↗

They may work for one of the big AI labs.

38m agoHN ↗

I guess I'm out of touch. What did china do to german cars?

28m agoHN ↗

Outcompeted them badly. Sent Porsche packing and BMW bawling.

1h agoHN ↗

Another piece of evidence on the pile that the sudden panic and desire to "slow down" is because they're hitting the plateau on capability

Which, honestly, is fine. A lot of juice to squeeze in efficiency and even if models got zero more capable, making the capability that is already here cheaper is a huge win for everyone (except Nvidia)

56m agoHN ↗

What universe do you live in that you can look at the past six months and see anything like a plateau in capability?

Edit: removed a comment that was uncharitable and rude, for which I apologize.

54m agoHN ↗

Have we honestly seen that great a leap in the last 6 months, or just better application of what we had 6 months before that.

We are seeing multiple frontier models dropping on the same day and no one bats an eye, because it's more of the same.

43m agoHN ↗

The difference between 6 months ago frontier and now frontier in 3d modelling, graphics and video editing is night and day.

4m agoHN ↗

Just because there are new capabilities, doesn't mean they've pushed passed the plateau, they've just expanded where the previous solutions work.

We've gone from 80% in some places to 80% in some more places.

54m agoHN ↗

Not that they've hit it but that they are approaching it. The time to panic and steer the narrative is before you hit the iceberg, not after

12m agoHN ↗

The whole "pace the frontier" thing originated from Anthropic.

With Opus 5.5, it doesn't seem like model improvement is plateauing. And Fable 5.5 will likely be dropped this or next week.

You can only imagine what they've got going internally.

42m agoHN ↗

Most of the impressive accomplishments we’ve seen in the last few months have been the result of huge agent swarms working together and brute-forcing solutions, not massive leaps in intelligence from standalone models. That is still an improvement in the usefulness and power of the technology, but it is NOT evidence that model intelligence is increasing faster than before.

34m agoHN ↗

I don't have any access to any agent swarms (and neither do most) and i still think the models have obviously improved massively in standalone intelligence. Of course they have, agent swarms are not magic. You can swarm all you want around GPT-4 era models and you'll get nowhere. And i've never seen the term 'brute-force' more abused than these LLM discussions. Basically none of the results have been brute force.

25m agoHN ↗

Agreed. You can’t “brute force” reality, which has an infinitely large state space. A million monkeys won’t write Shakespeare and all that

29m agoHN ↗

"This machine-intelligence stuff is overrated, they are just using <insert particular machine-intelligence technique here>" isn't the resounding verdict it may have sounded like when you typed it.

55m agoHN ↗

Another piece of evidence on the pile that the sudden panic and desire to "slow down" is because they're hitting the plateau on capability

if true then LLM related AI (post-post AI winter AI?) is probably one of the fastest inception-to-plateau tech sectors to have ever existed.

We're still improving transistors on a somewhat routine basis.

52m agoHN ↗

The plateau doesn't have to be perfectly flat, but it's not a straight line upward anymore either (kind of like our work on transistors, where we've kind of hit the bounds of speed in clock cycles but are improving on miniaturization and power efficiency)

45m agoHN ↗

It took 6 years to solve ARC-AGI 1, 1 year to solve ARC-AGI 2 and 6 months to solve ARC-AGI 3.

30m agoHN ↗

Those version numbers don't necessarily correspond to equal increases in "difficulty", though.

20m agoHN ↗

Correct, the benchmark became exponentially more difficult as it progressed from pattern matching puzzles to games.

54m agoHN ↗

Nvidia can start putting weights in silicon if model development slows down.

53m agoHN ↗

I think they are hitting compute restrictions. And buying compute right now can be 3-4X. And the costs are increasing. If they train a larger model and demand is high, that’s a lot of compute for Codex subscriptions, which is a loss leader for them. Especially Pro 20X which they just nerfed to 10X.

50m agoHN ↗

Another piece of evidence on the pile that the sudden panic and desire to "slow down" is because they're hitting the plateau on capability

I think it's more a token-cost-demand plateau. They've reached the scale and investor trillions to which they can't 10x the hardware cost of inference any more. They can't afford to compete by eating costs and there isn't appetite for more expensive inference.

So in order that they don't bankrupt each other they're looking for the legal cartel behavior coordinating a stop to growth by convincing governments to regulate them into stopping.

There's a lot of juice to squeeze in efficiency but only so much whereas it seemed like capability was going to continue to scale with parameter count.

Maybe it's good news for everyone that model capability is now going to scale on semiconductor cost meaning huge players are going to be very motivated to make semiconductors cheap.

47m agoHN ↗

It's not so much that they're hitting a plateau in capability, as we're saturating long horizon benchmarks and it's not greatly improving general usability. On the other hand, newer models have been amazing for people interested in 3d, graphics, video editing, etc. The difference between Opus 5.5/Astra and earlier models is night and day even if for many coding tasks they're not a revolution.

46m agoHN ↗

sudden panic and desire to "slow down" is because they're hitting the plateau on capability

I don't think that's the motivation, it's because both companies want to IPO and the _only_ way to even hope to be profitable is to do a whole lot less training, which costs a fortune. But unless Chinese labs go along with this gentleman's agreement (they won't), slowing down on training will bring about the inevitable Chinese model parity date more rapidly. At which point the game is well and truly over for OpenAI and Anthropic. Bit of a pickle they've gotten themselves into with the emphasis on being best, with premium prices to match.

34m agoHN ↗

Another piece of evidence on the pile that the sudden panic and desire to "slow down" is because they're hitting the plateau on capability

People were talking about plateau for years already.

26m agoHN ↗

Is there anything that could happen that you wouldn't use as evidence that they are hitting a plateau?

It just seems like these claims are constant and looking back the calls of 'plateau' between 2023 and 2025 were clearly false, why should we think it's different now?

20m agoHN ↗

Some version of this claim has been made for the past 4 years. There's a data cliff, there's no more compute to buy, the financials don't make sense and all of these orgs will be out of business by end of quarter.

Not once has any of these predictions come true, the pace of progress has continued on it's exponential trajectory since ChatGPT first came to the public's attention.

So why now? What is special about today that suggests all of this is coming to a screeching halt despite all evidence to the contrary?

11m agoHN ↗

Those points were true at the time and most are still true now. But they aren’t predictions.

- it’s correct there isn’t much fresh data anymore

- it’s correct that compute is scarce, that was 100% the case and a huge issue at the beginning of the year, it is better now but still scarce, and hardware is now way, way more expensive

- it’s correct the finances don’t make sense

But there is no way to know when a bubble pop, because it’s a psychological phenomenon across an extremely complicated distributed system (ie the stock and bonds markets)

56m agoHN ↗

This is literally the plan, open weight models are something like 60% of token spend, and it will get worse. many companies now have model gateways where you can slot in cheaper models via cli for cheaper. we've been using glm 5.x and it's pretty close to SOTA frontier models.

it's also why there have been so many calls for regulation and slowdowns.

13m agoHN ↗

Yup.

I see posts about OpenAI and Anthropic latest and don’t even care looking at what they do better. I just read the comments here.

I use DS4.1 Flash and GLM 5.3 Flash, pay peanuts per day and get more than acceptable results.

44m agoHN ↗

Pretty standard business to identify and compete on every axis (cost, speed, intelligence, etc). Often, nobody will be able to maximize every axis so you end up with a polyhedron derived from the axes where there’s a niche for everyone.

DeepSeek understands that. Grok understands it. Every other AI company thinks they need to be the best at everything all the time and it’s weird.

23m agoHN ↗

?! this model launch was around 10% of the dev day and the other time was spent on Dots and things other than models.

10m agoHN ↗

That makes sense. There's not really much else you can say about it.

1h agoHN ↗

If these models are so smart, can't _they_ select the right model for each task?

1h agoHN ↗

the right model for the task is the one that transfers the maximum amount of USD from your pocket to the provider's bank account.

1h agoHN ↗

No because then I'll go to the competition.

1h agoHN ↗

Why don't you simply ask the respective model which model is best for a specific task? :-)

1h agoHN ↗

Switching models is _very_ expensive in compute (you have to rerun everything from the beginning), and highly variable in cost. Cursor tried doing this for awhile, but inconsistent performance/usage means most users turned it off and pick models specifically.

1h agoHN ↗

I guess the question is, does the Dunning Krueger effect apply to models? The dumb ones might think they're up to the task.

50m agoHN ↗

These models have a knowledge cutoff that don't just prevent them from knowing about themselves (especially since most data about the model doesn't even exist until after the model is created), but they also don't know about other recent models. Sure, they can search and use other sources, even make some guesses based on the models they do know, but their default stance is more akin to "User asked about model X, model X doesn't exist, maybe it was an hallucination or mistake, let me do a web search...", but that assumes they have web search and are willing to spend tokens on it.

Personally I've taken to having a list of 3 to 4 models in default context with some ordering on which to prefer. Things like GPT 6 Luna is cheap very cheap, use it. Because otherwise the model will assume Haiku or such is the good cheap model to use.

The speed I'm having to update that document has not gone unnoticed.

1h agoHN ↗

What's driving the increase in release cadence here? We seem to get new models every week or so now, is this RSI?

1h agoHN ↗

Wanting to have the newer model than the competitor, presumably.

58m agoHN ↗

The old "the bigger number is better", GPT announces model 6.1, the obvious thing to do next is to announce Gemini 27, and after that Claudé 3000, then a flute album.

1h agoHN ↗

Response to DeepSeek’s technical paper and competition.

1h agoHN ↗

No, we're pacing ourselves to have the time to evaluate the impact each new model could have, obviously.

1h agoHN ↗

Chinese model pressure. Many of my SWE friends switched to Chinese models. I also use QWEN and GLM for many of the api requiring projects and dropped OpenAI and Anthropic. The only reason was the cost.

EDIT: I love getting downvoted by openai and anthropic employees or their bots.

51m agoHN ↗

I can't recommend Chinese models enough. My personal favorite is DeepSeek v4.1 Flash but I have tried Qwen 3.8, Kimi 3 and GLM 5.3 which are equally impressive but DeepSeek is the cheapest and fastest regularly hitting 270 token per second.

And yeah I have worked with Anthropic and OpenAI models, they're good but they cost a fortune while Chinese models are already really good at a fraction of the cost.

26m agoHN ↗

I was working exclusively with DS 4.1 Flash until Opus 5.5 got me back to a sub. I was disillusioned with what was available.

1h agoHN ↗

Probably just the singularity, no big deal

1h agoHN ↗

Its a news cycle more than anything, and its ONLY going to get much, much worse. Daily releases, or multiple daily, 30-45, by EOY. Welcome to RSI!

1h agoHN ↗

What I don't understand is how much people have to say about every single one. Aren't we at the diminishing returns stage yet? Is there really that much to discuss?

52m agoHN ↗

If you look closely at various benchmarks, you'll see that often models will improve in certain areas while regressing in others. It suggests we're already at the point of diminishing returns.

1h agoHN ↗

I do wonder if people switch back and forth between primary models (GPTvsClaude) that it may be a better idea to simply keep releasing updates as soon as possible in order to keep users from bouncing back and forth.

55m agoHN ↗

Probably one of the factors. Signed up to openai pro a few days ago, deciding between openai and anthropic, then sonnet 5.5 was released and am wondering whether I made a mistake.

Luckily it's not a mistake as now we have access to . . . dots.

(and sol 6.1, it seems)

52m agoHN ↗

This is it.

It's because they need subscription money and interaction data and so keeping a version bump in the wings to stop the bleeding from your competitor's version bump is the logical thing to do. It has nothing to do with RSI.

35m agoHN ↗

jokes on me, I pay for all the subscriptions.

17m agoHN ↗

Maybe process maturity too.

Like think about a software org with good CI/CD versus one without. The mature org can do consistent incremental releases because each one is safe and low overhead, the messier org will do fewer big releases because each release requires a big effort on its own.

As model developers mature we might expect to see more frequent point releases rather than the big bang evolutions.

59m agoHN ↗

New models are distill from the actual unrelease frontier models. They are just giving us better checkpoints.

59m agoHN ↗

The initial response to 6 Sol was bad, and Opus 5.5 was definitely winning the public vibes war. Makes sense to rush something out

57m agoHN ↗

They're releasing Sol 6.1 because 1. Astra 6.1 got postponed 2. Sol 6 is shitty 3. They have to release _something_ in response to Opus 5.5

57m agoHN ↗

Versions is marketing, snapshots/minor variations are easy and the number must go up. Release timing is another OAI's marketing tactic.

RSI

Recursive improvement doesn't imply increased rate, another word for it is "iterative" but this probably sounds too boring for some.

55m agoHN ↗

It is the only way to reduce prices while making it look like a good thing.

47m agoHN ↗

Productivity is increasing as models get smarter; we are ascending the singularity. I'm serious.

47m agoHN ↗

Opus 5.5 is better than they anticipated, it's faster, smarter, cheaper. I'm about to change provider for claude and I'm not the only one

32m agoHN ↗

It feels like an updated 4.6. It's fantastic.

30m agoHN ↗

I'm not the only one

See, that's an/the issue. As soon as people start to flee to the improved model, they start to serve degraded models to keep up with the demand.

32m agoHN ↗

and seems like they're claiming Sol/Opus are not frontier (and only Astra/Fable are)

27m agoHN ↗

Mature training pipelines, plus ever expanding RL datasets of increased quality, and mega GPU clusters to finish training in a few weeks. Automated safety and reliability testing.

1h agoHN ↗

Aka "we made an oopsie last week and released what should have been called GPT 6 Terra with the name GPT 6 Sol"

1h agoHN ↗

GPT‑6.1 Sol matches GPT‑6 Astra at roughly one-fifth of the cost

Astra is a pretty impressive model. Excited to try this.

1h agoHN ↗

I love free market competition. We're getting insane advancements every day. I remember when llms used to cost an arm and a leg for decent intelligence

1h agoHN ↗

This is great. But maybe part of the motivation is that 6-Sol wasn't as good as initially advertised so they needed to tweak it. I felt a clear degradation in quality in some simple refactoring tasks vs 5.6-Sol.

35m agoHN ↗

Yes, obviously. They're both working to make it cheaper, faster, and better at different industries (3d animations, etc). The only direction they are slowing is raw intelligence.

1h agoHN ↗

I wish they'd list the environmental cost. My employer has an unlimited AI budget so I don't care about using Astra if it's just more profit for OpenAI. I care more if it actually uses 5x more energy.

1h agoHN ↗

I don't understand the point of this, why just now when it comes to llms. Why wasn't anyone enraged with the environmental costs of kids playing video games. I would not be surprised the environmental cost of that is an order of magnitude bigger than what llms have.

Edit: for context, just Steam alone has ~200million monthly active users.

57m agoHN ↗

Considering Nvidia's hard shift to crypto and now AI, I doubt videogames are even in the same ballpark.

How many DCs are devoted solely to gaming?

53m agoHN ↗

Considering many games make use of cloud computing for online play and similar functions, they probably make up a pretty goot bit of global cloud compute capacity. Likely quite a lot less than the big AI players, but not an insignificant amount.

49m agoHN ↗

The energy costs of the cloud computing required for gaming are substantially less in power - not to mention overall demand - than LLMs. Come on, we're not in the same energy ballpark here.

39m agoHN ↗

The energy costs of the cloud computing required for gaming are substantially less in power

Yes but it adds up when you consider that just on Steam alone there are 200 million monthly active users.

26m agoHN ↗

You're not including the physical supply chain energy consumption of distributing video game equipment in this analysis. Nobody ships LLMs to big box stores and tries to sell them to consumers.

40m agoHN ↗

How many DCs are devoted solely to gaming?

An entire planet. Just Steam alone has one or two hundres million monthly active users.

47m agoHN ↗

If you recall history past the last 5 minutes, you will remember that people have indeed been enraged with the environmental costs of things for a long time. Its just that AI seems to have induced a mass amnesia, and people tend to forget about what happened pre 2024.

37m agoHN ↗

you will remember that people have indeed been enraged with the environmental costs of things for a long time

Yeah? Show me the big movements against computer gaming.

17m agoHN ↗

There are movements against consumerism and the environmental impacts of industry in general. Greenpeace is over half a century old.

The differences with AI are: 1) we are starting off (mid 2020s) from a baseline point of already being in a hopelessly shitty situation, past the 1.5C warming target; and 2) Electronics, chips, data centers etc were already a thing for a long time, but industry took _decades_ to ramp up production to pre-AI levels, and these things are used everywhere for a huge number of things. Now we're consuming electronics/data centers/water/power at an unheard-of rate, and for a single purpose (AI) with questionable benefits, besides the private interests of a handful of people.

44m agoHN ↗

Because people find video games fun, though I suppose there's some vocal people that think of them as bad for society. In contrast the AI companies are promising a torment nexus future.

I'd be curious as to how much of internet infrastructure is dedicated to gaming though.

14m agoHN ↗

I don't think video games consume nearly as much power. A PS5's power consumption is apparently around 200W. That's not enough to run even one GPU, let alone the armada it presumably takes to run Astra.

Even then people do care about the power consumption of non-AI things. Look at the energy label on your TV or tumble drier for example.

1h agoHN ↗

Given that the number one cost of inference is memory and compute, and the incremental cost of each is energy, cost per inference is roughly proportional to energy consumption.

55m agoHN ↗

Did you care about this when it came to your other computing needs? What PC/laptop are ypu running and how efficient is that?

36m agoHN ↗

Likely orders of magnitude different, this is a weak whataboutism.

16m agoHN ↗

Yes I do. I've got a spare desktop that isn't too efficient (probably ~100W idle but annoyingly I've lost my power meter) so I don't leave it on even though I would like to use it as a server.

Laptops use very minimal power - you don't need to worry about them. If they didn't their battery life would suck.

52m agoHN ↗

Look at Neuralwatt. They report energy usage with every call as well as aggregate statistics.

40m agoHN ↗

I want to energymaxx. Every home should have a nuclear generator for free limitless clean energy. Do not energysimp, we want prosperity for all we must energymaxx and invest heavily in solar/battery/nuclear.

1h agoHN ↗

Okay, now price cut 6 Sol (and rename it to Terra again).

1h agoHN ↗

GPT-6 Sol released a week ago. Shortest model life ever?

1h agoHN ↗

Taking GPT-6 "Sol" outside behind the shed and giving it a merciful end is about the best outcome possible.

Huge misstep releasing it.

1h agoHN ↗

Sorry, I'm GenX. Growing up they showed us "Old Yeller" in the school gym every year like that was some kind of treat.

55m agoHN ↗

The misstep was naming it Sol - it was Terra-level all along.

1h agoHN ↗

I wonder if releasing this soon sort of validates the rumor that Sol 6 was just the Terra model they bumped up and slashed the price.

Then Opus 5.5 caught them off guard and now they're actually releasing the correct sized model.

45m agoHN ↗

If that were true, they’d have axed their margins.

41m agoHN ↗

Whether it was or wasn't, Terra's absence shows Sol has replaced it as the new middle model.

1h agoHN ↗

Looking at the token prices, if this is half as good as 6-Astra for 3D model creation in Blender, it's going to be an absolute game changer.

Opus 5.5 is definitely better at coding, but nothing even comes close to 6-Astra for work in 3D graphics...

1h agoHN ↗

How is it with animations?

I have played around a little bit with fixing some rigging problems and was impressed, but Opus even warned me it was bad at animations cause it can only really grab screenshots to process static content.

57m agoHN ↗

You need to use the Blender MCP. There is an official plugin for this now, so the third party one can be avoided.

I've only dabbled but yes with SOTA models it is very good at animating and really most Blender tasks you can think of. Certainly if you are coming at Blender at below expert level it makes it far more accessible and fun to work with.

There are still rough edges of course. But try the official MCP out with Astra and judge for yourself.

52m agoHN ↗

Have you tried fable? (I did small experiements and was satisfied, but maybe there are reasons to switch?)

41m agoHN ↗

From the results of a lot of YouTubers in the space, I think Opus 5.5 is pretty competitive with Astra in 3D. It's slightly worse at spatial detail but better at aesthetics and little touches.

33m agoHN ↗

After all the hype, I’ve been kinda disappointed tbh. Modeling specific models are so much better (eg. Tripo3d). Astra still models some janky crap for me.

1h agoHN ↗

Cached input costs just $0.10 per million tokens—95% less than standard input pricing and 50% less than GPT‑6 Sol’s cached input pricing

This is the actual big announcement. 50% cheaper cache than GPT-6 Sol will get you far more mileage on Codex.

1h agoHN ↗

Exactly half as expensive as Opus 5.5 in every API pricing metric

54m agoHN ↗

And half as good. I didn't have great experiences with Anthropic models in the past, but Opus 5.5 seems to have turned a major corner. It is churning through tasks significantly more quickly and efficiently.

Suggest trying it out yourself: Ask for something difficult from GPT-6 Sol and Opus 5.5 and watch what each one does. The difference is stark.

Edit: Defining "difficult" as a complex coding or systems task (or even series of them in a single prompt).

52m agoHN ↗

I'm not an OpenAI simp, but how anyone can have any opinion on the performance of these models in less than a day - let alone a few hours - is beyond me.

50m agoHN ↗

Try it, it's that good compared to openai current offering.

I get better results and usage our of my $20 claude sub than my $100 openai sub... it's that ridiculous

43m agoHN ↗

The usage allowances are now insane, like they were when the Max plans were introduced. The $100 plan is usable again for real tasks.

49m agoHN ↗

Only takes 5-10 minutes to test your favorite one shot comparison prompt.

39m agoHN ↗

If 5-10 minutes is enough, you need a more ambitious one-shot goal.

35m agoHN ↗

Can you give an example? For me I find that one shot prompts are pretty good it’s only when working with large codebases and complex, multi prompt workflows, that I find the real limitations of models

29m agoHN ↗

While I have no experience comparing this brand-new model, OpenAI themselves call it "near-Astra" intelligence. I set Astra and Opus 5.5 independently working on the same large research/coding task in an experimental project (doing NURBS surface modeling stuff). They had the same starting repo state, same task packet, same test suite to try to meet. I have the $100 plan in both.

Astra used 215% of a week's budget (I burned 2 free resets) and took 13 hours. Opus used 20% of a week's budget and took 20 hours. Both were asked to use lesser sub-agents for implementation grunt work at their discretion (Luna, Sonnet) as long as they manage and review the output.

The timing comparison is not that interesting because the wall-clock speed mostly reflects how often they ran the (large, slow) test suite, not their coding speed. Although in the past my gut feeling is that OpenAI models do generally respond faster.

The quality of their implementation was more interesting. There turned out to be a bug in one of the unit tests the agents were trying to pass. Opus interpreted the natural-language requirements from the task packet, found the test bug, and fixed it. Astra tried hard to solve the problem without altering the test suite. In practical terms Opus got much, much farther into a useful implementation. Astra was still stubbing out and faking critical parts of the implementation (B-splines) and since it ultimately couldn't pass the full test suite, finally gave up on its implementation. Astra wrote some useful tooling in the process of its efforts which I ended up integrating into Opus's version of the code, but otherwise its approach was behind.

Now, this is just one comparison in one domain, and arguably Astra's strict adherence to the tests as-given is a good thing. But Opus wasn't merely loosening the rules / moving the goalposts to pass, it spotted an actual bug, and was more successful at doing what I actually wanted. And the cost difference was Astra-nomical.

Out of curiosity for an interpretation free from my personal bias, I gave Astra a hint from Opus and permission to change the test in question, which it did, and got a bit farther, but still ultimately didn't produce a working implementation (to be fair, Opus's was not completely working either, but was closer). I then fired up fresh agents to review the two repos. Predictably, an Opus agent thought the Opus-written repo was the better basis to build on, and an Astra agent thought the Astra-written repo was the one to keep. They were not explicitly told which was which nor did the commit trailers say, but I assume they can tell. However, after doing this twice each, I saved the 4 review reports into another folder and did yet another meta-review of the 4 reports, so each would see the arguments and critiques both directions. In this meta-review both Astra and Opus converged on preferring the Opus implementation.

19m agoHN ↗

I think it’s one of the reasons why you often see people decrying the lessening capabilities of the models a few weeks later, despite there being 0 proof of any changes, and evidence of the models staying the same from sites that track it.

They form these super strong opinions after a few prompts, then face reality over time.

People have been talking about how good whatever model is at “complex” tasks since the beginning, never mind that all of those models are now outperformed by Luna which many people consider unusable for complex work.

13m agoHN ↗

They’re comparing against the previous model, not the newly released one (6.1). Why do that on a thread about the new model, I don’t know.

45m agoHN ↗

For my personal experience, antropic model have better user experience except for 4.7 and 4.8 though. 4.7 and 4.8 feels like expensive downgrade of 4.6 to me (I didn't know why these two should even exist)

However it's less willing to obey your instruction so it's less usable for general runtine flows.

7m agoHN ↗

For me Anthropic models from 4.7 to 5 including where bad and ate tokens like crazy. Task delivery was worse than GPT 5.6 and token usage was 2-3x higher.

Looks like 5.5 is the new 4.6

42m agoHN ↗

A pity I have to use claude code to try this, that I can't use the tools I know and love and have built around (opencode).

(I did use some CC for Fable when it came out, and it was... ok. Not the worst thing ever.)

32m agoHN ↗

  > Ask for something difficult from GPT-6 Sol and Opus 5.5 and watch what each one does.

That's far too vague. I found Opus to be terrific at coding, but human text just seems so robotic with it. OpenAI models used to be the prototype for robotic text, but lately I've been finding them much more natural. What is "something difficult" in your workflow?

20m agoHN ↗

You HAVE to have a set of personal evals for each class of task you want to use models against at scale so you can test plausible candidates and compare output on your work against your evals.

There is way too much subtlety in what does and doesn't work for a given problem, context/prompt, tool set and eval. I can tell you Fable is generally better than Haiku, but comparing similar tiers really does depend on your exact context.

6m agoHN ↗

OpenAI models used to be the prototype for robotic text, but lately I've been finding them much more natural.

This was the biggest thing I noticed in the 6 models; their conversational prose is dramatically less grating.

24m agoHN ↗

Suggest trying it out yourself: Ask for something difficult from GPT-6 Sol and Opus 5.5 and watch what each one does. The difference is stark.

Oh yes, I know GPT-6 Sol is ... quite not up to par. At least it's not as bad as GPT-5.6 Terra I suppose.

14m agoHN ↗

Are you comparing Opus 5.5 with GPT-6 Sol or GPT-6.1 Sol? Because they are different models.

12m agoHN ↗

This news and thread is about 6.1 Sol, not 6 Sol. You haven’t even had time to do a fair comparison yet.

45m agoHN ↗

50% cheaper cache than GPT-6 Sol will get you far more mileage on Codex.

Cache doesn't help you much when you are compacting every 5 minutes...

I was shocked at how quickly I ran out my $100/mo subscription with a single agent (sol medium).

9m agoHN ↗

you can config codex to compact at a higher context limit

41m agoHN ↗

cache is typically 10%, is this OAI setting a new level at half, 5%?

26m agoHN ↗

The backdrop being deepseek offering 1% (I remember it was ~1% when 4-pro first came out early this year - 4-pro is now removed) / 2% (current for 4.1-flash).

1h agoHN ↗

So yesterday we were consumed with how this was being delayed because of safety, yada yada.

Guess not?

1h agoHN ↗

That model was implied to be GPT 6.1 Astra, not Sol.

1h agoHN ↗

I see where you are coming from. But 6.1 Sol seems like a new frontier in pricing, not intelligence. I do think the deceleration stuff was mostly bluster, but I don't think this release in particular contradicts it too much.

1h agoHN ↗

I can blow through my weekly on astra in a few hours; hopefully this really is as good.

58m agoHN ↗

I get decent results telling it to use Luna subagents for implementing commits.

1h agoHN ↗

Cache is priced at $0.1/M, 50% as sol 6 and sonnet 5.5.

1h agoHN ↗

$2/10 is pretty cheap for a frontier model...

1h agoHN ↗

I stopped using LLMs. I shit you not. My life got better.

1h agoHN ↗

So when does Anthropic answer? Tomorrow?

9m agoHN ↗

hopefully the answer doesn't include an increase in cost/decrease in usage.

1h agoHN ↗

Why they are not even benchmark model against Anthropic or anybody ?

1h agoHN ↗

GPT 6.0 Sol was so terrible—I wonder if 6.1 Sol will be good?

1h agoHN ↗

Let's all boycott and move to Claude until they release 6.1 Astra. I don't like to be teased.

When is the alleged "safety" concern satisfied? Does this mean releasing new capability to consumers is going to get a lot slower? Lower price for 6 Astra capability via this 6.1 Sol is exciting, but that is because of Astra capability not merely the low price point.

When do we get the next jump in capability? When is 6.1 Astra released?

52m agoHN ↗

Isn't Anthropic doing the same, with Opus 5.5 being out while Fable/Mythos is still on 5.1?

48m agoHN ↗

Is this due to a similar safety concern or just because it's not ready yet for one (or more) of a myriad of possible reasons?

The coverage around 6.1 Astra seems deliberately playing into the dubious, recently headline "safety" narrative in a way that feels distinct. But you may be correct in which case, I would take the correction on board and maybe suggest a different alternative.

Although in theory if OpenAI was boycotted in this way the market pressure would force them to release. Then everyone moves back over there. Then Claude faces the same pressure. So even so, I think it could still work even if you have to trade off who you are boycotting from time to time.

Without more details on the credibility of the "safety" concern this seems like a totally coherent action for customers to take. We shouldn't put up with teasing.

42m agoHN ↗

It's just vibe versioning, right? Fable 5 is a beloved product, it gets a .1 bump to feel close. Opus 5 and Sonnet 5 had a mixed reception, they get a .5 bump to create a sense of distance.

After what DeepSeek pulled with V4.1 Flash I've given up on trying to map LLM versions to semver.

1h agoHN ↗

The real announcement is the ultra fast mode ... Astra at 300t/s is insane!

56m agoHN ↗

I am afraid to ask for the price multiplier here. And it will burn your weekly allowance not in 1 day, but in 3 hours now? Or just one?

59m agoHN ↗

didn't 6 sol just come out a couple weeks ago?

57m agoHN ↗

Other comments have already addressed this.

58m agoHN ↗

the race to the bottom on models is well underway. huge IPO's only really make sense for DC/HW lockups, and going vertical.

58m agoHN ↗

It is hard to trust these scores. GPP 6 Sol has been so bad for few days.

56m agoHN ↗

They need to fix Astra first. My main issue is with GPT in general is that unless steered it goes into building AI “sloppiness”/machinery that is not “needed”.

The good part is that this kind of behaviour also makes it good to find subtle bugs or debug issues that Fable/Claude just cannot get/fix even when you point it.

55m agoHN ↗

The GPT 6 release was ... not great.

Sol 6 was so bad that I switched over to Opus 5.5 exclusively.

Huge regression compared to Sol 5.6, often doing really dumb things. Same for Luna.

Even Astra is very unreliable for coding. Brilliant for vision, sometimes just great, but it also often does very stupid things.

I'm a bit sour on OpenAI right now and skeptical that 6.1 will be much different.

(Note: this is after preferring and shilling Codex/OpenAI models for the last half year)

50m agoHN ↗

How does this jive with the exponential growth claims? Theoretically sol models are better than the 4 series models I was using at the beginning of the year, but in practice the results don’t seem to be much better. They always nerf the models over the course of the release so it _looks_ like the next version is better but I haven’t seen actual capability growth since ~January, and I’m pretty sure that was all tooling/harness improvements.

45m agoHN ↗

Eh. What? Is this common sentiment?

I mean Opus 5.5 is absolutely fantastic, unreasonably and unexpectedly so, but Astra was great and as far as I can tell SOTA until, when was it, 3 days ago, no?

(Sol 6 idk, have not used it much for coding really. Seemed to work just fine when Astra used it in Codex as subagents)

38m agoHN ↗

On r/codex the sentiment seems to be quite wide-spread.

34m agoHN ↗

When GPT 6 Sol & Luna were released, everything went down. I have been running Sol at max thinking and it is about the same as old Luna with max thinking, give or take. Sometimes feeling even dumber. I can't trust it to do anything big alone anymore without babysitting.

34m agoHN ↗

Opus 5.5 is so good that I don't want it to be replaced anytime soon. Stop training models, Anthropic, and just serve this thing without regressions for a year or three, can you?

22m agoHN ↗

They should etch it into an ASIC. The first model worthy of that honor.

21m agoHN ↗

Sol 6 definitely feels kind of dumb and worse than 5.6

Astra seems better though.

Showing one potentially saturated benchmark doesn't necessarily fill me with a lot of confidence in the coding results.

12m agoHN ↗

In my experience, no. There’s no way to know though. The whole conversation and industry are a combo of benchmaxing, faith, and mysticism.

Since like last December I haven’t had any issues getting work done with whatever the latest Anthropic or OpenAI models at the time were. Tooling and models have only gotten better since then.

13m agoHN ↗

That mirrors how disappointing Opus 5 and Fable were, for anything beyond one-shotted tasks or shiny demos. Maybe OAI is just a step behind Anthropic? Opus 5.5 seems like the real deal again, consistent good results on large, complex codebases.

53m agoHN ↗

DeepSeek and GLM made it impossible for "sota" to price any way they want.

51m agoHN ↗

GPT 6 Sol is obsolete after only one week! I am glad that they are not afraid to update the models more frequently. The Navier-Stokes thing revealed that it took them only a week or two to train a model more capable than Astra, and I want the pace of public releases to keep up with that.

9m agoHN ↗

I had it write some code the other day, boy was it awful-looking compared to 5.6. Worked perfectly, but ugly nonetheless.

51m agoHN ↗

I got a popup in my Codex just now saying "Try out 6.1 Sol!" and so I clicked the button to try it, and intriguingly, it set my model selector to "GPT-6 Astra Light" which makes me think 6.1 Sol may be in some way just a lighter/distilled version of Astra? defo interesting, not sure if I should read too much into it though. I see no option for directly selecting 6.1 Sol in my Codex Desktop UI.

45m agoHN ↗

Astra Light is the default option in the UI, so likely a bug

51m agoHN ↗

Came here only to check if the pelican spam has made it to the top again.

51m agoHN ↗

“OpenAI's new Pro 500 plan offers OpenAI's highest usage allowance and comes with access to its new "Ultrafast" feature — it also costs $500 per month.

At the same time, OpenAI is also making its existing $200 Pro plan less appealing. In Codex and Work, $200 Pro subscribers will see their included usage decrease from 20x of what the company offers to Plus users, down to 10x of that same allowance. In ChatGPT, meanwhile, GPT-6 Pro message caps will decrease from 200 to 100 per week.”

https://www.engadget.com/2272106/openai-adds-dollar500-pro-s...

Yikes

37m agoHN ↗

OpenAI is deeply unprofitable, particularly on those pro plans.

The only way is for prices to go up. Way up.

20m agoHN ↗

It really does look like OpenAI is trying to gradually get rid of their subscription plans. Every week there is noticeably less usage available to them while each new model release boasts substantially cheaper API token pricing. If this continues then the two pricing models will eventually be at parity.

37m agoHN ↗

This is pretty typical product positioning. You want to sell to both high-end and low-end users, so you offer products at a few price points. Then it turns out that that middle is a much better fit for most users. So you start making the middle a worse fit to push most of those users into the higher tiers.

Long term, this only works if you have a non-commodity, and if the higher tier is actually more profitable. We'll eventually learn whether both are true. For OpenAI right now, it's probably enough to just increase revenue, even if the higher tier is even less profitable.

22m agoHN ↗

The 200$ plan was appealing because you got 4x usage for 2x the price.

Now, as it's linear, it makes much more sense to downgrade to 100$ OAI and pick up a 100$ Claude sub. (without doing the numbers) the usage should remain the same, total paid the same, but having access to best of both worlds. It should be a win for the user, and a loss for OAI.

With this in mind, it sounds like a fumble by OAI.

37m agoHN ↗

They're really (finally?) starting to behave like a company bleeding money.

Our VC-backed subscription days are numbered

24m agoHN ↗

I'm ok with whatever price they give out given they are not a monopoly and have competition, the lock in is minimum for me. This means they have legit reasons to send us this price plan. I don't believe they would shoot themselves in the foot when there is cut throat competition (Claude/opensource) out there.

Lastly, I'd like to actually use it in the real world to see how far my plan goes or if its unusable.

23m agoHN ↗

Alas, I did enjoy burning investor money on my taxis, movies and tokens.

36m agoHN ↗

Wow, canceling my sub. Lets see how Claude is doing these days.

I can justify $200/mo but more than double is not appealing to me.

31m agoHN ↗

Well, here is a breaking-news for you: the 20x from Claude is not a 20x on the weekly usage, it's a 20x on the 5h usage, while the weekly usage is simply double the $100 plan...

Basically OpenAI aligned with Anthropic on the weekly usage with the caveat that OpenAI doesn't have a 5h limit.

22m agoHN ↗

OpenAI 20x wasn't 20x even before that change. I got a lot more from Claude 5x than Codex 20x...

20m agoHN ↗

You are literally completely flipping reality. Codex was, in fact, 20x. It was Claude that was not 20x until they got caught.

6m agoHN ↗

I'm describing what I got from 20x Codex vs Claude 5x. Codex is just not worth the money, at least for me. What's flipped is the value you get for each of those

20m agoHN ↗

"For antitrust reasons, it’s helpful for the US government to mediate or at least enable these discussions — they don’t need to participate, but do need to issue a narrow waiver for certain kinds of safety conversations. " - Dario a couple weeks ago.

Yes, he was talking about safety, but IMHO they're likely already IMHO pushing the boundaries of cartel type behaviour. And they will use safety as the cover to make it happen.

I suspect we'll see serious price fixing and the DOJ do nothing about it because of the inroads these people have with the Trump regime.

Whether that survives contact with Chinese open weight models is hard to say.

18m agoHN ↗

You are painting half of the picture, perhaps on purpose? The other half is this: Opus 5.5 is significantly better than both Sol 6.1 and Astra, and with the newly increased limits across the board, it is quite difficult to run out (unless you're spamming agents at Max effort). So it is a much, much better deal than OpenAI's Pro 100.

11m agoHN ↗

Opus 5.5 is significantly better than (..) Sol 6.1

Come on .. this is barely released and you can already make that assessment?

And no, the $200 Anthropic plan is not significantly better than the $200 OpenAI plan, it's just the same Marketing non-sense and anybody shall now rather stick to the $100 plan of both of these provider if the monthly budget is $200. Anthropic doesn't have a Luna Max equivalent, and frankly Sol 6.1 is yet to be thoroughly tested.

18m agoHN ↗

If you've used both you know the OpenAI plans don't compare to Anthropic plans _at all_. Claude code subscriptions are probably worth 4x as much in API spend compared to the same OpenAI subscription tier.

20m agoHN ↗

According to the company, existing subscribers will keep their current limits for a time, and will later receive a one-time credit to help them make the most of their new reduced allowances

Might want to hold off on canceling and continue to bleed them dry until the nerf hits

25m agoHN ↗

I think that's by design - they're going to IPO soon so if they can get a significant percentage of users to switch from the $200 to the $500, they can 2.5x projected revenue.

23m agoHN ↗

Very risky to do so especially considering how well is opus 5.5.

17m agoHN ↗

Yeah, that's not going to happen. They are more likely to lose a lot of customers, unless Anthropic does the same thing.

But $200 is likely the ceiling of what people will pay for a subscription with usage based on vibes.

10m agoHN ↗

For consumers they may as well buy GPUs and run local models. The cost is same over a year or two but infinite token usage, they get to keep the hardware, and local models continue to improve over that time too. I can't justify $200 on SOTA models for a personal subscription after Qwen3.8-27B. And it's only getting better from here.

13m agoHN ↗

At $500 per month, it's cheaper to just buy GPUs and use local models.

50m agoHN ↗

They released GPT 6 Sol literally 6 days ago. We've accelerated to a weekly model release cadence. That seems like...a big deal.

42m agoHN ↗

It's more like they released GPT 6 Sol too early because they were under pressure and now they are releasing the real version. You cannot do anything more than minor post-training in a week.

38m agoHN ↗

Implying they don't have like 3 or 4 "models" (different quants, post training, plain renaming) on the back burner at any point in time to do exactly that

35m agoHN ↗

That seems like...a big deal.

They can release a new version every day if they wanted to. The question is whether or not the new releases provide substantial improvements or not. It's not hard to just go through the motions, bump the minor version, then make an announcement to rile up the users who don't get that none of this is standardized or regulated in any way and it's literally all made up by the company trying to sell them the product.

50m agoHN ↗

Typically how long does codex take to update with the right model metadata for the release of a new model?

{"type":"item.completed","item":{"id":"item_0","type":"error","message":"Model metadata for `gpt-6.1-sol` not found. Defaulting to fallback metadata; this can degrade performance and cause issues."}}

49m agoHN ↗

OAI and Anthropic are locked in an unwindable race. Fighting over the same pool of users while cost is growing at geometric pace.

49m agoHN ↗

Impressive improvements, but GPT 6 Sol came out 7 days ago, and this one will behave differently. The panicked pace is becoming a liability, maybe they should have waited and released this as the 6.0 release

16m agoHN ↗

They are behind, hence the panic. On top of that, Sam has been trying to do another funding round, so he's desperate to make the company look good.

Opus 5.5 was a gut punch and my impression is OpenAI is still reeling.

10m agoHN ↗

People said Astra was a gut punch and that Anthropic was reeling. (Opus 5 was almost universally panned)

The best thing is that we benefit from these constant back and forth gut punches :)

48m agoHN ↗

I guess they released this because GPT-6 Sol was underwhelming, they didn't even release it to ChatGPT. It was basically GPT-5.6 Terra for the price of Sol. However, who doesn't like price cuts? Astra for the fifth of the price? Wow, OpenAI have been quite generous recently, I still have not forgotten their 90% price cut with GPT-5.6 Luna, and now this? Astra was truly a milestone, and now they are offering similar "intelligence" for cheaper price. Incredible.

One thing I wish was better communicated is the mileage we get for our subscriptions. I do not fully understand how much usage I get with each model and their reasoning effort on 5h and weekly limit in Codex. I am asking because I know switching to Astra would consume my 5h usage limit quite rapidly, so I avoid it. If I knew how much mileage I would get from each model and respective reasoning effort, then I would be able to plan my workflow better and know when to upgrade model for a task. In almost all cases, GPT-6 Luna (XHigh) have been enough. That's why I appreciate its discount, because its dirt cheap, yet highly capable.

In other news:

In the coming days, we’ll also offer GPT‑6.1 Sol Ultrafast , with up to 8x faster token generation compared to its standard speed in Codex.

46m agoHN ↗

It’ll be interesting to see what happens to the economics of this business if we hit a wall on peak intelligence but keep finding cool ways to lower prices.

37m agoHN ↗

Sam Assman needs money to buy a new super car or private jet?

37m agoHN ↗

Do these benchmarks have any meaning anymore? And do the announcements seem less exciting now? (Not taking anything away from the advances we are making but it seems more incremental now?) The reliable way to tell if you'll like a model is reliable collage/X reviews to gauge a model's capability and then trying it out to see if you like the style.

The last time a model announcement felt like a leap in capability beyond other things out there was Fable - which was promptly taken away. Sol and recently Opus 5.5 were strong because they approach that capability with a lot more efficiency and don't blabber incoherently (looking at you Opus 5.1).

Deepseek is a workhorse for those who prefer open and API usage. Other than that the model announcements all just seem like a blur and quite interchangeable but I wonder if that's just me tuning out or do others feel the same way?

5m agoHN ↗

I fully believe that these models perform better in benchmarks versus their predecessors, but in real world usage inside of real, production codebases? They feel just as flawed as ever. I honestly have not seen any significant improvement in a few months. The last thing where I felt "wow" was `/fast` mode and Deepseek.

34m agoHN ↗

Surprisingly (or maybe not) it matches the performance of Astra on my benchmark[1], but is much cheaper. It is also head to head with Opus 5.5 on both the price and pass rate, but edges it out slightly.

1 - https://bench.killswitch-lang.org/

30m agoHN ↗

Still too expensive. Needs to come down to $1 or less to compete with china.

27m agoHN ↗

Sol is so good, honestly - the sweet spot for me. I've only ever found it stumbles when you don't give enough direction. But for idea execution - Sol is the GOAT.

24m agoHN ↗

I’m a bit disappointed with Sol 6.1. I suspect they didn't show the benchmarks and test results because it would have been embarrassing to reveal that their flagship model can't compete with the capabilities of Sonnet 5.5. That said, I still think this model is useful for a great many things, but it looks like Anthropic has the upper hand this time.

23m agoHN ↗

Ok, but we need 6.1 Luna soon. 6 feels worse than 5.6 in our agentic use case.

19m agoHN ↗

I can appreciate they show that Opus 5.5 is objectively better intel. and cost wise on multiple benchmarks

19m agoHN ↗

At the same time, OpenAI is also making its existing $200 Pro plan less appealing. In Codex and Work, $200 Pro subscribers will see their included usage decrease from 20x of what the company offers to Plus users, down to 10x of that same allowance. In ChatGPT, meanwhile, GPT-6 Pro message caps will decrease from 200 to 100 per week.”

Fuck altruism, ammi right? lets make money, gobs of it by screwing the middle users as much as we can to push them into just two tiers: Ones that use it for recreation and others that pay through their noses.

18m agoHN ↗

very surprised by the sentiment against GPT 6.0 Sol, I've been using it exclusively since release and it feels like a cheaper astra to me. admittedly I haven't tried any anthropic models in a while other than small tests since i can't use my anthropic subscription in other harnesses (like OpenAI has supported natively for a long time).

If OpenAI cuts alternative harness support it will be a weird day trying to figure out what to do next, it's been so clearly the best bang for your buck (imo) for a while. maybe id finally have to give smaller models a try.

anything to avoid using the dogwater codex & claude code tuis.

anyways this seems like a nice cost improvement over GPT 6 Sol and I expect this will be my new daily driver.

11m agoHN ↗

This is the first time I've seen praise for GPT 6.0 Sol: it's widely disparaged on Reddit and here in the HN comments too. My own experience likewise shows 6.0 making loads of silly mistakes, both for things 5.6 Sol is good at and things 5.6 Luna Xhigh is good at.

13m agoHN ↗

Hardly any comparisons to Opus 5.5, which means it's not great

13m agoHN ↗

Excited for Gemini 4 at equal or better coding and a fifth of the price of 6.1 Sol

7m agoHN ↗

3.8 flash is more expensive than sol 6.0, google uses a LOT of reasoning tokens

12m agoHN ↗

I wonder if this is the reason GPT barely works now and some chats have become not accessible.

7m agoHN ↗

There was a model called Astra-Minor, found in the files a few days ago. I assume Sol 6.1 is this, as a last minute panic rename due to Sol 6 being underwhelming while Opus 5.5 turned out really strong. I can't really explain releasing Sol 6 in any other way, especially mere days ago.

4m agoHN ↗

For most sane people, OpenAI is the way to go... A lot of usage with very good models, but you know that Anthropic is laughing all the way to the bank with Opus 5.5 being "the best" model right now... There are a ton of people (and companies) that will just refuse to use anything else than the highest benchmarking model in existence.

3m agoHN ↗

If it's so good why is Dots, also released today, based on Astra?