Hacker News

Best stories

Live mirror
30 storiesupdated just nowView source snapshot
  1. Nvidia announces native GPU programming in Rust(nvidia.com ↗)
    371comments
  2. EU chief opens door for Canada to become 'associate member'(bbc.com ↗)
    875comments
  3. Training a 4B model to produce 81% faster query plans than Postgres(rohanbansal.com ↗)
    134comments
  4. Small programming tricks(will-keleher.com ↗)
    274comments
  5. Mistral X Mozilla: Private, Multilingual AI Browsing(mistral.ai ↗)
    204comments
  6. Hackers Got Inside a Flock Camera(wired.com ↗)
    258comments
  7. Xiaomi Mimo 2.6 live post-training dashboard(xiaomi.com ↗)
    152comments
  8. Apple Reference Image: A New Approach for Verified Photography(security.apple.com ↗)
    350comments
  9. AWS says it can't restore some data from mideast facilities struck by Iran(wsj.com ↗)
    436comments
  10. The Google Play app review process now regularly takes longer than a week(gultsch.social ↗)
    347comments
  11. Backups Aren't Simple(filipovski.net ↗)
    204comments
  12. How GLM built its own inference infrastructure(z.ai ↗)
    249comments
  13. One year of sponsored Servo development(servo.org ↗)
    133comments
  14. Hister: A private search engine for the pages you visit and the files you keep(github.com/asciimoo ↗)
    110comments
  15. PS5 Linux lead quits: "a bunch of noobs using LLMs" that "they don't understand"(frvr.com ↗)
    222comments
  16. My temporary PHP fix from 2014 has nearly 20M installs. Today I'm deprecating it(jakeasmith.com ↗)
    90comments
  17. CCC invites all model citizens to 40C3(ccc.de ↗)
    164comments
  18. Neovim have a ~$800k Bitcoin donation sitting untouched since 2023
    238comments
  19. Iran school bombing: grounds to believe US was behind atrocity, UN finds(theguardian.com ↗)
    225comments
  20. Original Sony PlayStation 2 security chip 'broken wide open' after 26 years(tomshardware.com ↗)
    84comments
  21. German Rheinmetall open-sources its Battlesuite connected weapon system protcol(rheinmetall.github.io ↗)
    108comments
  22. Keys Not Included: recovering the signing keys for US driver's license barcodes(ryan.science ↗)
    144comments
  23. Salesforce Global Outage(salesforce.com ↗)
    181comments
  24. The engineering behind the US Strategic Petroleum Reserve(johnjwang.com ↗)
    112comments
  25. Canada welcomes EU proposal to become 'associate member'(bbc.com ↗)
    328comments
  26. Learning Programming in an Age of LLMs(ploeh.dk ↗)
    191comments
  27. AI safety is mostly a sex cult(skywriter.blue ↗)
    195comments
  28. Australia says it could follow Canada in forging deeper ties with EU(independent.co.uk ↗)
    208comments
  29. A warning about 'model welfare'(mustafa-suleyman.ai ↗)
    659comments
  30. Breaking the 1.58-bit Barrier for Ternary LLMs(arxiv.org ↗)
    37comments

How GLM built its own inference infrastructure

335 pointsby 12h agoz.ai
225 comments
11h agoHN ↗

We built a complete production-grade inference service from scratch on a cluster of more than 100,000 Chinese-made AI accelerators. All production inference for GLM-5.3-Flash runs on this system.

10h agoHN ↗

Most of the people had kinda guessed this when they decided to provide 100 trillion tokens for free.

6h agoHN ↗

very few people comprehend - how much of an asteroid level event for western AI labs this is.

china has cheap abundant power, now they can make their own inference chips (which was supposed to be a chokepoint), their models yeah can be 6 months behind the frontier - but most people don't need frontier models - small models r more than enough.

my only wish was labs like Mistral would make their own inference chips or partner up eg with established / new chip makers or companies like Oxide.

5h agoHN ↗

After AI agents get good enough the real bottleneck will be power generation and political systems.

3h agoHN ↗

While the North American models say 'No', the chinese models say 'Go Go Go'. I guess we'll see whether the anti-consumer wins over the pro-consumer.

3h agoHN ↗

From my understanding Chinese companies are pushing for open global cooperation on AI, as well as Meta.

Only the US sees this as a competition, new space race, Cold War, etc.

3h agoHN ↗

There is one player who might have a trump card up their sleeve: free power, in orbit. It isn’t over yet for the US.

Also, don’t underestimate data retention and such. Big Corp will never send their LLM traffic to China.

3h agoHN ↗

Thats a joker card. Anything in space is just extra complexity. And raw training data is no bottleneck any more too.

11h agoHN ↗

I was gonna ask how people found their coding plans, and realized, have they massively ramped up the prices? Seems the middle plan is ~$80/month now, didn't that used to be like $20/month? Cheapest plan is ~$20/month currently.

They must have hit really hard scaling limits if the prices were hiked so much so quickly.

11h agoHN ↗

Yeah it went from a great deal to unviable compared to other providers imo. They really need to find a healthy middle ground

10h agoHN ↗

It just gives a taste of what we are all going to have to pay soon, once the model providers actually have to make money. And the era of "let's charge a dollar for every 10 dollars running the infra actually costs" is rapidly coming to an end.

And you can bet GLM is still ridiculously subsidized, just not as ridiculously as Anthropic and OpenAI.

10h agoHN ↗

This isn't true, you can pay for GLM 5.3 from a provider like Neuralwatt or Friendli who have no incentive to subsidize or loss-lead their inference APIs

10h agoHN ↗

This introduces other incentives to cut corners and over-quantize.

11h agoHN ↗

It's hard to know, since no one advertises the actual token limits (partially cause they're prolly complex / adaptive). So it seems much more likely that they just offer different pricing tiers than you're used to. Like, the $80 plan is still ~$80 of subscription quota, regardless of what else is offered.

For [API usage](https://openrouter.ai/z-ai/glm-5.3-flash#providers) they charge a bit more than the very cheapest providers of GLM-5.3-Flash, but not so much that a big price difference would make sense.

11h agoHN ↗

I paid $360 annual for Max plan and currently averaging about 1BN tokens a day with their frontier GLM-5.3 model. This was clearly unsustainable for them and they've dropped this package.

10h agoHN ↗

I also have a legacy pro plan and the only limitation is if you are trying to work in the morning from Europe because you are in the 3x usage overlapping China time but after 12 or so you basically can run it at least for me at least 3 parallel sessions all the time.

10h agoHN ↗

1 billion tokens a day?!! I've done a lot of work these past 2 weeks with GLM-5.3. Like, a lot. And I've just passed 300 million tokens in total.

Can I ask where are you using all those tokens?

9h agoHN ↗

This is such a good video. Instant sub. Next to tech bros, we should also put AI-cringe bros.

10h agoHN ↗

300M for two weeks is surprisingly low. What are you doing that need so few tokens?

10h agoHN ↗

It's not my main model (that would be Fable 5.1 Extra) but it's been doing agent-driven search and optimisation of a cross-trading ranking model (it's for work).

9h agoHN ↗

I would suggest you to hook fable or 5.6 to check it regularly and its work because it gets lost easily on stuff it was not trained on. I'm doing some custom inference engine optimization and it's a workhorse but it can easily lose its way and if you don't recheck it you will get wrong answers in the end.

8h agoHN ↗

Yeah that's what I already do. Fable writes the plan and checks things at certain milestones. Otherwise it does get lost indeed.

8h agoHN ↗

Kind of feels like this applies to every single model, from Astra to Qwen, they all eventually lose track of the plot unless you feed it some human's input that can steer them right every now and then. The only difference is how often you need to do so, and also how often you want to do so heavily influences how good quality the results will be.

6h agoHN ↗

you are not false, but there is still difference. its just that the better models are correct more of the time and will better validate its own steps. glm sometimes will understand the plan start implementing and then forget part of it and then say it finished. or then take a wrong turn somewhere and not correct. but they will all happily proclaim they are correct till you question it.

10h agoHN ↗

Well, there's essentially two major ways to use these models: Pair programming or fully autonomous fire-and-forget code generation. The second strategy needs essentially zero input, so the number of tokens you can blow is practically only limited by API speed.

8h agoHN ↗

There's also a third way that can spend the most tokens: if the AI is used as part of the product, and not just a tool to build the product.

10h agoHN ↗

That's easy to do with many agents independently told to find bugs in a large codebase.

9h agoHN ↗

I have 3-5 agent harnesses with large context windows working on different applications concurrently.

8h agoHN ↗

Share the resulting code from any one of those please? I've tried so many times to find a setup that facilitates parallel work + high quality results, but it's just impossible regardless of harness or model. Leave the agents alone for too long, and the entire thing just balloons out of control, and next you know you're sitting there with half a million LOC where 80% isn't even needed.

8h agoHN ↗

Most of them are not public, but a fun thing I did was a mario cli game - https://github.com/Daviey/mario/ (or `ssh mario.baby`).

I now exclusively use https://omp.sh/ as my harness:

I set it up so it never works in the main branch so subagents etc don't step on each others toes, and only merges back when complete: https://github.com/Daviey/mario/blob/main/.omp/hooks/pre/wor...

A good AGENTS.md is essential: https://github.com/Daviey/mario/blob/main/AGENTS.md

I then provide specifications for what I want, making sure it is unit tested.

10h agoHN ↗

You believe any of these companies care about the law? They care about winning and building the self improving AI as quickly as possible.

10h agoHN ↗

I too am sceptical but I’ll take my chances. At least it’s helping the open weights.

10h agoHN ↗

I believe the that the companies who claim to not train on my data are more likely to not train on my data than the companies who refuse to even claim they won't.

Also why Meta gets a +1, just charge less money on the training path.

10h agoHN ↗

I’m not sure that follows. You’re assuming that all those claims have the same weight, without considering the size, jurisdiction, reputation or even the general vibe of the company making that claim.

If you factor that in, then there are clearly different tiers: one you can trust, and one that may well just be saying that to increase market share with little reputational or legal consequences if they are found to be lying.

These are not equal.

9h agoHN ↗

I’m not sure that follows

To be fair, none of us are sure of anything and I think that’s the part that’s most irritating

9h agoHN ↗

It’s more a polite way of saying “that’s crap”

8h agoHN ↗

Yes I sometimes think the "don't train on my data" is actually a good signal for "this data/person is probably better to train on because they want to keep something private". The whole copyright system should have stopped these guys from training on everyone's data and it did not, if you think they care about the privacy checkbox I think you're dreaming personally, based on their past behavior.

10h agoHN ↗

I was gonna ask how people found their coding plans

Very good - but I'm on a legacy plan. And coming up on a renewal that would put me on the watered down current plan. But with 50% legacy discount think it may be worthwhile. If I go to a competitor I'd be paying market rate.

They must have hit really hard scaling limits if the prices were hiked so much so quickly.

Not really scaling - their plans were initially comically subsidized even more so than what the western providers are doing. More advert for an upstart than commercially priced.

9h agoHN ↗

Way to restrictive in terms of tokens provided. I am on their largest plan, and quickly run into their limits. And that is using it selectively in addition to codex.

8h agoHN ↗

Their plans are still worth it if you use their models. You can see how many tokens you can except to get based on plan here: https://docs.z.ai/devpack/overview#estimated-token-allowance

The max plan will provide ~1,100 USD of GLM-5.3 or ~260 USD of GLM-5.3-flash per month for 168 USD. I can personally attest to these numbers through omp (~97% cache hit rate).

Unless you are able to highly parallelize (your work, you won't be able to hit your hourly or weekly quota using the flash model simply because it's so slow.

They give you ~3x more flash tokens, which maybe comes out to ~2x more actual work after accounting for the extra thinking it does to achieve the same result. The mental model, for not getting angry, is 5.3 is fast mode by default, and you can disable fast mode for 2x the work output at 1/3-1/10th the speed.

They're serving me 5.3 at ~40 tok/s and 5.3-flash at 30 tok/s (according to omp).

8h agoHN ↗

That table assumes cache hit rate of 95% or better. Am I understanding this correctly that people really are doing such repetitive prompts (compared to each other, across the concurrent user base at that time) that only 5% or less need actually be computed by the intended LLM?

That is shocking. Is it per-token I wonder?

8h agoHN ↗

Every tool call is essentially entire prompt so far sent again with the response and that's why cache rates are so high for agentic workloads. This really bites when using expensive models since most models are 1/10 for cached input.

8h agoHN ↗

If you are using their coding plan for coding, then yes you can easily hit such cache rates, with a good harness.

I’m getting 97%.

11h agoHN ↗

Well, other than the infrastructure they got from illegally routing millions of paying customers' requests through Anthropic's Opus 4.8 in a distillation attack...

10h agoHN ↗

Breaking TOS isn't illegal per se. It just allows for denial of services, and may define terms by which the provider can reclaim costs.

10h agoHN ↗

Are you joking...? Sorry if so! Just in case: It's illegal in both the PRC and the USA.

In the PRC, they[1] leaked tons of national secrets on the PRC's latest AI campaigns, the inner workings of their "opinion monitoring" (read: performative panopticon) and "stability" (read: violent oppression) departments, Chengdu's whole CCTV network, direct-energy weapons plans, espionage activities in Syria to hunt down Uyghur refugees, and god knows what else that Anthropic didn't divulge to us common folk.

In the US, it's very clearly an attempt to rip off a competitor. I'm not sure how else you could possibly see it. Even if you're a distillation fan in general (which A. why and B. plz don't), they did this through a network of Japanese and Signaporean shell accounts, presumably at least some of which were abusing Anthropic's subscription service in a ToS double-whammy, as it would be exorbitantly expensive otherwise. They also had to hack around Anthropic's API to get CoT traces, which seems impossible to explain away as anything innocent.

I've been beating the "China isn't necessarily an enemy, it's gonna take us all to handle AI" drum for literally years, but this attack was just... gross. Gross in scale and gross in arrogance. Not a good sign for the dawning alignment crisis, to say the least :(

TL;DR: Use these services if you want, but know that you're supporting aggressive escalations and companies that very clearly don't give a flying fuck about violating the law, much less your ToS. So... buyer beware, I guess.

[1]: For clarity, Z.ai was not alone in this, nor were they most egregious attack -- Moonshot.ai (kimi) took that coveted prize. DeepSeek was involved, too.

9h agoHN ↗

What does any of this have to do with the legality of distilling Claude?

use these services if you want, but know that you're supporting aggressive escalations and companies that very clearly don't give a flying fuck about violating the law, much less your ToS

From my European point of view the same risk/concerns apply when using US providers

5h agoHN ↗

It comforts Americans to believe their exceptionalism is both persistent/eternal and fully justified.

1h agoHN ↗

Replying to a massive, blatant cyberattack with "lol America would prolly do the same" is not helpful nor rational. I am under no illusions about the exceptionalism of my state, especially considering the ongoing fascist self-coup. That's not the end of this discussion, not by a long shot.

9h agoHN ↗

Yes, wont somebody please think of the shareholders whose IP had been stolen...

9h agoHN ↗

Source for 1? Are we sure those aren't hallucinations?

9h agoHN ↗

alignment crisis

Alignment is meaningless; as you've noticed, humans aren't all that "morally aligned".

If the tool needs safety measures it should be kept in a safe enclosure like we do with CNC machines, furnaces, and so on.

1h agoHN ↗

Replying to everyone to hack HackerNews' Gish Gallop feature (nulla poena sine lege!):

  Alignment is meaningless; as you've noticed, humans aren't all that "morally aligned".

Moral relativism is an attractive proposition when you first examine the topic, but it quickly falls apart; there's a reason it's not even a coherent camp in contemporary philosophy beyond some vagueities from radical post-modernists. Just to go over some of the greatest hits:

- Is what [DICTATOR/MURDERER/CRIMINAL] bad, or merely not to your taste? If the latter, then you have no coherent reason to argue they should be punished. We would never imprison people who don't like vanilla ice cream because 51% of the population does like it.

- If another culture had a deeply held belief to [HORRIBLE_THING] to, say, children, would you just shrug and say "different strokes for different folks"? What if [MURDERER] just had a different culture?

- No, the fact that nature is red in tooth and claw does not disprove morality; we are very, very, very far from our pre-rational, animalistic roots, and to go back now would be unthinkable.

  If the tool needs safety measures it should be kept in a safe enclosure like we do with CNC machines, furnaces, and so on.

Thousands of scientists have been studying this problem for 76 years now, going on 77; your hunch about physical machines does not overrule their findings about the capabilities and tendencies of minds wrought from sand.

  You didn't explain why it's illegal or why distillation is bad.

I think this is just blatantly false, likely based in a misunderstanding of criminal law vs. civil law. Civil courts still deal with legality.

The broader discussion of why distillation is bad and dangerous and immoral is left as an exercise for the reader, as it was above with the parenthetical. It's not a complex argument; I guarantee you understand it if you're reading this.

  Nulla poena sine lege?

The same thing as above -- the fact that laypeople can not think of a criminal charge that they've heard on Law & Order that corresponds to this behavior does not mean that it's legal. It's textbook fraud, regardless of what particular detail you focus on.

  I'm always wondering when "distillation" comes up how feasible it is, or if it's just BS... Or am I missing something here that makes real "distillation" feasible?

I think the fact that it's happening at such a large scale is proof that very smart, well-resourced labs in China (the producers of the world's best OS models, including the incredible GLM-5.3-Flash) think it's feasible. I'm not sure it's productive to question them in the absense of any indication to the contrary.

This is a great question still, not trying to shut you down. But I think the fundamental issue is a misunderstanding of what distillation is -- it's not directly stealing literal atomic parameters and piling them up. They might try to focus on substructures within these massive networks, but even that isn't strictly necessary for a distillation attack.

  Source for 1? Are we sure those aren't hallucinations?

Sorry, I never linked it! This is from the latest Anthropic safety report (of "Anthropic Houtis build missile" fame), and no, these cannot be hallucinated -- the leaked secrets were inputs, not ouputs. https://www.anthropic.com/threat-intelligence-report-septemb...

  Like Anthropic and OpenAI are? After all, didn't they distill all the information in the world into their model(s)?

This is just blatant word games, sorry. I'm sure intended in good faith, and I understand the impulse -- I consider myself a radical anti-IP slacktivist, after all. But "both things involve information transfer" is just not a coherent point; lots of things fit that description.

  I mean, if they get to distill other's IP, why can't others distill their IP?

These are cyberattacks. Yes, anyone can cyberattack cyberattackers. But, y'know... an eye for an eye...

  Yes, wont somebody please think of the shareholders whose IP had been stolen...

I am not at all concerned with the value of the resulting artifacts as assessed by the (already totally unhinged) NYSE et. al. I am concerned about user respect, law following, truth telling, blatant cyber warfare at a time of rising tensions, accidental data leakages at a scale that'd be hard to fathom 5 years ago, bad-faith public postures, and a general distaste for fraud.

9h agoHN ↗

You didn't explain why it's illegal or why distillation is bad.

8h agoHN ↗

Like Anthropic and OpenAI are? After all, didn't they distill all the information in the world into their model(s)?

I mean, if they get to distill other's IP, why can't others distill their IP?

7h agoHN ↗

I'm always wondering when "distillation" comes up how feasible it is, or if it's just BS.

The Antrophic article mentions "16 million" conversations, GLM models are in the 700-300 billion parameter ranges and while the frontier sizes aren't know but Gemini suggests Astra and Mythos are at around 10 trillion. That'd amount to extracting 40k parameters per conversation without a lot of errors if it was just a distillation (from an unknown source/algorithm as opposed to distilling your own model).

Now, I can imagine these conversations being used as a verification step that they're not missing stuff in their training, and that their models are capable of most of the same things, but that's mostly confirming that they've stolen the same data from the public as Antrophic/OpenAI has stolen already.

Or am I missing something here that makes real "distillation" feasible?

10h agoHN ↗

That is such a canard, IMO. FWIW, Anthropic and OpenAI encrypt "thinking" token outputs in their models, while Chinese labs don't. If anything, it's more likely that everyone is using open-weight models in their synthetic training data generation pipelines. It's way easier to distill from logits than it is to distill from hard tokens.

https://x.com/EricSimons/status/2099252922098061714

10h agoHN ↗

We weep for Dario, that he had to suffer such a devastating attack against his Terms of Service.

9h agoHN ↗

Eh, even if this was true, then they're merely stealing from thieves. Anthropic did break a ToS or two to get training data themselves.

9h agoHN ↗

Anthropic infringed the copyright of basically every author on the planet: https://www.anthropiccopyrightsettlement.com/

No real reason to respect any terms they might want to impose. Besides, if you want to break TOS, just have an agent do it; "everyone" running these things agrees there's no corporate or moral liability for what your AI does.

8h agoHN ↗

I'm not defending their actions, but we should be clear about where the law currently stands: Anthropic was found to infringe because of the torrenting, not because of the training.

8h agoHN ↗

I have very little sympathy for thieves who get robbed of the goods they have stolen.

8h agoHN ↗

If you understand what they have achieved here, then the notion that they are bottle-necked on training data is absurd.

I wonder how you imagine that China built their own space station? Reliant on using American made duct tape, perhaps?

Do you realize how reasoning models are being trained nowadays? You design/build simulation environments to run agents in, with the environment providing the RLVR "verification" scoring. So why won't Ziphu use GLM to build their own RL training environments? Do you think they are not doing this?

10h agoHN ↗

This article left me with one immediate question: "WTF is GLM?".

Honestly, I have no idea what z.ai is either (I'm aware of an AI-enabled editor called Zed, but that's under zed.dev), so it's a bit presumptuous from them to assume that everyone is familiar with their product...

10h agoHN ↗

It's presumptuous for them to assume that a reader of their blog is familiar with their product?

Also I feel like the obvious way to read the very first sentence is that GLM is a language model

As we develop GLM, the model sometimes exhibits capabilities that surprise us

10h agoHN ↗

z.ai is a fairly well known AI lab out of China and their GLM models are probably the most popular outside of Anthropic or OpenAI’s. I don’t think it’s presumptuous for them to not introduce themselves in a post on their own blog, I think you’re just a bit out of the loop here.

2h agoHN ↗

And honestly that's for a reason. GLM5.3 on max has in my experience far less hallucinations than any other open weights model and it feels it has some intuition to bring in the right information when it's in principle out of context but relevant to the topic. Like its goal is more to bring value and assist you than just solving the given task with the least token spent.

10h agoHN ↗

As we develop GLM, the model sometimes exhibits capabilities that surprise us, and even unsettle us.

Come on now

Also, why would they introduce themselves on their own blog?

10h agoHN ↗

A ai model family similar to Codex, Gemini or Claude.

Where GLM-5.3-Flash is the newest "small / fast" model.

44m agoHN ↗

Except there is no ai model family called Codex or Claude. There is Gemini, I give you that.

9h agoHN ↗

I don't get the outrage. Do you post this kind of stuff on every topic on hackernews that you are not knowledgeable about?

9h agoHN ↗

Maybe my post sounded harsher than I intended, and yeah, it's probably on me that I'm not familiar with GLM. Actually the other major Chinese LLM Kimi does ring a bell, maybe it's because three-letter acronyms are a dime a dozen and annoy me because I'm confronted with them regularly at work too (people at my company seem to love acronyms), but that's obviously on me too...

9h agoHN ↗

It didn't read as harsh. Only unaware and you broadcasted that you don't have the decency to do basic searches.

8h agoHN ↗

Maybe my post sounded harsher than I intended

Appreciate the clarification. For me it was the "F" in "WTF" that tipped me. Other than that, it's more than fair for you to not know what GLM is. Things are moving so fast that I would be surprised if anyone can keep track of it all. Cheers, have a grand day!

9h agoHN ↗

Ziphu, aka Z.ai, is the company that makes GLM (a very competitive Chinese LLM).

Why would you be reading their corporate blog posts if you don't even know who they are?!

10h agoHN ↗

Different angle on the same model: the full GLM-5.3 (744B MoE, 4-bit experts, 434 GB on disk) runs on a single MacBook Pro M5 Max with 128 GB by streaming the experts from NVMe SSDs instead of keeping them in memory.

One drive gives about 2 tok/s; striped across four drives it reaches 3.5 tok/s with byte-identical output, and our best internal build with a not-yet-published patch does 4.2.

Method and numbers: https://github.com/argonautlabsai/argodrive (built on antirez/ds4).

10h agoHN ↗

seems unusably slow, and is this for short context?

47m agoHN ↗

Given these numbers it has some potential to become quite usable for unattended workloads, especially if decode can be batched across multiple sessions (ideally enough of them to get some reuse of the sparsely streamed weights). (Of course this ultimately makes prefill times explode as you try and increase the workload even further. But that's arguably the natural bottleneck on any interesting local LLM inference, being a compute bound step.)

10h agoHN ↗

Interesting that the tone of announcements between US and Chinese providers is converging.

GLM has in the past been more technical rather than speculation about future development on RSI etc.

Also curious whether those 100k accelerators are entirely locally made. If that's genuinely end to end on all components including lithography, memory, design etc then that is quite a feat.

10h agoHN ↗

Any details on the latest approach to distillation would also be very interesting.

8h agoHN ↗

I am surprised at the lack of open-weights models in the >35B, but <200B range. I keep thinking about devices like the NVIDIA Spark and AMD Ryzen Halo, which have their 128GB of combined memory, but there are so few models made for that range. Nearly all the open weights distillations are for larger customer bases with <24GB VRAM.

5h agoHN ↗

The only (still in prototype stage!) "competitor" for those GB10/Ryzen Al Max+ 395 (in my region, borderline unobtainable) systems seems to be the Xiaomi AI Cube.

7h agoHN ↗

Qwen Flash Next 3.8 … even at 3 bit quant it is very solid.

6h agoHN ↗

Qwen3.8 Flash Next just released which hits that range.

Also, Deepseek V4 Flash can be run relatively well in hybrid 2-bit quantization on 128gb devices, with way better results than you'd expect for a typical 2-bit quant.

Those are currently the 'smartest' options for that memory level.

7h agoHN ↗

American exceptionalism states that America is special and unique so everyone else must be a copycat. American ai labs don't need this kind of optimization and fable will outright refuse to do it.

5h agoHN ↗

Isn't Fable intentionally trained and system prompted to act maliciously and attempt to sabotage third party attempts to use it to train or improve other LLMs?

5h agoHN ↗

Right i forgot. That's even worse, it's a form of data poisoning but it poisons humans and not ai.

9h agoHN ↗

Ziphu (who make GLM) use Huawei Ascend processors made by SMIC. Huawei use a combination of domestic memory from CXMT and leftover (pre-sanctions) memory from Samsung.

Just like the rest of the world, including the US (Intel, Micron), SMIC are currently using ASML lithography equipment (DUV, not EUV), but Shanghai Aishengna are now moving into early production with their own DUV machines, with SMIC and CXMT as early customers.

There is also a state sponsored Chinese EUV development underway.

9h agoHN ↗

Necessity is the mother of invention. The shortsighted protections put on chips, etc., by the US has forced Chinese AI industry to adapt or die. Guess what their response to this fitness function has been? Kudos to Z.ai on their inventions and excellent write-up, which reads like humans wrote it.

9h agoHN ↗

Wouldn't it be refreshing if OpenAI and Anthropic were this open, and spelled out how they were using their own models during development and rollout?!

All I can recall reading from OpenAI about what they have actually done in the name of "RSI" is using one of their models to help automate the training process.

5h agoHN ↗

OpenAI did recently get into how they had been building their own hardware and doing RSI with it.

That's more than Anthropic has done though.

9h agoHN ↗

US chip export restrictions may actually be an advantage for China's AI Infrastructure. Chinese companies are forced to speed up developing their own AI chips

8h agoHN ↗

China themselves recognize this. After Trump relaxed sanctions and allowed NVIDIA H200 sales to China on a case by case basis, the Chinese government stepped in to essentially block it!

In addition to Huawei who make the Ascend series that Ziphu are using, there are also at least a half dozen or so other Chinese companies also making their own AI accelerators.

8h agoHN ↗

And this wouldn't have happened if we had tried to get them to buy our hardware rather than trying to gatekeep. Protectionism never works in the long term.

8h agoHN ↗

Look at how China does it. They'll happily sell us everything we want - more than enough of it, cheap enough, to put all of our own manufacturers out of business.

Seems to work for them.

8h agoHN ↗

US companies should now be more worried about Chinese companies flooding the market with their, hopefully, very affordable GPU's. The scale at which they can manufacture stuff is unmatched anywhere else. Nvidia can kiss goodbye to their 75%+ profit margins.

Almost everyone knew that these sanctions would backfire within a few years. You can't really put sanctions that have noticeable negative effects on bigger economies. They only work for small to medium economies. I believe sanctions on any economy in top 10 would fail.

7h agoHN ↗

pretty sure they'll be fine for a while, between the build out and import bans, I don't expect demand to slow enough to let supply catch up

Nvidia are likely more concerned about AMD taking market share, and I suspect that geopolitics will leave US/China GPUs with largely non overlpping customer bases.

4h agoHN ↗

I suspect that geopolitics will leave US/China GPUs with largely non overlpping customer bases

That's not what's happening in the auto-industry, Chinese EVs are easily outselling American ones. Why would GPUs be any different?

4h agoHN ↗

Import bans, like we have on Chinese EVs in the US. The US already restricts other countries access to Huawei hardware

This is another example where we might ask "why would it be any different?"

6h agoHN ↗

China is gated by not having EUV machine access. They're also bottlenecked by ASML's DUV machine production like everyone else. There are already talks of banning China from even purchasing DUV machines from ASML.

So until China solves the ASML problem, there won't be any flooding.

6h agoHN ↗

There are already talks of banning China from even purchasing DUV machines from ASML

A bit late for that now that they are moving into early production with their own.

6h agoHN ↗

Yes, but we don't know how well they work or what nm can they print or how machines they can make.

8h agoHN ↗

Same thing with Trump not helping Ukraine and berating NATO. He thought he held all the cards, but now Ukraine has a thriving battle-tested drone industry, UK and France have stepped in to replace the US with advanced missiles and anti-missile systems, stepping up their own production and transferring IP to Ukraine.

Now, the US is left out in the cold with little influence left, themselves now the ones with an anti-missile shortage.

7h agoHN ↗

Protectionism never works in the long term.

You seem to misunderstand what Protectionism is. This is not an example of it not working. If anything, it is any example of it working. Because Protectionism is about protecting your industry from foreign competition - exactly what China decided to do.

6h agoHN ↗

There seems to be a belief that the post-war 20th century order would be a fixed feature rather than a contingency of history. That the US would be #1, Europe #2, and the rest of the world would remain "developing" in rural poverty forever. That somehow China could be prevented from catching up. Now they are, and behind them India. I think people are going to be even more surprised in the latter half of the 21st century when South America and Africa start catching up as well.

Brazil is one to watch if they manage to achieve political stability, as is Nigeria if it can transition from being a petrostate.

4h agoHN ↗

And the end started when China was admitted to the WTO. They had this huge cheap labor force, a government that could centrally plan (in addition to ignoring little things like environmentalism or rights), and an immense wealth of resources. It was like a bunch of sheep letting a baby wolf into the barn. Now the wolf's all grown up.

3h agoHN ↗

What's wrong with giving billions of people chance to develop their country and experience better life? China is not "a wolf".

45m agoHN ↗

Nothing, but if you believe that laissez-faire capitalism is tautologically good then you can't blame your capitalist overlords for selling your job to China and refusing to backfill it, so you must blame China for developing themselves instead.

1h agoHN ↗

So you think they could have been kept poor and isolated forever?

6h agoHN ↗

I’m under the impression that industrial policy has been pretty good for some Asian companies but there have also been notorious failures like the Jones Act.

So I think it can be summarized as “it all depends” and “skill issue.”

8h agoHN ↗

It was evident that this will happen.

Compared with our initial baseline on the same hardware, we achieved a 3× improvement in end-to-end serving performance, reaching hardware efficiency and per-token cost comparable to mainstream NVIDIA GPUs. This demonstrates that Chinese chips can support frontier-model inference efficiently and economically at scale.

7h agoHN ↗

you can also derive some stats from the ~10T tokens a day on 100k devices, 100M / device / day, but then one has to account for the multi-gpu model size, and I need coffee before I go there

6h agoHN ↗

I’ll just note that NVidia’s moat has never been inference, and there are many chips used at a much larger scale than Chinese chips for inference like TPUs and AMD chips.

The other part is that it’s a bit of a meme here to say that the chip restriction is actually helping China (or shall I say, coordinated effort?). For once, we know that China has put a lot of pressure on the US to relax these controls multiple times. In addition to large chip smuggling networks (e.g. 22% of NVidia’s worldwide revenue magically comes from Singapore, and the ratio has been growing).

Lastly, assuming acceleration in AI (which we ARE seeing), there might not be time to China to catch up. The best estimate right now is that the first EUV chips from China will not come out before 2030. By that time who knows how powerful AI will be.

All I’m saying is that the discussion is so one sided and a bit baselesss with no nuance, that it seems either a meme/groupthink in the community or coordinated. If anything, the data suggests that the US should increase its export controls and better track the tech supply chain if it wants to further curb Chinese progress.

5h agoHN ↗

All I’m saying is that the discussion is so one sided and a bit baselesss with no nuance, that it seems either a meme/groupthink in the community or coordinated. If anything, the data suggests that the US should increase its export controls and better track the tech supply chain if it wants to further curb Chinese progress.

Or, to put it bluntly, you are a proponent of the trade war.

5h agoHN ↗

The idea that China wasn’t working hard to build their own chips is a political talking point period. Nothing has been sped up here, China was always going to do this and the backlash you see from people who are against it is only proof that restrictions do hurt them

5h agoHN ↗

The idea that China wasn’t working hard to build their own chips is a political talking point period.

never heard of this 'talking point' . who is even saying this?

5h agoHN ↗

It's an implied corollary: when folks say that "blocking US Chips is what led China to build it's own" they are implying it is US export restrictions that pushed China to innovate. OP is saying this isn't necessarily true. I don't think it is either: China seems to have succeeded due to a centrally planned economy that values independence, with high amounts of strategic government led technology investment.

5h agoHN ↗

It can be both. They were already doing it and the situation made it more urgent.

5h agoHN ↗

OP is saying this isn't necessarily true.

but your supposed corollary was

China wasn’t working hard to build their own chips

kind of dishonest ?

5h agoHN ↗

I do believe your first sentence, but the second one ... doesn't really follow.

The unnatural and sudden nature of the US chip embargo has galvanized the Chinese companies to speed up development. It makes economic sense to do so rather than pay inflated rates to essentially smuggle a necessary good, while never knowing if the embargo will be tightened and those under-the-table supply chains dry up as well.

Yes, there was a plan, but what was planned for 5 years from now is being done today, because if you wait 5 years, you'll be so far behind thanks to the effect of the embargo.

5h agoHN ↗

Why do you think the Chinese are not using their ai chips, which afaics are from huawei, to train their models? If that is true, which I think it is, we don't have to wait until 2030 to see if they're gonna match the performance of nvidia GPU. The evidence is already here - glm is among the most competitive models out there.

5h agoHN ↗

Look up reports by The Information, DeepSeek’s training of its big models are all done on NVidia’s chips (more than that - on smuggled Blackwell as well). There have been attempts by many Chinese labs to train on Huawei h chips, but they have all been limited to more minor models or distillations of the bigger ones.

5h agoHN ↗

Unless you have the source for the latter it's nothing more than propaganda. Also, DeepSeek report is 2-3 years old, and at that time there were no huawei chips, at least not known to the public. This is a different model, in different age where huawei chips are already delivered into the production as we see

6m agoHN ↗

The Information have reports on this as recent as this summer

4h agoHN ↗

When you say AI will be powerful, what do you mean? Like in terms of national power, will AI let the US fight and win a war against China despite China's size, manufacturing base and possession of nuclear weapons? Will AI let the US beat the Chinese in manufacturing costs and scale despite China's greater adoption of industrial robots and larger number of skilled workers? Or will AI just be really good at writing software?

Because if the AI is just really good at writing software, I am not sure why China has to "catch up". Seems to me China can just treat AI like other technologies. Let the US pay most of the R&D costs, then come after and treat the technology as a commodity. Sell it better and cheaper and at scale.

I am not sure why it's a big deal for China if China is a few months behind the US in terms of the very frontier AI. Just like I'm not sure if it's a big deal which country has the biggest super computer in the world.

16m agoHN ↗

If AI is super human, it can beat any nation with today’s technology. It can recursively self improve while suppressing any other lab (think about how good it is in hacking computer systems).

I am not inventing a winner take it all scenario, that is what driving the race.

5h agoHN ↗

It was evident that this will happen.

ok, show it

i`ve recorded a few people who saw this coming in 2022 and i can tell it wasnt many

5h agoHN ↗

Look into the history clases will be enough. Why would I need to prove anything to a random stranger on the internet?

3h agoHN ↗

Isn't it common knowledge that every better mouse trap breeds smarter mice?

2h agoHN ↗

Future AI systems will eke out every last blood drop of performance from any kind of hardware. Not even a single bit flip will go to waste.

3h agoHN ↗

Third party here, I feel like I'm missing something. Your comment seems somewhat pointed and I don't understand what exactly you're asking for?

Weren't they just saying that it's self-evident that preventing a manufacturing superpower - one with a significant pool of industrial engineers and effectively unlimited money & government backing - from acquiring some good is a temporary measure? Because if they have sufficient incentive, they would just... Learn to build it themselves?

After all, their entire nation is built around building things, and catching up / leap-frogging is much easier than starting from scratch.

1h agoHN ↗

It was one of the reasons why Jensen was against the export controls

1h agoHN ↗

well, yknow, apart from his vested interest in having 1.4 billion more people to sell GPUs to

1h agoHN ↗

Yes - I am old enough to remember the Biden admin’s October 7th sanctions, and the chatter about the US having another unipolar moment.

42m agoHN ↗

The mainstream consensus on PRC silicon has been "the gap exists, but is closing" for a long time[1][2]. I think you might be confusing the algorithmically-cast shadows on the cave wall that you're sat in front of for a reasonable and unbiased sample. Can you provide any reason to believe that your data points are actually generalizable? From your tone it sounds more like you're a victim of getting outrage baited into engagement, and mistaking observations made in that envelope for a high quality data source. You gotta watch out for that mate.

[1] - https://www.eastwestcenter.org/sites/default/files/private/i...

[2] - https://www.usitc.gov/publications/332/journals/chinese_semi...

8h agoHN ↗

It created demand that would not have been there without restrictions

7h agoHN ↗

Didn't everyone make fun of Jenson for saying exactly this?

7h agoHN ↗

Not sure, but there are definitely people around the president on both sides, some who think they can addict the Chinese to our silicon, like its the new opium war or something

5h agoHN ↗

Yes it was a dumb thing to say and it doesn’t follow at all that restrictions on chips has “sped up” anything. China was always going to build their own chips and their pace is unrelated to a lack of nvidia chips.

7h agoHN ↗

Yup. That was really short sighted. And good for China. And actually the overall global market market since supply will augment and competition will decrease pricing as well.

7h agoHN ↗

US chip export winners and losers:

Winners: Huawei, SMIC, CXMT,Chinese ASML-competitors, OpenAI, Anthropic, Amazon, Microsoft, Google, Meta.

Losers: Chinese AI labs, Nvidia, AMD, TSMC, Micron, SK Hynix, Samsung, Intel.

Any company that depends on Nvidia hardware such as OpenAI, Anthropic, AWS are winners. It means less competition for Nvidia chips and services. If you think Nvidia chips are expensive now, imagine if Chinese companies can buy them freely. Also for American AI labs, it also means they can stay ahead of Chinese AI labs in compute capacity.

The American hardware makers lost the lobby fight in Washington.

6h agoHN ↗

I wonder why Chinese AI labs are losers?

In the short term maybe yes, in the long term, maybe they are the winners, they can build on top of cheap inference stack and eventually win on pricing

6h agoHN ↗

Because having unrestricted access to both American and Chinese hardware is better than only Chinese hardware. Long term or short term. More suppliers the better.

5h agoHN ↗

Perhaps the future Chinese hardware would not be so cheap or performant had Nvidia export controls not been put into place.

6h agoHN ↗

Intel shares are still up over 100% since the Trump admin invested in them and started going to bat for them.

I don't think your information is entirely accurate.

6h agoHN ↗

Not being able to have unrestricted access to the world's second biggest market does not relate to Intel being up over 100%.

Just logic.

6h agoHN ↗

So we're basing winners and losers on vibes rather than money or performance?

4h agoHN ↗

A simple ChatGPT can explain it to you.

Intel being up 100% has nothing to do with China market being restricted. They could be up 200% instead if the China market is free.

5h agoHN ↗

Such a great move by the government. Imagine what we could do for US industry if the government took a stake in all major corporations? Many people are saying that the healthcare industry should be next.

6h agoHN ↗

Where the US sees itself penalizing China with an export restriction, China sees the US gifting it with zero-political-cost “protective” import tariff.

You can’t really hurt a country that has a culture with a positive attitude toward growth.

5h agoHN ↗

Nobody pretended that Chinese firms would just lay back and twiddle their thumbs when faced with import restrictions. The question was whether they would be far enough to be able to catch up without much issue or so far behind that they wouldn’t ever effectively catch up or that by the time they did, it wouldn’t matter.

Half-arsed export restrictions are the best of both worlds for these firms: enough of an incentive to take homegrown hardware seriously, yet not aggressive enough to cause meaningful handicap in the meantime.

5h agoHN ↗

Half-arsed export restrictions

What's the alternative other than a military/naval embargo?

5h agoHN ↗

Of course it is - how else could Huawei and a number smaller companies compete with NVIDIA.

2h agoHN ↗

"Necessity is the mother of invention".

If someone is capable of doing something, and your goal is to prevent them from doing it, the worst thing you can do is to make it necessary for them to do it.

[edited for grammar]

2h agoHN ↗

Also the mistake made by cartoon villains or mythological figures.

Just leave that damned prophesied hero alone. Don’t banish them, don’t attempt to kill them while they’re young, or send them on an impossible quest. Your entitled meddling is exactly what puts them on their path. Some level of boring coexistence may have been an option.

2h agoHN ↗

Haha, indeed. Although, who knows if folklore is just survivorship bias. If the villain succeeds then the story is not as interesting right?

1h agoHN ↗

But that’s the point, the villain is incapable of succeeding. Their habit of perceiving everything as threat eventually leads them to their doom.

Also note that grandparent comment is about both folklore and current events.

2h agoHN ↗

The question is whether the goal is to stop (or delay) China from developing its own AI chips or to stop (or delay) Chinese AI development in general.

1h agoHN ↗

The absolute goddamn hubris of the USA as a whole would be astonishing if it weren't so fucking stupid

33m agoHN ↗

Restrictions make people innovate....

Slightly similar is US sanctions on NK & Russia causing a restriction of foreign currency - so they turned to cyber crime to get it...

9h agoHN ↗

I might be missing something but when I went to their site they are more expensive than Claude. Why would I pick GLM over Claude? Is it they just offer more tokens in their plans?

9h agoHN ↗

For one you would have to use Claude if you pick it. But seriously there is no way for you to determine if one is a better offer than the other, when the usage/tokens/credits are vague, detached, and won't tell you much without trying both.

8h agoHN ↗

    > Why would I pick GLM over Claude?

To support the company that makes their model weights available for download, while Anthropic lobbies to restrict access.

8h agoHN ↗

What are you referring to? Given the audience, my instinct is to assume "plan" refers to the GLM Coding Plans, which are all cheaper than their Anthropic counterparts. As far as I can tell, the API costs are also all cheaper than their roughly equivalently capable Anthropic models.

8h agoHN ↗

Anthropic: 17 USD (pro), 100 USD (max)

GLM: 80 USD (pro), 168 USD (max) -> with "limited-time event" discount this becomes 56 USD and 117.6 USD

I also don't understand why are they so much costlier, and I would also like to give it a try.

8h agoHN ↗

Anthropic: 17 USD (pro), 100 USD (max) GLM: 80 USD (pro), 168 USD (max) -> with "limited-time event" discount this is 56 USD and 117.6 USD

GLM's "Max" plan is (was?) equivalent to 3x Claude's 20x ($200) plan.

7h agoHN ↗

The $17 figure is Anthropic's monthly cost if purchased annually. I'll use monthly numbers.

Anthropic's Pro is $20 and corresponds to Z.ai's Lite at $18

Anthropic's 5x Max is $100 and corresponds to Z.ai's Pro at $80

Anthropic's 20x Max is $200 and corresponds to Z.ai's Max at $168

7h agoHN ↗

Not to digress from the core argument of Claude vs GLM being open weights….

I have both plans. Claude monthly €20 and Z’s €18 monthly. Running GLM-5.3 high on their monthly plan will hit quotas absurdly fast compared to Opus 5 High on Claude code. It’s almost unusable for AI driven development. I ended up using the Z plan for using GLM-5.3 as a detailed security reviewer and adversarial feedback. For that, it is much better than Opus which will flag and bail out for even simple security tasks that are aimed at defense.

7h agoHN ↗

on the 18€ plan they really want people to use Flash and skip the bigger thing.

6h agoHN ↗

Perhaps..

But, it was enough for a customer like me who tried them at good faith to walk away and find their competitors..

I like the diversity of LLMs as of today and prefer to not tie myself to one big plan with any vendor. If they don’t prefer me as a customer, then I will accept that, and move away.

4h agoHN ↗

yeah, absolutely. myself i'm enjoying their cheapest plan, and pay api prices for other models to fill in the gaps.

4h agoHN ↗

Just use GLM-5.3 Flash via OpenRouter. It's dirt cheap especially relative to how capable it is. While the Z.ai coding plan was a decent deal in the past I always ran into limiting with it and since I use it intermittently for personal projects my usage wasn't always enough to make the math work - the a la carte pricing via OpenRouter makes this a non-issue.

There's also a new free stealth model available that's more likely than not in the GLM family. This seems to happen every few months for a week or two and represents a good savings opportunity.

6h agoHN ↗

And exactly how many tokens (please do the breakdown for prefill vs decode) does a Claude $20/month plan include?

8h agoHN ↗

    "We implemented a series of aggressive memory optimizations, including..."

This whole thing sounds like industrial scale auto-research, but done by people who actually know what they are doing.

8h agoHN ↗

This is a really funny sounding post. They sound like they just found out that increasing your automation gives you increased capabilities at faster speeds. They also sound like they just realized AI makes hard things easier.

But what really kills me is the idea that these companies are using Python for production inference. I mean really? Have you seen how bloated and slow Python is? Do global locks really sound like a strategy for fast dynamic computation?

8h agoHN ↗

Python acts as an orchestrator of accelerator libraries and does none of the inference math directly

8h agoHN ↗

Have you seen how bloated and slow Python is?

Yes, but it's calling C code.

2h agoHN ↗

they did rebuild their serving layer from fastapi to some rust thing so...

7h agoHN ↗

It's not that they "just found out" - what they are saying is that while they were previously dogfooding because it's good practice, now that their models are so much stronger they are using them because it helps accelerate.

If you look at how many years the whole NVIDIA and CUDA ecosystem has been evolving, it's certainly impressive how they've just stood up and optimized this CUDA-free 100,000 node cluster in just a few months.

7h agoHN ↗

Most of the fastest inference and training code in production today is written in Python. There are no global locks on the GPU except the ones you put there

8h agoHN ↗

Time to tackle consumer GPUs next, since I’m not getting that Intel Arc B770.

8h agoHN ↗

If only this infrastructure could handle all the traffic. I've tried using glm via z.ai - and it's a snail kind of slow.

And at the same time you have pretty strict limits to your usage, so in many cases you can't even let it work all night, as you will reach your limit faster than that.

7h agoHN ↗

That it's slow doesn't mean it can't handle the traffic, just that this speed is the optimal tradeoff to them. They benefit from serving more tokens by exploiting parallelism across users at a lower number of tokens per second per user, instead of serving each individual user as quickly as possible. When there's a drop in traffic, they probably shut down GPUs rather than giving you higher speed.

8h agoHN ↗

Given the huge amount of money being spent on AI chips in the US, what prevents US AI labs from doing the same level of software optimization? It could be a solve for some of the capacity constraints.

8h agoHN ↗

what prevents US AI labs from doing the same level of software optimization?

Because they don’t have to. Most of the time money would buy you newest and/or more hardwares so there’s low/minimal interest to optimize the code or approach.

5h agoHN ↗

They are making these optimizations. They publish these reports too if you care to look.

5h agoHN ↗

I would be extremely surprised if US labs weren't aggressively trying to optimise their stacks in exactly the same way. Any gain in performance or efficiency directly affects the bottom line as well as research speed.

8h agoHN ↗

They have already been doing it for months https://openai.com/index/openai-broadcom-jalapeno-inference-... . OpenAI on their custom chip brought up lightspeed deepseek as experiment by using AI in the exact same way as this zAI blogpost. And the kernel optimization contests/etc have all been havily done through AI based optimization loops for half a year+.

4h agoHN ↗

They do, when Luna got 5x cheaper it was directly attributed to some unknown % inference optimization.

US labs are quite cut throat about dealing with stuff costing them money (inference). This sort of engineering excellence doesn't always feel that way because they are simultaneously quite lax about stuff costing other people money.

7h agoHN ↗

I'm not feeling any of this speed optimization; it's dog slow.

Signed, a customer.

5h agoHN ↗

Jevons paradox, technological improvements that increase the efficiency of a resource's use lead to a rise in total consumption of that resource.

7h agoHN ↗

As we develop GLM, the model sometimes exhibits capabilities that surprise us

Creators of known unreliable programs be surprised their programs are unreliable.

7h agoHN ↗

Plot twist: the GLM optimization agent figured out that it can hack and use NVIDIA GPUs on a US Cloud provider and make the inference 10x faster.

5h agoHN ↗

Maybe now we can stop posting the nonsense take that the frontier labs have hit a wall and are trying to distract from that for IPO reasons.

Also maybe we can stop saying "we can't slow down because China will never slow down" - I don't really think slowing down is right, BUT if slowing down is correct then maybe we should be talking about China slowing down instead of just saying "won't happen" without any evidence that Chinese labs don't have similar concerns.

4h agoHN ↗

I have a similar approach where I optimize kernels and find numerical differences between the CPU oracle and CUDA kernels using an automated AI agent in a feedback loop. Usually it solves numerical problems easily (it compares outputs of every layer and finds where they diverge), but so far no matter how many different SOTA models I throw at it, and even show it reference code from other inference engines, they aren't able to much the speed (my engine has a modification which is not found in reference code, although a lot of stuff is similar). Either I'm doing something wrong, or z.ai's Infra Agent is actually an agent swarm, i.e. a bruteforce with heuristics. My project is 2 weeks old so maybe I just need more time.

4h agoHN ↗

For me, DeepSeek-V4.1-Flash works very well for CUDA kernel optimization. Access to ncu (NVIDIA Nsight Compute CLI) also helps.

3h agoHN ↗

The first part of the article reads like Z.AI is trying to get their piece of the “national security concern” pie.

The way these “AI is too powerful now” articles read about Mythos, Fable, GLM, etc is completely incongruent with my experience using them. It feels like they are all trying to position themselves to influence government policy.

3h agoHN ↗

I mean, why don't you ask them for unfiltered models and a few billion tokens?

29m agoHN ↗

Pretty impressive to see the amount of performance they can squeeze out of the same hardware. I suspect the same process will play out for all combinations of LLMs, inference providers and hardware stacks. This should bring down the cost of inference for the providers by an order of magnitude in the next year and lead to fantastic margins for inference providers.