Hacker News

Top stories

Live mirror
30 storiesupdated just nowView source snapshot
  1. Cloudflare Quick Tunnels(cloudflare.com ↗)
    179comments
  2. Saving another 100TB of RAM with math (and Rust)(cloudflare.com ↗)
    1comments
  3. Android 17 is the first since 3.x to add new APIs without releasing to the AOSP(grapheneos.social ↗)
    2comments
  4. Photon-Emission-Guided Laser Fault Injection Enables RP2350 Secure Debug(ledger.com ↗)
    25comments
  5. Show HN: Cactus Needle 3: 8-29MB automation models can match DeepSeek V4 Flash(cactuscompute.com ↗)
    50comments
  6. Apple releases iPhone Duo simulator and Xcode 27.1 beta(developer.apple.com ↗)
    1comments
  7. US Military had close call after using AI for hallucinated intelligence report(cnn.com ↗)
    134comments
  8. OpenJev(openjev.com ↗)
    228comments
  9. North Korean nuclear test sets off years of earthquakes(science.org ↗)
    113comments
  10. The Implications of Linguistic Illegibility for LLM Security(arxiv.org ↗)
    3comments
  11. Cache-to-Cache: Direct Semantic Communication Between Large Language Models(arxiv.org ↗)
    discuss
  12. Systemd is a suite of basic building blocks(systemd.io ↗)
    33comments
  13. C++26: Trivial infinite loops are no longer undefined behaviour(sandordargo.com ↗)
    131comments
  14. Show HN: Ax-check.com – Can agents use your product?(ax-check.com ↗)
    14comments
  15. I vibed a proof of Conway's conjecture(overreacted.io ↗)
    141comments
  16. A heap overflow and SSO misconfiguration to compromise OpenAI internal repos(hacktron.ai ↗)
    189comments
  17. The first new cat species discovered in 100 years(nationalgeographic.com ↗)
    6comments
  18. Inside ZCode: Silently uploading your Git history to the cloud(ferstar.org ↗)
    84comments
  19. Our brain evolved from two primitive nervous systems that merged: Study(newscientist.com ↗)
    12comments
  20. A search-and-inference database from scratch in pure Zig(antfly.io ↗)
    4comments
  21. Cekura (YC F24) Is Hiring(ycombinator.com ↗)
    discuss
  22. Mathematicians Build Long-Awaited Graph Sandwich(quantamagazine.org ↗)
    11comments
  23. How SpaceX streamlined the Raptor engine(construction-physics.com ↗)
    3comments
  24. Minimal Phone 2(minimalcompany.com ↗)
    74comments
  25. Border agents can search cellphones without a warrant or reasonable suspicion(lawandcrime.com ↗)
    28comments
  26. Warez: The Infrastructure and Aesthetics of Piracy (2021)(archive.org ↗)
    4comments
  27. How to Write with an LLM(sockpuppet.org ↗)
    207comments
  28. Show HN: Scry, programmable internet search w/ congestion pricing(scry.io ↗)
    11comments
  29. Jemalloc 5.4.0(github.com/jemalloc ↗)
    79comments
  30. I don't like passkeys(hawksley.dev ↗)
    633comments

The Moat or the Commons

47 pointsby 4mo agowarman.life
27 comments
4mo agoHN ↗

DeepSeek, Qwen, Kimi, GLM — running on the LangChain, vLLM, llama.cpp, and Ollama stack

"running on the LangChain" ??

EDIT: look, I think the general discussion is important, so I don't want to denounce the article. I, for one, am excited for better control, ownership, and accessibility of models. The ride labs take us on can be quite frustrating. Maybe there's even signal that the model progression is stalling (ie Opus 4.7). If that's true, then some of the notions made in the article are important to discuss. Ref https://x.com/ClementDelangue/status/2046622235104891138?s=2...

EDIT: this is not a complaint about the grammar. Look at my reply in the comments.

4mo agoHN ↗

Running on the stack consisting of Langchain etc

Yeah not sure about both ollama and llama.cpp though lol

4mo agoHN ↗

There was a bigger opportunity here to mention OpenCode, Pi, etc - open Harnesses that provide accessibility to the oss models, a platform others can build around, and something enterprises can adopt in reliable ways; for the most dominant use case of AI today.

4mo agoHN ↗

more than that, its pretty clear that there is an insane underinvestment in the harness layer. ive been iterating on my own ideas in that area through the lens of increasing reliability. and holy crap is there so much low hanging fruit. i literally can’t figure out a sustainable way to do the work without commercializing at that layer

4mo agoHN ↗

"the" is connected to "stack", not "LangChain". "LangChain" is a adjective that modifies stack.

Cutting it off at "the LangChain" is like if I took the first sentence of your edit and said "look, I think the general" ?? You think the general?

4mo agoHN ↗

I wasn't complaining about grammar. I didn't even notice that.

4mo agoHN ↗

I think it's good to note that there are scenarios where the trend of open models keeping up will not continue forever.

AI development speed is increasingly influenced by the quality of the model you are able to use internally. The frontier labs could easily pull ahead again if they increasingly withhold their best models (e.g. Claude Mythos) from the public entirely. They will benefit from increased R&D speed internally that cannot be matched by open labs.

Also, it seems plausible that frontier labs will eventually accumulate architecture-level improvements in their models that make them significantly more efficient than open LLMs. In that case, if the open labs cannot reverse engineer and replicate that, then their models will forever fall behind. On the other hand, any architectural innovations in the open model space can be used freely by closed frontier labs.

There's a counter-trend that favors open models however, which is that the design of model harnesses and multi-agent systems is more and more important to AI quality today relative to the intelligence of the model itself. This MIGHT mean that having a bunch of dumb, but cheap models in the right harness can actually compete very well against raw frontier intelligence in most practical tasks. (In other words, better harnesses makes models more efficient at improving their task completion by spending extra tokens.) This would give cheaper open models an advantage in any task where they're smart enough to complete at all, since a good multi-agent harness might mean they can do these tasks reliably and typically for cheaper than frontier models even if pushing a higher number of raw tokens.

4mo agoHN ↗

This seems like a wildly unlikely risk. Innovations in this space are just mathematical ideas, easy to write down in a paper and replicate.

It’s much more likely that performance will plateau and open weights will catch up asymptotically

4mo agoHN ↗

It’s much more likely that performance will plateau and open weights will catch up asymptotically

I really don't think so. This almost never structurally happens.

I think it'll be more like Linux on the Desktop.

Or Ubuntu on the smartphone.

Or Firefox.

We'll have open weights, but 99% of everything will go through hyperscalers.

4mo agoHN ↗

I really don't think so. This almost never structurally happens. > I think it'll be more like Linux on the Desktop.

I think it will be Linux on the server, or the one that runs your watch, your phone, the radio or infotainment system in your car, maybe your thermostat, a bunch of medical devices and military devices, running in space shuttles and space stations and... You get the point. It's on everything.

4mo agoHN ↗

The smartphone is the most important piece of infrastructure in the modern world, yet we have basically two vendors.

Unless something dramatically changes, that's the world we're in for.

Chinese foundation model providers are releasing fewer weights as they "catch up", not more. There's little incentive for anyone to dump on the market if they can't collect the proceeds.

4mo agoHN ↗

The smartphone is the most important piece of infrastructure in the modern world, yet we have basically two vendors.

Unless something dramatically changes, that's the world we're in for.

If you limit LLM use to cellphones, maybe, but that seems awful silly right now. And why would you when there's so many B2B or B2C tools and products for it to go in. No reason to consider the market to be that constrained IMO.

4mo agoHN ↗

There's little incentive for anyone to dump on the market if they can't collect the proceeds.

Foreign state actors are not lacking incentives when the entire US economy is propped up by overvalued and overhyped AI. Like dumping a model that runs at Opus 4.6 brains at a fraction of the price on non-nvidia hardware.

4mo agoHN ↗

Mathematical ideas are very difficult to protect. But models can also be improved with brute force improvements in size. Imagine Mythos is a 32 trillion parameter model for example. That could be very difficult to replicate even though everybody knows exactly how it works.

4mo agoHN ↗

Superior architectures will leak pretty quickly via engineers. Withholding your best models doesn't work unless you have no competition.

4mo agoHN ↗

Withholding your best models doesn't work unless you have no competition.

It could also work if you DO have competition but your compute capacity is overbooked anyway, so releasing the better model doesn't actually make you that much more money (except for raising prices for the same amount of compute, which would give limited gains).

This is pretty much the situation Anthropic is in today.

4mo agoHN ↗

That just means that Anthropic is fucked unless they get more capacity.

4mo agoHN ↗

Superior architectures will leak pretty quickly via engineers.

I agree with the outcome of your premise (i.e., openness), but for different reasons:

First, isn't it the case that these bleeding edge 'newfangled' LLMs are basically variations on the same core ideas from "Attention Is All You Need" from 2017? [1]. Different scale, but still the same basic architecture. Even the "MoE" innovation keeps the Transformer attention stack while replacing or augmenting the dense feed-forward/MLP part with routed expert blocks.

And, I would argue that Engineers aren't working on new architectures. That would be Researchers, working on

  State-space models/Mamba (CMU/Princeton ecosystem), 
  Diffusion Language Models (Inception Labs), 
  Long-convolution architectures/Hyena (Stanford etc.), 
  RWKV/Recurrent LLMs (open-source community), 
  Memory-augmented architectures (Google Research/DeepMind?), 
  World models/spatial intelligence (LeCun/Fei-Fei Li/DeepMind), 
  Symbolic/neurosymbolic alternatives, 
  Thousand brains (Numenta).

That research is still open, so the outcome that you propose (openness) is likely to come to pass. Researchers/Scientists gotta publish, otherwise it's not science (to quote LeCun [2])

[1] https://arxiv.org/abs/1706.03762

[2] https://x.com/ylecun/status/1795589846771147018

4mo agoHN ↗

A bag of weights is not a product, but a component of a product. The models won't be that important at the end of the day. They are already good enough, or almost so.

The agentic tooling and harnesses will be what's important... and nobody has a moat with those, either. At least not until the easy money runs out and the patent suits start flying back and forth.

4mo agoHN ↗

I can't tell if this was written by AI or by an author that has absorbed all of its worse tendencies, but past the bullet list at the front this was terrible writing. It's like the author was trying to meet a page requirement.

4mo agoHN ↗

Def AI. Interesting perspective nevertheless

4mo agoHN ↗

There's a fourth option: the frontier labs are used by businesses who require (or think they require) the very best models and want to outsource model compliance such as HIPAA to someone else, and then individuals/smaller companies use the open source models.

4mo agoHN ↗

One anecdote: I've been using Kimi K2.6 exclusively for code work for about 1 week now (since it was released) with opencode. I don't really miss claude/codex. Kimi is much faster and cheaper and good enough for all my uses cases writing elixir and ruby code.

And it'll only get better/cheaper.

4mo agoHN ↗

Grab the weights. Seize the means of production. Tech workers of the world unite.

4mo agoHN ↗

American AI was financed on a particular bet. The bet was that frontier models would be the next great monopoly business

The collision between those two facts — that American capital paid for a moat, and that the technology no longer provides one — is the most important force in the AI industry today.

The open-weight ecosystem did not arrive in stages. It arrived in a wave. In late 2024, a Chinese lab named DeepSeek released a model

Looking at the assertions above, anyone passingly familiar with AI over the past few years will tell you that open weights and open research were the norm until OpenAI GPT-3 came along, and even then they were forced to release GPT-OSS by the market. So what technology moat? There has never been one in AI. Training 100B+ or trillion+ parameter models in expensive runs was potentially a moat, until the chinese startups showed in short order that it could be done for $6 million a run. Even the CUDA monopoly seems to be ending.

Also, no evidence referenced to back up any of the assertions. How do they know that the bet was that the frontier models would be the next great monopoly business? Especially when there were many from the outset: GPT, Anthropic, Llama, Deepmind, etc. etc.

I'd argue that the wholesale replacement of labor was and is the driver behind the capex, not monopoly dreams.

The starting premises appear to be, well, faulty. Whither the rest of the article?

4mo agoHN ↗

The moat is absolutely about integration, not the underlying tech/model.

Azure Copilot can charge whatever it wants because you can't use anything else.