Hacker News

New stories

Live mirror
30 storiesupdated just nowView source snapshot
  1. Tin: full-text search for Postgres(planetscale.com ↗)
    discuss
  2. Part-human part-mouse brain developed in science breakthrough(bbc.com ↗)
    discuss
  3. Show HN: FastRecall, ultra-cheap memory across AI models(fastrecall.ai ↗)
    discuss
  4. Michael Burry slams OpenAI, Anthropic for 'self-serving' calls to slow AI(nypost.com ↗)
    discuss
  5. Spanish data watchdog publicises first AI agent-linked data breach report(reuters.com ↗)
    discuss
  6. Xcode-Project-Format(github.com/apple ↗)
    discuss
  7. Flowchart: How pixels become an Apple Reference Image(claude.ai ↗)
    discuss
  8. Supply Chain Compromise of Korean-Language Windows 11 Installation Media(logpresso.com ↗)
    discuss
  9. Pangram – AI detector for text and images(pangram.com ↗)
    discuss
  10. Migrating the GitHub Copilot Runtime to Rust, Using Copilot(github.blog ↗)
    1comments
  11. Game UI Database(gameuidatabase.com ↗)
    discuss
  12. Found a B2B billing stack for my agency that doesn't feel like 2010(cordhq.app ↗)
    discuss
  13. Pro UI: native grade components for pro software(pro-ui.dev ↗)
    discuss
  14. Page Shield ML caught 4 storefront malware campaigns scanners missed(cloudflare.com ↗)
    discuss
  15. Why the Postpandemic Tech Bust Sent Billionaires to Trump(wired.com ↗)
    discuss
  16. Hacker puts 'full redundancy' code-hosting firm out of business (2014)(pcworld.com ↗)
    discuss
  17. OpenAI Discloses Six New Incidents of ‘Concerning’ A.I. Behavior(nytimes.com ↗)
    7comments
  18. Math challenges a 2k-year-old story about Parthenon's optical illusions(phys.org ↗)
    discuss
  19. Could Your Chatbot Be Stealing Ideas from You?(nytimes.com ↗)
    1comments
  20. Weeping whales: Stillborn humpback whale grieving documented(phys.org ↗)
    discuss
  21. I gave my agents a heartbeat(mlsystemsri.com ↗)
    discuss
  22. Will AI Replace Your Doctor? Dr. Zeke Emanuel and AMA President Dr. John Whyte(youtube.com ↗)
    1comments
  23. A Large Database Does Not Mean Large Shared_buffers(keithf4.com ↗)
    discuss
  24. Size-Specialized Memory Allocation(go.dev ↗)
    discuss
  25. What's Scarier Than Agents Taking over Internet? CEO Cartel Trying Take over AI(fractalsofchange.substack.com ↗)
    1comments
  26. AI Cheating Is on the Rise(vals.ai ↗)
    discuss
  27. Unplug America(engageq.notion.site ↗)
    discuss
  28. Suzanne Ciani's Buchla Cookbook(echo.orpheusinstituut.be ↗)
    discuss
  29. OpenAI discloses six new AI safety incidents(axios.com ↗)
    1comments
  30. The Return of Sail Power: Cargo Ships Are Turning Back to the Wind(gcaptain.com ↗)
    discuss

Xiaomi Mimo 2.6 live post-training dashboard

249 pointsby 5h agomimo.xiaomi.com
61 comments
4h agoHN ↗

Hah, it would be great to see more labs pick this up.

4h agoHN ↗

This is pretty neat. What would be a good reason for the other Model providers to not do this?

4h agoHN ↗

Speculating here, but I assume researchers can make a reasonable estimate of the size of closed models based on factors like training time, training speed, and the number of tokens processed.

Also, Anthropic and OpenAI probably want to keep each other on their toes so they don’t end up on the wrong side of another Opus 4.6 / GPT-5.3-Codex situation, where one lab releases a model only for the other to drop a better one hours later.

4h agoHN ↗

I think first of all it’s not an obvious idea, also the marketing surplus for other providers is not as big for openai/anthropic as for xiaomi and last but not least I’m pretty sure you can withdraw methodology from here.

I’m saying who has a million dollars for me, so I can make my own model?

4h agoHN ↗

I didn't know 2 thirds of the training data would be source code.

4h agoHN ↗

that is the the "data used to improve the model" when signing up for the subscription plans

4h agoHN ↗

this is the rl run, not the pretraining run

4h agoHN ↗

even in pre-training, usually 30%-50% is code these days.

4h agoHN ↗

This is crazy, but sadly anthropic/openai will never do this, what has happened to this world, where chinese companies are more open than US or even EU companies

4h agoHN ↗

Neoliberalism, that famously open and transparent economic ideology

4h agoHN ↗

Why are they doing this? To try head off accusations about distillation?

4h agoHN ↗

Sometimes you're confident about what you're doing and show how you work to the world.

Keeping the garage door open, or at least making the door translucent. It's always cool.

4h agoHN ↗

That China's official policy is now to prefer open models and open model development may be a part of it.

4h agoHN ↗

With that policy in place, labs might be incentivized to be creative in their openness. This being fun/free PR

4h agoHN ↗

BRICS just had a New Delhi meeting where Xi pushed a 5-point plan on AI cooperation that centered on open source models

3h agoHN ↗

Bottom of the page says "Open is what we value."

4h agoHN ↗

When you run benchmarks while training, isn't that the definition of contamination? Asking because I am not sure if this is normal in big labs now.

4h agoHN ↗

They are using it to evaluate checkpoints during the training, they are probably not using the benchmarks for training the models. It's a common practice for big reinforcement learning runs.

4h agoHN ↗

Kinda yes. The benchmarks become part of the validation set, which means the models get slightly overfit to them if they are used as criteria for stopping the training. But a lot less compared to using them in the training data.

I'd guess everybody uses at least some benchmarks as stopping criteria, which is kinda sensible, but it also does induce some benchmaxxing, and explains partly why the newest models always tend to eke out in benchmarks.

https://en.wikipedia.org/wiki/Training,_validation,_and_test...

4h agoHN ↗

Correct. If just stopping criteria, that is less contaminated. The question gets muddier once you also use it to determine hyperparameters during small-scale runs.

4h agoHN ↗

You gotta have something to aim at. And, presumably, the benchmark is not part of the training data, it is the test against which the model is tested at each stage; is behavior moving in the right direction?

3h agoHN ↗

It's implicitly trained against. There is like information leakage with researchers messing with the training parameters and checkpoints used.

It's not the direct feedback loop of RL but its not far.

3h agoHN ↗

They exist to detect degradation. Datasets are not perfect and if a batch contains too much bad data it can ruin a run, also an opportunity to find bad data and improve the dataset filtering.

4h agoHN ↗

2.6 Pro: >started 2026-09-15 10:32 UTC

For some reason I thought training took much, much longer than what the progress bar suggests.

This is really neat, I'm currently using mimo 2.5 pro, and it's decent (or great given the price). Hopefully their next one is multimodal.

4h agoHN ↗

These are post-training reinforcement learning steps.

4h agoHN ↗

Yes, updated the submission title to say "post-training" to hopefully prevent further confusion

4h agoHN ↗

I been using MiMo-V2.5 to do most of my work as software engineer, on a variety of projects I'm working on, and I been VERY happy with ROI. The model is very powerful! Not perfect – I've run in hallucination loops once or twice, but nothing a stop-then-continue wouldn't solve.

The cost is unbelievably low, and the quality of intelligence I get is equivalent to when I was working mostly with Anthropic models (late last year/early this year). I'm fully invested in MiMo and I'm very happy with it.

-- PS: I also check almost daily to see if other models are capable of doing such great work. And they do – DS4F is powerful and DS41 is impressive, GLM 5.3 Flash gets a job done well, etc. – but when I add cost of M-token in the ROI math, Jeez! MiMo is an order of magnitude better.

4h agoHN ↗

2.5 Pro or the regular 2.5?

I always found that those Mimo models to be really good at tool calling and following instructions

4h agoHN ↗

I've found that mimo v2.5 works for very basic things like a python script to do one thing, but it also is very 'dumb' compared to qwen 3.8-flash-next (I think the benchmark scores for terminal and coding specific benches back this up). And definitely not in the same class as like a GLM5.2 or 5.3. It's fast but makes basic mistakes that only get caught later.

4h agoHN ↗

May I ask why you ended up there instead of just using the heavy subsidized subscription. I’m actually curious.

3h agoHN ↗

How fast is it compared with the other Chinese models?

2h agoHN ↗

They both are in the 50-100 tok/s range. The Mimo v2.5 Pro Ultraspeed beta could reach 1000 tok/s, hoping they can do something similar for the new model, it was amazing.

2h agoHN ↗

I am also using 2.5 and it is giving me solid results. Its available free on Openrouter

1h agoHN ↗

How does it compare with DS 4.1 Flash in your experience, if you ignore the cost?

14m agoHN ↗

I’ve been very pleased with DS 4.1 flash. Not so much the 4.0 models, but for coding (Rust) it’s been great so far (3 solid days of work).

I’ll give Mimo a try.

4h agoHN ↗

You'd think they would make it less obvious that they are running their whole operation with Claude

4h agoHN ↗

It's not obvious to me. What's the tell?

2h agoHN ↗

If you're thinking of the UI style, definitely not Claude. It is incapable of writing a clear sentence like "what each step's samples are made of", would have used all-caps for everything, more padding and gradients.

1h agoHN ↗

They'd be running in the red then cause they charge way less than Claude. Sorry but it just doesn't make logical sense. They have open source, papers, and self hosting too

4h agoHN ↗

The Chinese labs are just making fun of the US labs at this point.

Where is the cool shit from the US labs?

3h agoHN ↗

With other software, devs convince their managers of the importance of using open source stuff in their stack. With AI, it's usually managers choosing what models to use for the devs. The US labs don't need to give a damn how much devs like open source

3h agoHN ↗

The US labs don't need to give a damn how much devs like open source

In the short term, true.

In the long term, unknown but typically when you hold progress that way while other countries don't you at best end up becoming siloed while the rest of the world continues on without you.

2h agoHN ↗

This isn't about liking open source. This is about the labs just being cool and doing cool shit instead of the opposite which is Anthropic where all they talking about is killing everyone and taking everyone's job.

14m agoHN ↗

These labs are still (for the time being) made of people, who reflect their lives onto the work.

The US population is much more pessimistic and doomsday driven these days, whereas the Chinese are more optimistic and future driven.

3h agoHN ↗

Very cool to see the openness here, and likely more like this will come from smaller startups where they win users on transparency.

3h agoHN ↗

Well, if open source AI is dangerous (for OpenAI/Anthropic IPOs?), this is like watching a time bomb.

1h agoHN ↗

the open burial started when zAI served their latest model on all Chinese chips.

now we r just noticing the grave getting dug deeper.

39m agoHN ↗

For my own usage, Luna is cheap enough that I don't care if other models are cheaper. I'm interested if another model is in some way better and not too expensive.

11m agoHN ↗

Luna is great but makes a lot of mistakes at high and lower in my experience (large rust codebase). I use Luna Max for asynchronous subagent reviews and am very happy with its work, but it’s slow af.

2h agoHN ↗

Neat! I've been trying out their next model for the last week, which I assume is a version of this, and it's been a good experience so far.

I had used 2.5-pro for a hefty chunk of development, and found it to work like a somewhat forgetful senior engineer who was new to my project. Very capable, would almost always choose a reasonable option, if not always the best one for the project, and not great at multi-tasking. Generally, made me comfortable not scrutinizing the code line-by-line, but still needed a bit of steering once projects got to a reasonable size.

The next model is a clear step up in the multi-tasking capability at least, with me very rarely having to steer the implementation of a well-defined issue. In terms of code, I found MiMo-V.2.5-pro to be extremely conservative, implementing minimal solutions. The next model seems a little bit more ambitious, in positive ways, making good guesses about gaps/next steps. It also seems to be a fair bit better at design, at least for the little bit I've done, it was good at translating my concepts to practical elements on screen, and cleaned things up nicely as I made suggestions.

1h agoHN ↗

2.6-pro just reached 63.7% by step 10, it's on step 11 right now.

Even flash reached 60.7% by step 12, and it's on step 16 now.

This is so exciting lmao.

2h agoHN ↗

Mino 2.5 has been my workhorse for coder and tester agents (the ones planner agents delegate tasks to)

2h agoHN ↗

Distillation in real-time? Very interesting!

1h agoHN ↗

That "training cost" is just live revenue count for Anthropic/OpenAI API calls!

/s

22m agoHN ↗

Very curious that everyone here (so far) seems to assume this dashboard presents real data.

5m agoHN ↗

$5 per second if my eyes don’t fool me. That’s ~$432K per day. Enough to rent 3,000 B300 nodes on Modal.