Hacker News

Top stories

Live mirror
30 storiesupdated just nowView source snapshot
  1. AMD's random number generator can't generate a 0?(flatassembler.net)
    54comments
  2. Can gzip be a language model?(nathan.rs)
    77comments
  3. 9 Ads per Minute: FIFA Cup 26 – "the price of the beautiful game"(bristol.ac.uk)
    90comments
  4. MiMo v2.6(xiaomi.com)
    419comments
  5. Type Punning in C and C++(pwkf.org)
    2comments
  6. Spymarks, Not Watermarks(brand.io)
    120comments
  7. Verda (Finland) raises $189M in Series B(verda.com)
    10comments
  8. JetBrains Air: A System of Products for Agentic Software Development(jetbrains.com)
    23comments
  9. Transformers Explained Visually(poloclub.github.io)
    71comments
  10. Attention is all you have(alicegg.tech)
    255comments
  11. MiMo-v2.6-Pro: Intelligence, Performance and Price Analysis(artificialanalysis.ai)
    30comments
  12. I said no and Apple said yes(dbushell.com)
    219comments
  13. What Sun got wrong(dtrace.org)
    351comments
  14. A font that reads what you wrote(rohanadwankar.github.io)
    16comments
  15. JavaFX 27 Native Image on a Raspberry Pi 5(ennerf.github.io)
    4comments
  16. I don't want to read what you didn't write(colinbreck.com)
    306comments
  17. AI coding has made CI a bottleneck, so we reworked ours to keep up(linear.app)
    298comments
  18. A build graph that rolls dice(fzakaria.com)
    1comments
  19. Engineering Memory: On learning to memorize first 100 digits of pi (2024)(gregorygundersen.com)
    12comments
  20. Divide by depth for instant 3D(gabrieloc.com)
    30comments
  21. World Wide Words(worldwidewords.org)
    2comments
  22. Looking forward to Git 2.56 – and 3.0(lwn.net)
    74comments
  23. NASA’s Mars Sample Return mission is dead(science.org)
    338comments
  24. What It's Like to Work in One of America's Data Centers(wsj.com)
    13comments
  25. Claude Status – Elevated errors for multiple models(claude.com)
    92comments
  26. The Advisory Group on Mathematics and Artificial Intelligence(terrytao.wordpress.com)
    71comments
  27. Python Workers are now generally available(cloudflare.com)
    38comments
  28. Socrates vs. the Written Word (2011)(wondermark.com)
    29comments
  29. HERMES radio enables voice and data communication over vast distances(ieee.org)
    61comments
  30. How do traffic signals work? (2019)(practical.engineering)
    64comments

MiMo-v2.6-Pro: Intelligence, Performance and Price Analysis

96 pointsby 8h agoartificialanalysis.ai
30 comments
5h agoHN ↗

"When evaluating the Intelligence Index, it generated 140M tokens, which is somewhat verbose in comparison to the median of 140M."

5h agoHN ↗

Nowadays these error can be a good thing :)

Human error means this wasn't just stopped together by some bot.

5h agoHN ↗

My bet is that it's a bot error, but of a rule based one.

3h agoHN ↗

It feels suspicious that MiMo-V2.6 Pro gets 46 in de index while DeepSeek-V4.1 gets 39.

Why?

2h agoHN ↗

The main AA benchmark keeps changing, and had to be radically changed when Astra came out and showed zero improvement over GPT 5.6 Sol in their benchmark. Opus 5 is still 1 point ahead of Fable 5.0 on the index, if you manually add Fable 5.0 back into the list, so it hasn't actually been "corrected". It's only Fable 5.1 that is shown as ahead of Opus 5.

The AA benchmark is a weighted average of other benchmarks and some internal ones. I think the difficult part is finding benchmarks that reflect your own use of the models.

1h agoHN ↗

The way Artificial Analysis keeps changing their weights feels kind of like deciding who the winner should be and making the weights reflect that. They’ve been changing their weights to add more weight to improved long-running agentic capabilities, but doing so means they’re reducing the relative importance of world knowledge and of writing ability.

I’ll grant that maybe world knowledge isn’t that important for these models. But writing ability is important for human understanding, and I think the weird turns of phrase and word choices reflect the labs’ underweighting of the importance of human understanding.

4h agoHN ↗

It is an impressive model. Agreed on most that is written on this page, with the exception of it being fast. I ran it on my own LLM benchmark suite[1] and it is faster than DeepSeek but still much slower than leading models. But it's pricing is where it really shines.

KillSwitch-Bench 1.0

  Claude Opus 5           66.9
  GPT-6 Astra             57.9
  Claude Fable 5.1        46.7
  MiMo-V2.6-Pro           38.8
  Muse Spark 1.3          36.5

1 - https://bench.killswitch-lang.org/

1h agoHN ↗

Speed seems to vary a lot with demand. Last night it was reaching 80+ tok/s

2h agoHN ↗

the graph has a dropdown for selecting models

1h agoHN ↗

Not the graph at the top, the one further down.

4h agoHN ↗

OpenAI usage limits have been severely cut, and intelligence appears to be markedly declining, so I'm going to start trying these Chinese models seriously now. I don't mind if it takes longer. I just need the intelligence to predictably work the same way from day to day.

3h agoHN ↗

Same- I pay $200/mo for Codex but whereas I used to get a week's work out of a weekly limit, now I get roughly 1~2 days.

I've stopped using Astra entirely and remain on Sol orchestrating Luna Xhigh, but it's still not nearly a week's usage for a week's allotment.

And even then, whenever a new model is about to come out, it feels like the model I'm using is being dumbed down substantially.

I have no evidence for this and can have no evidence for this, but I can vote with my wallet regardless.

2h agoHN ↗

I strongly agree. Check out the Codex subreddit. Many empirical examples of Astra silently downgrading the models. One found Astra was silently using Luna Max (but still billing for Astra).

Even when I try to stick with Sol X/High, my limits are at best half of what they were before Astra launched, and the intelligence has declined markedly.

I cancelled my $100 plan. This is absolutely absurd and frankly unusable now.

2h agoHN ↗

It feels bizarre reading about the amounts spent on it here and paying like 10 eurobucks a week for DS

2h agoHN ↗

These tools were pretty great if you could afford them, but now they are expensive and shit, and that combination doesn't work.

52m agoHN ↗

For me, DS Pro is still behind Sol. But I do think many people hitting the Codex limits in 1 or 2 days are doing something wrong.

13m agoHN ↗

My spending with Deepseek was more than a Codex subscription, though. How are you keeping it to $10/month? Super light usage?

1h agoHN ↗

and intelligence appears to be markedly declining

Serious question: does anyone have evidence of this?

It’s something that’s constantly asserted, and has been since 2023. Every time someone posts a site that tries to track this though, I look at it and it’s just a flat line.

1h agoHN ↗

Sol 5.6 xhigh had been a very reliable workhorse for coding for me via the 200 bucks sub.

But this week they seem to have tweaked the system to a point at which all models (Astra, Sol, Luna) hit rate limits all_the_time without me being anywhere close to the weekly limit.

Early results with MiMo 2.6pro are quite encouraging for anything that's non-UI work so likely switching spend for the time being

3h agoHN ↗

Per Xiaomi, MiMo v2.6 training run cost $3.47m. A far cry from the estimated costs ($100m+) for the Big 5 (MSL, xAI, GDM, OAI, Ant). I wouldn't be surprised if salaries and R&D costs have similar drastic disparities.

For a model that matches Muse Spark 1.3 in benchmarks, MiMo v2.6 Pro is incredibly cheap, given its cache rates will remain $0.0036 per million.

2h agoHN ↗

I sorta got the impression that the $3.47 million only covered post-training , given that few of the graphs start at zero. Is a barely-trained model going to score 48 on DeepSWE v1.1 ?

https://mimo.xiaomi.com/rl/

2h agoHN ↗

My understanding of tech salaries in China is that they are pretty decent, but not as high as in SF; closer to typical European salaries.

Mostly due to lower cost of living; Shenzhen is way cheaper than SV

50m agoHN ↗

I seriously doubt salaries are included. It must be just the electricity and GPU costs.

54m agoHN ↗

where's the flash model? it's out already isn't it