Hacker News

Top stories

Live mirror
30 storiesupdated just nowView source snapshot
  1. Can gzip be a language model?(nathan.rs)
    48comments
  2. I said no and Apple said yes(dbushell.com)
    26comments
  3. Study: Young users (9 to 18Y) ditch Google for AI, with unknown consequences(norwegianscitechnews.com)
    6comments
  4. MiMo v2.6(xiaomi.com)
    399comments
  5. Spymarks, Not Watermarks(brand.io)
    109comments
  6. Transformers Explained Visually(poloclub.github.io)
    64comments
  7. Attention is all you have(alicegg.tech)
    240comments
  8. What Sun got wrong(dtrace.org)
    341comments
  9. MiMo-v2.6-Pro: Intelligence, Performance and Price Analysis(artificialanalysis.ai)
    16comments
  10. I don't want to read what you didn't write(colinbreck.com)
    269comments
  11. AI coding has made CI a bottleneck, so we reworked ours to keep up(linear.app)
    266comments
  12. Looking forward to Git 2.56 – and 3.0(lwn.net)
    54comments
  13. NASA’s Mars Sample Return mission is dead(science.org)
    323comments
  14. Divide by depth for instant 3D(gabrieloc.com)
    27comments
  15. Engineering Memory: On learning to memorize first 100 digits of pi (2024)(gregorygundersen.com)
    8comments
  16. World Wide Words(worldwidewords.org)
    1comments
  17. Claude Status – Elevated errors for multiple models(claude.com)
    80comments
  18. Used ThinkPad Buyer's Guide (2019)(bobble.tech)
    3comments
  19. A font that reads what you wrote(rohanadwankar.github.io)
    2comments
  20. PDF Forgeries Are Surprisingly Rare (2022)(gwern.net)
    29comments
  21. The Advisory Group on Mathematics and Artificial Intelligence(terrytao.wordpress.com)
    66comments
  22. What It's Like to Work in One of America's Data Centers(wsj.com)
    2comments
  23. Socrates vs. the Written Word (2011)(wondermark.com)
    20comments
  24. How do traffic signals work? (2019)(practical.engineering)
    63comments
  25. Python Workers are now generally available(cloudflare.com)
    38comments
  26. HERMES radio enables voice and data communication over vast distances(ieee.org)
    59comments
  27. Apple Copland D11E4 Booting in the Browser(pagetable.com)
    40comments
  28. Frontier AI on Your Own Hardware(timdettmers.com)
    78comments
  29. Grok 4.7(x.ai)
    482comments
  30. Turn off and restrict access to Apple Intelligence features on Mac(support.apple.com)
    198comments

MiMo-v2.6-Pro: Intelligence, Performance and Price Analysis

61 pointsby 5h agoartificialanalysis.ai
15 comments
3h agoHN ↗

"When evaluating the Intelligence Index, it generated 140M tokens, which is somewhat verbose in comparison to the median of 140M."

3h agoHN ↗

Nowadays these error can be a good thing :)

Human error means this wasn't just stopped together by some bot.

3h agoHN ↗

My bet is that it's a bot error, but of a rule based one.

38m agoHN ↗

It feels suspicious that MiMo-V2.6 Pro gets 46 in de index while DeepSeek-V4.1 gets 39.

Why?

27m agoHN ↗

The main AA benchmark keeps changing, and had to be radically changed when Astra came out and showed zero improvement over GPT 5.6 Sol in their benchmark. Opus 5 is still 1 point ahead of Fable 5.0 on the index, if you manually add Fable 5.0 back into the list, so it hasn't actually been "corrected". It's only Fable 5.1 that is shown as ahead of Opus 5.

The AA benchmark is a weighted average of other benchmarks and some internal ones. I think the difficult part is finding benchmarks that reflect your own use of the models.

2h agoHN ↗

It is an impressive model. Agreed on most that is written on this page, with the exception of it being fast. I ran it on my own LLM benchmark suite[1] and it is faster than DeepSeek but still much slower than leading models. But it's pricing is where it really shines.

KillSwitch-Bench 1.0

  Claude Opus 5           66.9
  GPT-6 Astra             57.9
  Claude Fable 5.1        46.7
  MiMo-V2.6-Pro           38.8
  Muse Spark 1.3          36.5

1 - https://bench.killswitch-lang.org/

1h agoHN ↗

OpenAI usage limits have been severely cut, and intelligence appears to be markedly declining, so I'm going to start trying these Chinese models seriously now. I don't mind if it takes longer. I just need the intelligence to predictably work the same way from day to day.

1h agoHN ↗

Same- I pay $200/mo for Codex but whereas I used to get a week's work out of a weekly limit, now I get roughly 1~2 days.

I've stopped using Astra entirely and remain on Sol orchestrating Luna Xhigh, but it's still not nearly a week's usage for a week's allotment.

And even then, whenever a new model is about to come out, it feels like the model I'm using is being dumbed down substantially.

I have no evidence for this and can have no evidence for this, but I can vote with my wallet regardless.

29m agoHN ↗

I strongly agree. Check out the Codex subreddit. Many empirical examples of Astra silently downgrading the models. One found Astra was silently using Luna Max (but still billing for Astra).

Even when I try to stick with Sol X/High, my limits are at best half of what they were before Astra launched, and the intelligence has declined markedly.

I cancelled my $100 plan. This is absolutely absurd and frankly unusable now.

56m agoHN ↗

Per Xiaomi, MiMo v2.6 training run cost $3.47m. A far cry from the estimated costs ($100m+) for the Big 5 (MSL, xAI, GDM, OAI, Ant). I wouldn't be surprised if salaries and R&D costs have similar drastic disparities.

For a model that matches Muse Spark 1.3 in benchmarks, MiMo v2.6 Pro is incredibly cheap, given its cache rates will remain $0.0036 per million.

27m agoHN ↗

I sorta got the impression that the $3.47 million only covered post-training , given that few of the graphs start at zero. Is a barely-trained model going to score 48 on DeepSWE v1.1 ?

https://mimo.xiaomi.com/rl/