- 29comments
- 395comments
- 102comments
- 62comments
- 232comments
- 330comments
- 10comments
- 246comments
- 252comments
- 7comments
- 49comments
- 1comments
- 320comments
- 25comments
- 28comments
- 63comments
- 1comments
- 76comments
- 18comments
- 62comments
- 38comments
- 1comments
- 58comments
- 77comments
- 22comments
- 39comments
- 481comments
- 195comments
- 37comments
- 300comments
"When evaluating the Intelligence Index, it generated 140M tokens, which is somewhat verbose in comparison to the median of 140M."
Nowadays these error can be a good thing :)
Human error means this wasn't just stopped together by some bot.
My bet is that it's a bot error, but of a rule based one.
Yes, it seems they have a template that they fill with numbers. Similar issues spotted on Grok's performance page: https://news.ycombinator.com/item?id=49789558
It feels suspicious that MiMo-V2.6 Pro gets 46 in de index while DeepSeek-V4.1 (https://artificialanalysis.ai/models/deepseek-v4-1-flash) gets 39. According to the appendix at the bottom of https://mimo.xiaomi.com/mimo-v2-6 the deepseek model sometimes surpasses mimo and it's not so far behind in capabilities. A week ago opus 5 appeared 1 points ahead of fable 5 despite fable being a much smarter model (this has been corrected already)
It is an impressive model. Agreed on most that is written on this page, with the exception of it being fast. I ran it on my own LLM benchmark suite[1] and it is faster than DeepSeek but still much slower than leading models. But it's pricing is where it really shines.
KillSwitch-Bench 1.0
1 - https://bench.killswitch-lang.org/
Why sol is not in the comparison?
OpenAI usage limits have been severely cut, and intelligence appears to be markedly declining, so I'm going to start trying these Chinese models seriously now. I don't mind if it takes longer. I just need the intelligence to predictably work the same way from day to day.
Same- I pay $200/mo for Codex but whereas I used to get a week's work out of a weekly limit, now I get roughly 1~2 days.
I've stopped using Astra entirely and remain on Sol orchestrating Luna Xhigh, but it's still not nearly a week's usage for a week's allotment.
And even then, whenever a new model is about to come out, it feels like the model I'm using is being dumbed down substantially.
I have no evidence for this and can have no evidence for this, but I can vote with my wallet regardless.
Per Xiaomi, MiMo v2.6 training run cost $3,470,000. A far cry from the estimated costs for the Big 5 (MSL, xAI, GDM, OAI, Ant). I wouldn't be surprised if salaries and R&D costs have similar drastic disparities.
For a model that matches Muse Spark 1.3 in benchmarks, MiMo v2.6 Pro is incredibly cheap, given its cache rates will remain $0.0036 per million.