- 10comments
- 378comments
- 91comments
- 224comments
- 59comments
- 327comments
- 8comments
- 221comments
- 227comments
- 312comments
- 46comments
- 4comments
- 24comments
- —discuss
- 60comments
- 17comments
- 28comments
- 58comments
- 36comments
- 75comments
- 57comments
- 477comments
- 193comments
- 38comments
- 22comments
- 76comments
- 8comments
- 36comments
- 23comments
- 300comments
"When evaluating the Intelligence Index, it generated 140M tokens, which is somewhat verbose in comparison to the median of 140M."
Nowadays these error can be a good thing :)
Human error means this wasn't just stopped together by some bot.
My bet is that it's a bot error, but of a rule based one.
Yes, it seems they have a template that they fill with numbers. Similar issues spotted on Grok's performance page: https://news.ycombinator.com/item?id=49789558
It feels suspicious that MiMo-V2.6 Pro gets 46 in de index while DeepSeek-V4.1 (https://artificialanalysis.ai/models/deepseek-v4-1-flash) gets 39. According to the appendix at the bottom of https://mimo.xiaomi.com/mimo-v2-6 the deepseek model sometimes surpasses mimo and it's not so far behind in capabilities. A week ago opus 5 appeared 1 points ahead of fable 5 despite fable being a much smarter model
It is an impressive model. Agreed on most that is written on this page, with the exception of it being fast. I ran it on my own LLM benchmark suite[1] and it is faster than DeepSeek but still much slower than leading models. But it's pricing is where it really shines.
KillSwitch-Bench 1.0
1 - https://bench.killswitch-lang.org/
Why sol is not in the comparison?