687 points by volf_ about 9 hours ago | 325 comments | View on ycombinator
rao-v about 8 hours ago |
stymaar about 8 hours ago |
Pro [2]:, 1.02T total / 42B activated parameters
margorczynski about 6 hours ago |
No matter how much cash you throw you can't just materialize a 100 nuclear reactors to power the data centers.
user43928 about 8 hours ago |
Maybe Terminal Bench 4.0 and ExploitGym are reasonable.
Terminal Bench 4.0
GPT 6 Astra 59.6
Claude Fable 5.1 55.1
Claude Opus 5 49.0
MiMo-V2.6-Pro 34.9
MiMo-V2.6-Flash 28.8
DeepSeek V4.1 Flash 26.8
MiMo-V2.5-Pro 1.5
ExploitGym GPT 6 Astra 42.4
Claude Fable 5.1 30.4
Claude Opus 5 22.1
MiMo-V2.6-Pro 17.8
MiMo-V2.6-Flash 6.0
MiMo-V2.5-Pro 0.1
DeepSWE v1.1 DeepSeek V4.1 Flash 74.2
Claude Opus 5 74.0
GPT 6 Astra 74.0
MiMo-V2.6-Pro 71.9
Claude Fable 5 70.0
MiMo-V2.6-Flash 67.9
MiMo-V2.5-Pro 19.0simonw about 8 hours ago |
Pelicans for Pro: https://tools.simonwillison.net/markdown-svg-renderer?url=ht...
nemothekid about 8 hours ago |
vatsachak about 9 hours ago |
Some features of the release I like:
- Demonstration of diverse tasks, such as using a DAW
- Graphs from various benchmarks and price ranges
- Real world use of the model in scientific environments
toephu2 about 7 hours ago |
No moat and competition is good for consumers though.
volf_ about 8 hours ago |
Averages ~25-35tok/s which isn't bad for a first attempt.
GodelNumbering about 6 hours ago |
Interesting but not surprising trend across the board seems to be, the flash models seems to have caught up with the pro-sized models of H1'26. No surprise all labs are rushing to bigger models.
EDIT: Wow, took a detailed look at the benchmarks. Mimo 2.6 pro, the 1T model leads Kimi K3, a 2.8T param model in 14 out of 15 benchmarks (and the last one is near tie)!! Good to see they also kept the price the same, and landed in the greenest quardrant of the intelligence vs speed of AA.
paradox460 about 4 hours ago |
Also when I was using it, I managed to do quite a bit on a few bucks worth of OpenRouter credits. Not sure how well it keeps up in the modern world against things like Luna, but I hope it remains competitive
XCSme about 5 hours ago |
thrownawaysz about 8 hours ago |
It's because offpeak electricity is cheaper?
Funnily it's perfect if you are in the Pacific Time Zone because you can use it daytime 9am to 5pm
wren6991 about 1 hour ago |
> we also removed the cyber dataset from the upcoming pro run, since we observed some bad patterns in the rollout logs.
And this in the model card (emphasis mine):
> Aligned RL: Cold start from self-correction — the model reflects on and rewrites its own misaligned turns into grounded next steps. Throughout RL, environment hardening, adversarial screening, and verifier cross-checks keep the loop honest against reward hacking.
Environment hardening during the RL runs. Uh oh, did someone start making a few too many paperclips?
syntaxing about 8 hours ago |
jjcm about 6 hours ago |
Design: https://image.non.io/78795662-8bfc-4e14-8d72-3738392aa6b3.we...
MiMo 2.6 Pro Ultraspeed (36min): https://html.non.io/annui-mimo/
Grok 4.7 (25min): https://html.non.io/Annui-grok/
Astra (19min): https://html.non.io/annui/
Overall this felt like the weakest of the three. Ultraspeed was fast as far as tokens per second goes, but it overthought quite a lot of things resulting it in having one of the longest build times. That overthinking didn't lead to better results either - note the statue with the cropped off arm. It's also the worst implementation of the dynamic lighting effect / displacement effect - the background especially has some significant distortion. Astra was the only one that seemed to understand that displacement should happen less the further something is in the distance.
Here's a vid of all 3 side by side with the source design: https://non.io/video/annui-comparison.mp4
Havoc about 5 hours ago |
wkcheng about 6 hours ago |
With some models you can find hosting companies based in the EU or US, but then you don't know how they're quantizing the models, so you're not sure about the actual output quality.
How are people actually using this? Or are people just experimenting with side projects?
lwansbrough about 8 hours ago |
dom96 about 6 hours ago |
KillSwitch-Bench 1.0
Claude Opus 5 66.9
GPT-6 Astra 57.9
Claude Fable 5.1 46.7
MiMo-V2.6-Pro 38.8
Muse Spark 1.3 36.5
1 - https://bench.killswitch-lang.org/undefined about 9 hours ago |
informal007 about 5 hours ago |
Zaraif13 about 3 hours ago |
Gotta love a capable open model. BUT, how can they just casually throw in that they're actively exploring RSI as if it's just another technique? Is this not alarming at all?
bonsai_spool about 6 hours ago |
I’ll be trying these models out and may end up switching my subscriptions if this craziness continues
wmedrano about 4 hours ago |
est about 3 hours ago |
xlayn about 4 hours ago |
Twice, trice or quadrix(tm) are just approximations, we have already paid with:
- more expensive electronics
- less work
- all the retirement money put into gpus
- all the "fair use" of all the books, all the images,
- and then taxes to bail them?heyjstn about 2 hours ago |
ddxv about 8 hours ago |
pulkitsh1234 about 6 hours ago |
nivance about 3 hours ago |
stemlord about 6 hours ago |
eriquesito about 7 hours ago |
DanMcInerney about 8 hours ago |
drob518 about 7 hours ago |
MisterMunchkin about 8 hours ago |
Just tried 2.6 flash on a really niche topic I specialise in and it has done a really good job. They’ve definitely polluted their training data with claudeslop, but looking past the slop there is a decent model.
algoth1 about 9 hours ago |
coss about 5 hours ago |
esafak about 7 hours ago |
https://artificialanalysis.ai/models/mimo-v2-6-pro#intellige...
That's pretty fast; I think I'll try it: https://openrouter.ai/xiaomi/mimo-v2.6-flash
One concern I have is that they allegedly do not discount cached tokens: https://www.reddit.com/r/opencodeCLI/comments/1t37dz3/xiaomi...
Can anyone comment?
system2 about 4 hours ago |
bertili about 8 hours ago |
alfalfasprout about 8 hours ago |
And as these models get better the pace of training is quickly speeding up too.
This doesn't bode particularly well for anthropic/OAI after they go public.
varispeed about 8 hours ago |
gigatexal about 8 hours ago |
NooneAtAll3 about 8 hours ago |
so weird to acknowledge someone being on the front edge, but not name it
spwa4 about 8 hours ago |
MiMo-V2.6-Flash-310B-A15B roughly GPT-5.6 Luna / Claude 4.9 according to benchmarks MiMo-V2.6-Pro-1.02T-A42B roughly GPT-5.6 Sol / Opus 5 according to benchmarks
Perhaps with IQ2 flash will run on 128G M5?
cmrdporcupine about 5 hours ago |
I deliberately left things wide open and ambiguous to see what it would do. I think with better upfront planning this would be excellent value.
legions-love about 1 hour ago |
paidx about 4 hours ago |
16t96 about 8 hours ago |
nlcs about 8 hours ago |
nofpu about 8 hours ago |
EtienneDeLyon about 3 hours ago |
unpopularopp about 8 hours ago |
omani about 8 hours ago |
but now I got my "proof".
jwpapi about 8 hours ago |
It looks like the "Frontier Line" to me, which is also often misinterpreted. frontier does not mean the best models. It means all models that are not strictly dominated, meaning in most cases: Not same price or cheaper and more intelligent.
I personally would like the word frontier to be used with more criterias: Open Weights, per use-case, etc etc. This would make model selection easier, but I understand it’s not an easy thing to do.
paperboy10000 about 3 hours ago |
They are excellent in marketing, I guess that is something.
The realtime dashboard they shared during training (https://mimo.xiaomi.com/rl/) was an incredible learning and teaching tool for me, and they’ve been unusually comprehensive in sharing details about their methodology (check out that tech report - it's got lots of clever behind the scene tricks like Google or Deepseek writeups) and benchmark scores (even the stuff they didn’t do well on).
If you’re releasing an open model going forward, please consider offering the community more of this transparency!