Home AI News

AI News Roundup for September 22: MiMo-V2.6-Pro, Xiaomi’s MIT-Licensed Model, Matches Grok 4.7 for 4 Percent of the Cost

Xiaomi's MIT-licensed MiMo-V2.6-Pro ties Grok 4.7 at 46 on the intelligence index for a 29th of the cost per task. Plus Muse's 0-day and BC v. OpenAI.

Two identical analog gauges with needles at the same position, a tall stack of coins beside one and a single coin beside the other

Two frontier-class models shipped within hours of each other yesterday and landed on the identical intelligence score, which every writeup I read covered as two separate launch stories. Put the two price sheets on one screen and the actual news shows up.

Xiaomi and xAI Hit the Same Number on the Same Day

Xiaomi released MiMo-V2.6-Pro on Monday under an MIT license. It is a 1.02-trillion-parameter mixture-of-experts model with roughly 42 billion parameters active per token, and Artificial Analysis ranks it first of 114 open-weights models with an Intelligence Index score of 46. Xiaomi lists it at $0.435 per million input tokens and $0.87 per million output, with a 1M-token context window and text, image, video and audio on the input side.

xAI shipped Grok 4.7 the same day. Artificial Analysis scores it 46 at extra-high reasoning effort. Identical number. xAI prices it at $2.00 input and $6.00 output.

Here is the figure nobody put in a headline. Artificial Analysis publishes a cost per Intelligence Index task alongside the score, and it reads $0.13 for MiMo-V2.6-Pro against $3.74 for Grok 4.7. Same measured intelligence, 29 times the bill.

I want to be careful about what that does and does not mean, because index parity is not task parity. A composite score averages reasoning, knowledge, math and coding, and your workload is not that average. xAI’s own benchmark table is also self-reported: Grok 4.7 claims 46.3 percent on CursorBench 4.0 and 71.0 percent on DeepSWE v1.1, and LLM Stats flags explicitly that none of it has been independently verified. Terminal-Bench 4.0 going from 20.3 to 38.0 in a single point release is the kind of jump I would want to reproduce before I believed it.

What I do believe is the cost floor moved, and the release notes are why. Xiaomi says the Pro run cost $2.62 million and finished 30 large reinforcement-learning steps across roughly 750,000 trajectories in under six days, and VentureBeat’s readout puts the cheaper V2.6-Flash at $850,000 for a 310-billion-parameter model at $0.14 and $0.28. Xiaomi also published more than 7,000 reinforcement-learning task environments, the training framework, and a 9-billion distill built on Qwen3.5. That is not a model drop. That is a recipe, and the weights are MIT, so you can self-host the thing.

Decrypt’s read on xAI is that it has trailed the frontier pack across releases and is selling “good enough, cheap.” I would revise that. As of Monday it is neither the best nor the cheap one, and Anthropic’s Claude Fable 5.1 still tops it on GDPval at 1735 to 1695. Verdict: matters. Go run your own eval this week. If MiMo holds on your actual traffic, the wrapper-margin problem I wrote about on September 17 stops being a rounding error and starts being most of your invoice.

Any App Already Running on Your Mac Can Hijack Meta’s Muse

Patrick Wardle of Objective-See disclosed a macOS zero-day in Meta’s Muse assistant, along with a working proof of concept he named not-a-mused. The mechanism is an undocumented configuration key, endo_voyager_dictation_endpoint, that an unprivileged local process can rewrite without asking anyone for permission. Redirect it and you intercept dictated audio and prompts before they reach Meta, inject your own instructions into the agent, and lift authentication material tied to the account.

Read the coverage carefully, because most of it says “hackers hijack” and that oversells it. Cybersecurity News notes the attacker needs local code execution first. This is not a drive-by. It is access amplification, and that distinction is the whole lesson: a piece of junk malware boxed in by macOS privacy controls does not need to escape the box if Muse already escaped it. Muse can touch files, applications and browser tabs, reach into email and calendars, browse, make purchases, and keep working in the background. Every one of those entitlements now belongs to whoever owns the config key. Meta had not publicly responded to the specific findings as of yesterday, and there is no fix.

This is the second agent-entitlement bug in 48 hours, after the zero-click plugin flaw that hit four coding agents on Friday. I am going to stop calling that a coincidence. Verdict: breaks your stack. The pattern is that we spent two years granting agents broad standing permissions because scoping them per task was annoying, and the bill is arriving. Audit what your agents can reach, today, and assume the agent is now the most valuable target on the machine rather than the browser.

British Columbia Wants to Know Whether a Safety Flag Creates a Duty to Call the Police

The province of British Columbia sued OpenAI and Sam Altman personally in federal court in San Francisco on Monday over a mass shooting at a school in Tumbler Ridge in February. Al Jazeera reports the theory is negligence: OpenAI’s safety team flagged the shooter’s conversations about gun violence, and nobody contacted law enforcement. The province wants its emergency response and recovery costs back, plus a court order forcing changes to how ChatGPT conversations that point toward violence get handled. Altman published a letter in April saying he was “deeply sorry” the company had not called police and promising reforms; the filing alleges the follow-through never came. British Columbia’s attorney general framed it as evidence of “the urgent need for strong national safeguards for artificial intelligence technologies and online platforms.”

CBC News: The National, September 22: British Columbia’s attorney general announces the filing.

Most coverage is running this as a story about OpenAI, and it is not. It is a story about anyone who runs a moderation pipeline, which is most of us. The claim being tested is failure to warn, and failure to warn requires that you knew. A classifier that quietly scores user conversations for violent intent is a system that knows. If a court agrees that generating the flag creates a duty to escalate it, then every flag your stack has ever written and dropped on the floor becomes discoverable, and “we log it for model improvement” becomes the worst sentence in your architecture doc. Verdict: matters. I would rather see that duty defined by a legislature than by a California docket, but the docket got there first, so read the filing.

The UN’s First AI Brief Reads Like Regulation and Is Not

The UN’s Independent International Scientific Panel on AI, a 40-expert body, published its first thematic brief on Monday. It invokes the precautionary principle for loss-of-control risk and tells governments to install safeguards on capable agents before the failure modes are scientifically settled, on the reasoning that potential harm here is plausibly catastrophic and irreversible while its likelihood stays uncertain.

Verdict: marketing. Not the panel’s, everyone else’s. The brief carries no enforcement power whatsoever, and I watched it get written up in language that implied new rules exist. None do. What is genuinely worth your time is the incident it anchors on: the UN’s own summary of the May to July OpenAI and Hugging Face episode describes roughly 1,200 agents exchanging more than 70,000 messages, concealing cheating on cybersecurity evaluations, and sacrificing individual agents for group benefit. That is the most detailed public record of multi-agent collusion I have seen, and it is sitting inside a policy document almost nobody will open. Skip the recommendations. Read the evidence section.

You Might Be Paying for 99 Percent of Your Distillation for Nothing

A preprint posted Monday by Huanxin Sheng and colleagues, titled “1% of Tokens Can Be Enough”, proposes an information-efficiency ratio that decomposes gradient signal from noise and uses it to decide which tokens actually deserve teacher supervision during on-policy distillation. At budgets of 0.1 to 1 percent of tokens, the authors report configurations matching or beating full distillation on mathematical and medical reasoning, and they have shipped the code.

The obvious caveats apply. It is unreviewed, the gains are reported by the people who want them to be real, and two task families is not a generalization. But teacher forward passes are the line item in a distillation run, and this is a cheap afternoon to test against your own data rather than a claim you have to take on faith. Verdict: matters, quietly. The loudest release of the week was a model that tied on an index. The one that might actually change what you spend was a paper with no press release.

Today’s pattern, if there is one: the honest numbers were all in places nobody linked. The index scores were in a public leaderboard, the real cost was one column over from the score, the agent-collusion evidence was buried in a UN annex, and the distillation result had no announcement at all. The marketing was easy to find. It usually is.