Mistral Large 4 Ships 1T Params, Costs 4x Rivals

Mistral released Mistral Large 4 in public preview on October 6, 2026 — a 1.05-trillion-parameter mixture-of-experts model with 49 billion active parameters, priced at $1.36 per million input tokens and $4.18 per million output. Artificial Analysis scored it 38 on its Intelligence Index, calling it the best model from outside the US and China, but measured its cost per task at more than four times comparable open-weights rivals.

Key takeaways

  • 1.05T total parameters, 49B active — about 4.7% fire per token.
  • Artificial Analysis Intelligence Index: 38, versus DeepSeek V4.1 Flash at 39.
  • $1.13 per index task against $0.25 for GLM-5.3-Flash.

What is Mistral Large 4?

Mistral Large 4 is the French lab’s new flagship model, internally abbreviated ML4 and nicknamed “le Chonk” in Mistral’s own announcement. It is a granular mixture-of-experts system with roughly 1.05 trillion total parameters and 49 billion active per token, natively multimodal, and fluent in more than 160 languages.

It shipped October 6 as a research public preview on Mistral’s API, served from the company’s European datacenters.

Mistral said it trained the model from scratch on 3,800 Nvidia Grace Blackwell accelerators. Each one pairs two Blackwell GPUs with a Grace CPU. The reinforcement-learning stack generated 33 billion tokens a day in parallel trial-and-error rollouts, according to Mistral, with just under half feeding the training workflow.

A 1.6-billion-parameter vision encoder handles image input. Expert count, routing and layer layout have not been published.

How much does Mistral Large 4 cost?

Standard API pricing is $1.36 per million input tokens and $4.18 per million output tokens, with cached input at $0.14 per million. A launch discount cuts that in half for the first two weeks, to $0.68 and $2.09. That promotional window closes around October 20.

Those headline numbers look competitive. The per-task numbers do not.

Artificial Analysis measured $1.13 to run its full Intelligence Index on Mistral Large 4 at list price, or $0.57 at the discount. GLM-5.3-Flash ran the same gauntlet for $0.25 and DeepSeek V4.1 Flash for $0.27. The reason is verbosity: the Mistral preview emitted 200 million output tokens during the run against a class median of 81 million.

Mistral Large 4 pricing against the open-weights field

ModelTotal paramsActiveInput / 1MOutput / 1MWeights
Mistral Large 41.05T49B$1.36$4.18Due end of October 2026
GLM-5.3Not publishedNot published$1.40$4.40Released
Kimi K32.8T~104B$3.00$15.00Released July 27, 2026
DeepSeek V4 Pro1.6T49BNot disclosedNot disclosedReleased, MIT
Sources: Mistral announcement, MarkTechPost model comparison, Artificial Analysis.

Mistral is pricing within a few cents of GLM-5.3 and well under Kimi K3. Against the cheap Chinese flash tiers, it is not close.

Is Mistral Large 4 better than GPT-6 or DeepSeek?

On aggregate intelligence, it is level with one frontier tier and a point behind another. Artificial Analysis put Mistral Large 4 Preview at 38 on version 4.3.2 of its Intelligence Index, matching GPT-6 Luna at maximum reasoning and trailing DeepSeek V4.1 Flash at 39. It ranked 64th of 225 models, against a class median of 26.

Artificial Analysis called it “the most intelligent AI model from outside the US and China.” That is a real milestone and a narrow one.

Mistral’s own benchmark card leans hard on security work. The company reports 93% on Cybench and 82% on a reproduce-and-patch test, the highest figure it lists. On the independent Artificial Analysis Cyber Index, though, the model scores 50 — tied with GLM-5.3-Flash and behind Xiaomi’s MiMo-V2.6-Pro at 56.

Coding results are the weak spot

Mistral reports 61.7% on DeepSWE v1.1, 59.4% on SWE-Atlas-QnA, 49.8% on the Coding Agent Index, and 28.3% on Terminal-Bench 4. That last number is the one to sit with. For comparison, Claude Sonnet 5.5 posted 70.6% on Terminal-Bench at launch.

In a Surge AI blind human evaluation, Mistral Large 4 scored 3.74 out of 5, placing second of five models behind Claude Opus 5 at 4.22. On agentic work it reports 59.9% on AutomationBench and 1,393 Elo on AA-Briefcase. On vision it claims 42% on Dense 200 against 41% for GPT-6-Astra — a one-point lead inside the noise band of most evaluations.

One caveat worth flagging: the coding scores were evaluated privately and are not yet independently reproducible. Mistral says architecture, benchmark and post-training detail arrives with the weights.

What is still missing from the launch?

Three things, and each of them matters to anyone planning a budget around this model. The weights are not out, the license has no name, and the context window is reported inconsistently across Mistral’s own materials and the platforms serving it.

  • Weights. Due by the end of October 2026. Until then there is no self-hosting, which is the entire reason enterprises care about an open-weight frontier model.
  • License. Unnamed. DeepSeek V4 Pro ships under MIT and Kimi K3 under a modified MIT. A restrictive Mistral license would change the economics completely.
  • Context window. Mistral’s materials indicate 1 million tokens; Artificial Analysis lists 512k and OpenRouter’s hosted endpoint shows 524K.
  • Pricing page. As of this writing, Mistral’s public pricing page does not list Large 4 at all. It still shows Mistral Large at $0.5 in and $1.5 out.

A preview that announces a price the vendor’s own pricing page has not absorbed is a preview, not a product. Treat the $1.36 figure as a research-preview rate until it appears there.

Who wins and who loses?

Mistral wins a sovereignty argument, not a price war. The company closed a €3 billion Series D at a €21 billion post-money valuation in September 2026, and it calls Large 4 the first milestone funded by that round. For European buyers who must keep inference inside the EU, there is now a model at the 38-point intelligence tier that runs there end to end.

That is a procurement moat, and procurement moats are worth real money. They are not the same thing as a cost advantage.

The winners

Nvidia. Another 3,800 Grace Blackwell accelerators of demonstrated demand, from a lab whose entire pitch is European independence. Sovereignty does not extend to silicon.

European regulated buyers. Banks, insurers and health systems that could not legally route prompts to a US API now have a credible frontier-tier option with resident inference.

Security teams. If the 82% patch figure survives independent testing, Mistral has the strongest open-weight story in vulnerability remediation — and that work can run on-premise once weights land.

The losers

Anyone buying on cost per task. At $1.13 against $0.25, the verbosity tax is not a rounding error. High-volume inference workloads will keep routing to the Chinese flash tiers.

Coding-tool vendors betting on Mistral. A 28.3% Terminal-Bench 4 score does not displace anything in the agentic coding market right now.

The open-weights pricing floor. Rivals like Reflection’s Beam and the newly capitalized DeepSeek keep compressing what buyers will pay for this intelligence tier. Mistral just priced above that floor and will have to defend the gap on compliance alone.

What to watch next

  • End of October 2026: weights release, plus the architecture, benchmark and post-training detail Mistral deferred. The license name lands here too.
  • Around October 20: the two-week 50% launch discount expires, moving effective pricing from $0.68/$2.09 back to $1.36/$4.18.
  • Whenever weights drop: whether independent reproduction confirms the 61.7% DeepSWE and 82% patch figures evaluated privately by Mistral.
  • No date given: Mistral says it did not pause training after Large 4 and expects larger versions “in the coming months,” followed by use-case-specific models built on this base.

Frequently asked questions

How many parameters does Mistral Large 4 have?

About 1.05 trillion total parameters with 49 billion active per token, roughly 4.7%. It is a granular mixture-of-experts model with a 1.6-billion-parameter vision encoder for native image input.

When are the Mistral Large 4 weights released?

Mistral says by the end of October 2026. The license has not been named. Until the weights ship, the model is available only through Mistral’s API preview.

What does Mistral Large 4 cost per million tokens?

$1.36 input and $4.18 output, with cached input at $0.14. A 50% launch discount applies for the first two weeks, bringing it to $0.68 and $2.09.

Is Mistral Large 4 the best open-weight model?

Not on intelligence. Artificial Analysis scored it 38, behind DeepSeek V4.1 Flash at 39, and called it the strongest model from outside the US and China rather than the strongest overall.

How fast is Mistral Large 4?

Artificial Analysis measured 116.1 output tokens per second on Mistral’s API, against a class median of 86.1, with time to first token of 1.46 seconds versus a 3.81-second median.

What hardware trained Mistral Large 4?

Mistral says 3,800 Nvidia Grace Blackwell accelerators in its European datacenters. Training duration was not disclosed. Reinforcement-learning rollouts produced 33 billion tokens per day.

Can I run Mistral Large 4 on-premise?

Not yet. Mistral says self-deployment becomes possible once weights are released at the end of October, including private-cloud and on-premise deployment for security use cases.

The bottom line

Mistral Large 4 is the best non-US, non-Chinese model anyone has shipped, and that sentence carries more weight in a procurement meeting than in a benchmark table. The 38-point Intelligence Index score is real. So is the four-times cost gap against open-weights peers at the same tier.

The investable thesis here is not performance. It is that EU data-residency rules create a buyer who cannot shop on price, and Mistral has just become the only vendor serving that buyer at this capability level from inside the bloc. A €21 billion valuation needs that moat to hold.

For everyone else, wait for the weights and the license. A frontier open-weight model you cannot download yet, under terms nobody has read, priced above its peers, is an announcement — not a decision.

Free guide · 12 pages

AI Tools Every Investor Should Use in 2026

10 tools for research, stock signals, charts and crypto, what each one costs, three ready-made stacks from $0 a month and 7 copy-paste prompts for filings and earnings calls. Enter your email and we will send it to you.

Plus one email a week: AI Money This Week, every Sunday. Unsubscribe anytime.

Sources

Wealth Engine researches and drafts with AI tools and checks every figure against the sources above. How we report.

Comments

Leave a Reply

Discover more from Wealth Engine

Subscribe now to keep reading and get access to the full archive.

Continue reading