Mistral Large 4: Its Place Beyond The US And China, And Its Limits For Agents
AIThis post was created with the assistance of artificial intelligence (AI).

🔍 Read the full analysis: Mistral Large 4: Its Place Beyond The US And China, And Its Limits For Agents on ThorstenMeyerAI.com

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get office and shipping supplies delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

TL;DR

Mistral has released Large 4 as a research preview, scoring 38.4 on the Artificial Analysis Intelligence Index. That makes it a leading model from outside the United States and China in the cited comparison, but it trails major US and Chinese models and costs more per benchmark task than two Chinese models that score higher. Its long-run agent suitability remains uncertain, particularly given the benchmark results, output volume and reported hands-on hallucinations.

Mistral has released Mistral Large 4, a research-preview model that scored 38.4 on the Artificial Analysis Intelligence Index, a substantial rise from the company’s previous flagship but below the top US and Chinese models in the same comparison. SenseTime’s push into multimodal AI agents offers related context on the expanding AI landscape. The result gives France a prominent position among AI developers outside those two countries, while leaving questions about Large 4’s cost, reliability and suitability for long-running agent tasks, an area also explored in Claude’s agent teams.

Artificial Analysis Index version 4.3.2 places Large 4 below every leading US and Chinese flagship listed in the source material, against the backdrop of the broader US-China AI competition. Its score trails the top-ranked model, Anthropic’s Claude Opus 5.5 at 57.6, by 19.2 points. It is also below several Chinese models, including GLM-5.3 at 44.8, Kimi K3 at 43.6 and DeepSeek V4.1 Flash at 39.5. The source describes Large 4 as the highest-scoring model from outside the US and China in that comparison.

Mistral says Large 4 has 1 trillion total parameters, with 49 billion active, accepts text and images, and has a 512,000-token context window. It is available through the company’s API as a research preview. Mistral has said model weights are expected at the end of October, but the source says the licence has not been published. Listed pricing is $1.36 per million input tokens and $4.18 per million output tokens, with cached input at $0.14; a 50% discount applies for the first two weeks.

The improvement over Mistral’s earlier models is pronounced: the source gives Large 3 a score of 9 and Medium 3.5 a score of 14 on the same Index version. Mistral has also said reinforcement learning is ongoing, meaning results could change. Artificial Analysis estimates Large 4 costs $1.13 per Index task. For comparison, the source lists GLM-5.3-Flash at $0.25 per task with a 41.8 score, and DeepSeek V4.1 Flash at $0.27 with a 39.5 score.

At a glance
reportWhen: Released yesterday, according to the so…
The developmentMistral released Large 4 as a research preview, posting a major score increase over its previous model while remaining behind leading US and Chinese competitors.
Mistral Large 4: Not a Frontier Model — Reality Check
AI Dispatch · Reality Check · 7 October 2026

Mistral Large 4: best outside the US and China — and still not a model to run your agents on

The headline is true: France has the most intelligent model outside the US and China. The independent data says the rest: every US and Chinese flagship scores higher, the best by 19 points. It costs 4× more per task than Chinese open models that outscore it, and it’s 2.5× as verbose as the median model.

Artificial Analysis Intelligence Index v4.3.2 — same version, like for like
Claude Opus 5.5 US57.6
Claude Sonnet 5.5 US56.0
Claude Fable 5.1 US53.4
GPT-6 Astra US52.7
Gemini 4 Argon US52.6
GPT-6.1 Sol US51.8
GLM-5.3 CN · open44.8
Kimi K3 CN · open43.6
GLM-5.3-Flash CN · open41.8
DeepSeek V4.1 Flash CN · open39.5
Mistral Large 4 (Preview) FR38.4
GPT-6 Luna US · small model~38
DeepSeek V4 Pro 0813 CN36.0
GLM-5.2 CN33.7
vs US frontier
−19.2 pts

~two-thirds of Opus 5.5. Level with OpenAI’s small model, Luna.

vs China open
8th

Eighth among open models once weights ship — behind seven Chinese ones. Beats GLM-5.2 and V4 Pro, loses to their successors.

vs Canada
n/a

Cohere doesn’t compete at this tier — reported ~14% hallucination at ~9% accuracy, because it declines most questions. A field of one.

The cost problem is worse than the intelligence problem — $ per Index task
Mistral Large 4
$1.13
Index 38.4 · $0.57 launch promo
GLM-5.3-Flash
$0.25
Index 41.8 · 4.5× cheaper
DeepSeek V4.1 Flash
$0.27
Index 39.5 · 4.2× cheaper
Gemini 4 Argon
~$1.99
Index 52.6 · +14 points
Per-token pricing looks competitive ($4.18/M output, well under the $10 median) — but it burns 200M output tokens on the Index vs an 81M median. Cheap tokens × 2.5 as many tokens is not a cheap model.
Why not for agentic or long-running work
The gap compounds
19 pts behind

The Index is now agentic-heavy — Briefcase, GDPval, AutomationBench, Terminal-Bench. Errors multiply across steps: tolerable in chat, fatal over a two-hour run.

AA v4.3.2
Verbosity
200M vs 81M

Output tokens to complete the Index. On an agent, verbosity is cost and latency on every step.

AA
Hallucination is back
observed

Confident false assertions in hands-on use. US frontier has largely moved past this — Gemini 4 Argon: 15%. In fairness Chinese open models are worse (Kimi K3 51%, DeepSeek V4 Pro 94%). In an agent, a fabrication is a wrong premise every later step builds on.

AUTHOR’S TESTING · not an AA figure
✓ What it’s genuinely good at
  • Cyber defence: 50 on the AA Cyber Index; 82% CyberGym-E2E (ahead of Luna’s 78%). Likely top-3 open model on cyber.
  • Documents & images: 19% GDP.pdf (+18 vs Large 3); 100 images per request.
  • Speed: 116 tok/s, 1.46s TTFT — well above median.
  • The jump: Large 3 scored 9 on this Index. 9 → 38 is real progress.
  • Jurisdiction: French parent, EU hosting, weights promised end of October.
▸ Who should actually use it
  • Legally bound buyers (defence, classified, DORA, health data): now the best European option by a wide margin. Wait for the weights, check the licence, pilot on cyber and documents.
  • Everyone else, for agentic or long tasks: don’t. A US frontier model is meaningfully more capable; GLM-5.3-Flash is more capable and 4× cheaper.
  • Note: Preview — Mistral says RL is still running, so scores may move. That changes next month’s decision, not today’s.
The take

Mistral says it has “essentially closed the gap.” It has closed the gap to where the Chinese open-weights field was a few months ago, while that field and the US frontier have both moved on. On every independent measure that matters for agents — intelligence, cost per task, verbosity and factual reliability — Large 4 is not a frontier model. “Most intelligent outside the US and China” is true mainly because almost nobody else outside those two countries is competing. Use it if you have to. Don’t use it because of the headline.

Sources: Artificial Analysis — Mistral Large 4 article & model/provider pages (6 Oct 2026), Index v4.3.2, comparison data; Trending Topics independent-ranking analysis; AA-derived reporting for frontier scores and AA-Omniscience rates (Argon 15%, Kimi K3 51%, DeepSeek V4 Pro 94%); Cohere profile as reported by Suprmind. Mistral Large 4’s AA-Omniscience result isn’t published in text — the hallucination point is the author’s own testing. Preview scores may change. Not investment advice.
thorstenmeyerai.com

The Cost of Agent Work

The Index includes benchmarks for agentic knowledge work, software workflows and coding, so its score speaks to tasks involving multiple steps, not only short question-and-answer exchanges. A lower result does not directly quantify how often a particular customer workflow will fail, but it is relevant evidence for organizations considering Large 4 as an autonomous worker.

The source also reports that Large 4 generated 200 million output tokens across the Index, compared with a median of 81 million for comparable models. That figure indicates substantially higher output volume in the benchmark; in deployed agent workflows, more generated tokens can add cost and time, depending on the task and how a system is configured. The source’s task-cost estimates also put two lower-scoring alternatives below Large 4’s cost, complicating the case for choosing it on price-performance grounds.

For European buyers seeking a model from outside the US and China, Large 4 may still be relevant as a regional option. But the source’s rankings do not establish that it is the strongest practical choice for every workload. Enterprise requirements, data handling, latency, reliability and the final model licence will also affect procurement decisions.

Amazon

AI model API access

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

From Earlier Mistral Scores

The release follows a period in which Mistral’s previous models scored well below the current result on the same Artificial Analysis Index version. Large 4’s move from Large 3’s 9 to 38.4 marks a considerable reported gain. The source characterizes it as a major step for a European lab, while cautioning that a large improvement does not place the model at the top of the broader field.

The phrase “most intelligent model outside the US and China” depends on the comparison set. In the cited results, Large 4 sits above models from other regions, but the source argues that the closest competitors at the frontier are concentrated in the US and China. It also says Cohere is not a direct match for this comparison because its enterprise model is oriented toward retrieval and tool use rather than frontier reasoning. That is the source’s assessment, not a universal measure of model quality.

Artificial Analysis’ Index combines results from several task categories, including AA-Briefcase, GDPval-AA, AutomationBench and Terminal-Bench 4.0. Because benchmarks measure selected tasks under defined conditions, their rankings can inform comparisons without predicting performance in every customer’s live environment.

“Reinforcement learning is still running.”

— Mistral

Preview, Licence and Reliability

Several details remain unsettled. Mistral’s expected end-of-October weights release has not yet happened, and the source says the licence is unpublished. It is not clear what restrictions will apply, whether the release date will hold, or whether the weights will match the API preview.

The source says reinforcement learning is ongoing, so benchmark scores may shift. It does not provide a controlled hallucination rate for Large 4. Its account of confident false claims comes from hands-on testing by the author and should not be treated as a measured frequency or a finding that applies to every prompt. Nor do the benchmark figures establish how the model will perform across all production agent systems.

The task-cost figures are estimates tied to the Artificial Analysis Index evaluation. Actual costs can vary with prompt length, output volume, caching, discounts and workload design. The source material also ends before completing its comparison of Large 4’s price with Gemini 4 Argon, so no conclusion about that comparison is included here.

Weights and Further Testing

The next reported milestone is Mistral’s planned release of Large 4 weights at the end of October. Buyers will be able to assess the release more fully once the weights and licence terms are public, rather than relying only on an API preview. Mistral’s ongoing reinforcement learning may also lead to updated results, though the company has not supplied a revised score in the source material.

For teams evaluating Large 4 for agents, the practical next step is to test it against their own tasks, tracking completion rates, factual errors, latency and total token costs across multi-step runs. The current evidence supports a clear conclusion about the benchmark snapshot: Large 4 has improved sharply, but it remains behind several leading alternatives, and the preview data alone does not resolve whether it is dependable or economical for a specific workflow.

Key Questions

What did Mistral announce?

Mistral released Large 4 as a research preview through its API. The source describes it as a text-and-image input model with a 512,000-token context window.

How does Large 4 rank in the cited benchmark?

It scored 38.4 on Artificial Analysis Intelligence Index version 4.3.2. That is below the leading US and Chinese models listed, while the source identifies it as the highest-scoring model from outside those two countries in its comparison.

Are Large 4’s weights publicly available?

Not yet, according to the source. Mistral has promised weights for the end of October, but the licence has not been published in the material provided.

Is Large 4 a good choice for AI agents?

The available evidence does not settle that for every use case. Its benchmark score covers agent-related tasks, while the source also reports high output-token volume and an author’s observations of confident errors. Teams should test it on their own workflows and compare cost, accuracy and completion rates.

Source: ThorstenMeyerAI.com

This content is for general information only and is not financial, tax or legal advice. Consult a qualified professional for decisions about your money.
HALLOWEEN

Halloween Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

The Memory Squeeze: Why Your RAM Bill Doubled

DRAM prices have surged up to 600%, driven by AI-focused chip reallocation, with supply constrained and new capacity delayed until 2027–2028.

The Real Price Of AI Quantization To Four Bits

New research reveals that quantizing language models to four bits causes significant, selective performance loss, especially in reasoning and math capabilities.

A Nudge Here And There, Rather Than Heavy Handed Interventions To Unleash Growth: ALEX BRUMMER

Andy Burnham’s proposed First Home scheme drew a positive response from housebuilders as UK home completions fell and builders cut targets.

Aleph Alpha. The retrospective case.

Analyzing Aleph Alpha’s strategic pivot, funding, and acquisition to understand the pitfalls of late structural adaptation in European sovereign AI development.