🔍 Read the full analysis: Mistral Large 4: Its Place Beyond The US And China, And Its Limits For Agents on ThorstenMeyerAI.com
Get office and shipping supplies delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
TL;DR
Mistral has released Large 4 as a research preview, scoring 38.4 on the Artificial Analysis Intelligence Index. That makes it a leading model from outside the United States and China in the cited comparison, but it trails major US and Chinese models and costs more per benchmark task than two Chinese models that score higher. Its long-run agent suitability remains uncertain, particularly given the benchmark results, output volume and reported hands-on hallucinations.
Mistral has released Mistral Large 4, a research-preview model that scored 38.4 on the Artificial Analysis Intelligence Index, a substantial rise from the company’s previous flagship but below the top US and Chinese models in the same comparison. SenseTime’s push into multimodal AI agents offers related context on the expanding AI landscape. The result gives France a prominent position among AI developers outside those two countries, while leaving questions about Large 4’s cost, reliability and suitability for long-running agent tasks, an area also explored in Claude’s agent teams.
Artificial Analysis Index version 4.3.2 places Large 4 below every leading US and Chinese flagship listed in the source material, against the backdrop of the broader US-China AI competition. Its score trails the top-ranked model, Anthropic’s Claude Opus 5.5 at 57.6, by 19.2 points. It is also below several Chinese models, including GLM-5.3 at 44.8, Kimi K3 at 43.6 and DeepSeek V4.1 Flash at 39.5. The source describes Large 4 as the highest-scoring model from outside the US and China in that comparison.
Mistral says Large 4 has 1 trillion total parameters, with 49 billion active, accepts text and images, and has a 512,000-token context window. It is available through the company’s API as a research preview. Mistral has said model weights are expected at the end of October, but the source says the licence has not been published. Listed pricing is $1.36 per million input tokens and $4.18 per million output tokens, with cached input at $0.14; a 50% discount applies for the first two weeks.
The improvement over Mistral’s earlier models is pronounced: the source gives Large 3 a score of 9 and Medium 3.5 a score of 14 on the same Index version. Mistral has also said reinforcement learning is ongoing, meaning results could change. Artificial Analysis estimates Large 4 costs $1.13 per Index task. For comparison, the source lists GLM-5.3-Flash at $0.25 per task with a 41.8 score, and DeepSeek V4.1 Flash at $0.27 with a 39.5 score.
Mistral Large 4: best outside the US and China — and still not a model to run your agents on
The headline is true: France has the most intelligent model outside the US and China. The independent data says the rest: every US and Chinese flagship scores higher, the best by 19 points. It costs 4× more per task than Chinese open models that outscore it, and it’s 2.5× as verbose as the median model.
~two-thirds of Opus 5.5. Level with OpenAI’s small model, Luna.
Eighth among open models once weights ship — behind seven Chinese ones. Beats GLM-5.2 and V4 Pro, loses to their successors.
Cohere doesn’t compete at this tier — reported ~14% hallucination at ~9% accuracy, because it declines most questions. A field of one.
The Index is now agentic-heavy — Briefcase, GDPval, AutomationBench, Terminal-Bench. Errors multiply across steps: tolerable in chat, fatal over a two-hour run.
AA v4.3.2Output tokens to complete the Index. On an agent, verbosity is cost and latency on every step.
AAConfident false assertions in hands-on use. US frontier has largely moved past this — Gemini 4 Argon: 15%. In fairness Chinese open models are worse (Kimi K3 51%, DeepSeek V4 Pro 94%). In an agent, a fabrication is a wrong premise every later step builds on.
AUTHOR’S TESTING · not an AA figure- Cyber defence: 50 on the AA Cyber Index; 82% CyberGym-E2E (ahead of Luna’s 78%). Likely top-3 open model on cyber.
- Documents & images: 19% GDP.pdf (+18 vs Large 3); 100 images per request.
- Speed: 116 tok/s, 1.46s TTFT — well above median.
- The jump: Large 3 scored 9 on this Index. 9 → 38 is real progress.
- Jurisdiction: French parent, EU hosting, weights promised end of October.
- Legally bound buyers (defence, classified, DORA, health data): now the best European option by a wide margin. Wait for the weights, check the licence, pilot on cyber and documents.
- Everyone else, for agentic or long tasks: don’t. A US frontier model is meaningfully more capable; GLM-5.3-Flash is more capable and 4× cheaper.
- Note: Preview — Mistral says RL is still running, so scores may move. That changes next month’s decision, not today’s.
Mistral says it has “essentially closed the gap.” It has closed the gap to where the Chinese open-weights field was a few months ago, while that field and the US frontier have both moved on. On every independent measure that matters for agents — intelligence, cost per task, verbosity and factual reliability — Large 4 is not a frontier model. “Most intelligent outside the US and China” is true mainly because almost nobody else outside those two countries is competing. Use it if you have to. Don’t use it because of the headline.
The Cost of Agent Work
The Index includes benchmarks for agentic knowledge work, software workflows and coding, so its score speaks to tasks involving multiple steps, not only short question-and-answer exchanges. A lower result does not directly quantify how often a particular customer workflow will fail, but it is relevant evidence for organizations considering Large 4 as an autonomous worker.
The source also reports that Large 4 generated 200 million output tokens across the Index, compared with a median of 81 million for comparable models. That figure indicates substantially higher output volume in the benchmark; in deployed agent workflows, more generated tokens can add cost and time, depending on the task and how a system is configured. The source’s task-cost estimates also put two lower-scoring alternatives below Large 4’s cost, complicating the case for choosing it on price-performance grounds.
For European buyers seeking a model from outside the US and China, Large 4 may still be relevant as a regional option. But the source’s rankings do not establish that it is the strongest practical choice for every workload. Enterprise requirements, data handling, latency, reliability and the final model licence will also affect procurement decisions.
As an affiliate, we earn on qualifying purchases.
From Earlier Mistral Scores
The release follows a period in which Mistral’s previous models scored well below the current result on the same Artificial Analysis Index version. Large 4’s move from Large 3’s 9 to 38.4 marks a considerable reported gain. The source characterizes it as a major step for a European lab, while cautioning that a large improvement does not place the model at the top of the broader field.
The phrase “most intelligent model outside the US and China” depends on the comparison set. In the cited results, Large 4 sits above models from other regions, but the source argues that the closest competitors at the frontier are concentrated in the US and China. It also says Cohere is not a direct match for this comparison because its enterprise model is oriented toward retrieval and tool use rather than frontier reasoning. That is the source’s assessment, not a universal measure of model quality.
Artificial Analysis’ Index combines results from several task categories, including AA-Briefcase, GDPval-AA, AutomationBench and Terminal-Bench 4.0. Because benchmarks measure selected tasks under defined conditions, their rankings can inform comparisons without predicting performance in every customer’s live environment.
“Reinforcement learning is still running.”
— Mistral
Preview, Licence and Reliability
Several details remain unsettled. Mistral’s expected end-of-October weights release has not yet happened, and the source says the licence is unpublished. It is not clear what restrictions will apply, whether the release date will hold, or whether the weights will match the API preview.
The source says reinforcement learning is ongoing, so benchmark scores may shift. It does not provide a controlled hallucination rate for Large 4. Its account of confident false claims comes from hands-on testing by the author and should not be treated as a measured frequency or a finding that applies to every prompt. Nor do the benchmark figures establish how the model will perform across all production agent systems.
The task-cost figures are estimates tied to the Artificial Analysis Index evaluation. Actual costs can vary with prompt length, output volume, caching, discounts and workload design. The source material also ends before completing its comparison of Large 4’s price with Gemini 4 Argon, so no conclusion about that comparison is included here.
Weights and Further Testing
The next reported milestone is Mistral’s planned release of Large 4 weights at the end of October. Buyers will be able to assess the release more fully once the weights and licence terms are public, rather than relying only on an API preview. Mistral’s ongoing reinforcement learning may also lead to updated results, though the company has not supplied a revised score in the source material.
For teams evaluating Large 4 for agents, the practical next step is to test it against their own tasks, tracking completion rates, factual errors, latency and total token costs across multi-step runs. The current evidence supports a clear conclusion about the benchmark snapshot: Large 4 has improved sharply, but it remains behind several leading alternatives, and the preview data alone does not resolve whether it is dependable or economical for a specific workflow.
Key Questions
What did Mistral announce?
Mistral released Large 4 as a research preview through its API. The source describes it as a text-and-image input model with a 512,000-token context window.
How does Large 4 rank in the cited benchmark?
It scored 38.4 on Artificial Analysis Intelligence Index version 4.3.2. That is below the leading US and Chinese models listed, while the source identifies it as the highest-scoring model from outside those two countries in its comparison.
Are Large 4’s weights publicly available?
Not yet, according to the source. Mistral has promised weights for the end of October, but the licence has not been published in the material provided.
Is Large 4 a good choice for AI agents?
The available evidence does not settle that for every use case. Its benchmark score covers agent-related tasks, while the source also reports high output-token volume and an author’s observations of confident errors. Teams should test it on their own workflows and compare cost, accuracy and completion rates.
Source: ThorstenMeyerAI.com
Halloween Picks
halloween
As an affiliate, we earn on qualifying purchases.
