Mistral AI put online this Tuesday 6 October the public preview of Mistral Large 4, which the company itself nicknames "the Chonk". The model has 1,050 billion parameters, of which 49 billion are activated per token, and accepts text as well as image input. It is served from Mistral's European data centres, via the Mistral Studio API, at $1.36 per million input tokens and $4.18 per million output tokens. The weights will follow "by the end of the month" according to the announcement post, on 27 October according to The Next Web.
Mistral itself situates the model in one sentence. Large 4 would be "competitive with the strongest open source models globally" and would "clearly surpass any open-weight model developed in the United States or Europe". Speaking to CNBC, the company talks about the most capable open model outside China, "with a substantial margin". The two formulations do not say the same thing, and independent measurements do not validate them to the same degree.
3,800 Grace Blackwell GPUs and a generation leap from Mistral Large 3
Large 4 was trained from scratch on 3,800 Nvidia Grace Blackwell GPUs installed in the data centres that Mistral owns in Europe, over approximately two months according to the company's statement reported by CNBC and The Next Web. The technical sheet describes a so-called granular mixture of experts, a 1.6 billion parameter vision encoder and a context window of one million tokens, which Artificial Analysis measures for its part at 524,000. The training data covers more than 160 languages, including all the official languages of the European Union.
The previous flagship model, Mistral Large 3, released on 2 December 2025, had 675 billion parameters of which 41 billion were active. It had been trained on 3,000 H200 GPUs and published the same day under the Apache 2.0 licence, its "reasoning" version being then announced for later. In ten months, Mistral has changed chip generation, increased total size by 56% and merged instruction and reasoning into a single hybrid model.
On 6 October, the same technical sheet displays the launch prices crossed out, replaced by $0.68 input and $2.09 output, i.e. half the announcement prices, with no indication of the duration of this discount. The licence of the weights, however, does not appear there. The sheet bears the mention "Open" without naming a licence, whereas the same documentation displays Apache 2.0 for Large 3 and a modified MIT licence for Medium 3.5.
An independent index at 38 points, eight below the best open model
The results published by Mistral have not yet been replicated. On DeepSWE v1.1, an agentic programming test, Large 4 obtains 61.7%, ahead of GLM-5.3 and DeepSeek V4 Pro in the announcement graphs. The table published the day before by Reflection AI for its own model takes the same values for GLM-5.3 and Qwen 3.8 Max, but adds Kimi K3 at 68% and DeepSeek V4.1 Flash at 74.2%, two models absent from Mistral's graph. CNBC notes that the model "remains behind the frontier" in areas such as code.
The most readable measurement comes from Artificial Analysis, which runs its own evaluations. Its intelligence index, consulted on 6 October, gives the following positions:
- closed models: Claude Opus 5.5, 58 points; GPT-6 Astra, 53;
- Chinese open weights: MiMo-V2.6-Pro from Xiaomi, 46; GLM-5.3, 45; Kimi K3, 44; GLM-5.3-Flash, 42; Qwen3.8 2.4T, 40; Qwen3.8-Flash-Next, 40; DeepSeek V4.1 Flash, 39;
- Mistral Large 4 Preview, 38;
- DeepSeek V4 Pro, 36;
- open weights outside China: Motif 3 (South Korea), 34; Inkling Small from Thinking Machines (United States), 26; Mistral Medium 3.5, 14; gpt-oss-120b from OpenAI, 12; Mistral Large 3, 9.
On this index, Large 4 is the eighth open-weight model, eight points behind MiMo-V2.6-Pro and behind six other Chinese models. "Competitive with the strongest" is defensible at the scale of a generation rather than a ranking. The second half of the sentence holds without reservation, with a twelve-point lead over the best measured American open model. The formula given to CNBC, "outside China", is however played out at four points from the Korean Motif 3. Beam, presented on 5 October by Reflection AI, was not yet measured. The clearest leap remains internal. Mistral Large 3, a model without reasoning on an index that favours reasoning, stood at 9. Mistral specifies that reinforcement training "is still ongoing", that the model "shows no sign of saturation" and that this first stage is financed by the €3 billion fundraising of September. The preview figures are therefore provisional.
Artificial Analysis moreover classifies Large 4, for the moment, among proprietary models, for lack of published weights. And the documentation that showcases Large 4 still lists, among its generalist models and with the mention "third-party", GLM 5.3 from Z.ai, which will replace GLM 5.2 on the platform upon its withdrawal on 31 October. Mistral had begun serving GLM-5.2 on its own servers in August, a configuration that ActuIA described in early September as that of a Europe that assembles what China provides.
Cybersecurity, the ground where several closed models refuse the task
Mistral builds its commercial argument on a specific domain, cybersecurity, and it is there that the independent figures are most favourable to it. On the Cyber Index of Artificial Analysis, which combines three evaluations of discovery, reproduction and correction of vulnerabilities in real code repositories, Large 4 Preview obtains 49.5 points and ranks fifth of the 18 models evaluated on 6 October. Ahead of it, Grok 4.7 (56.4), MiMo-V2.6-Pro (56.1), GPT-6 Luna (52.7) and GLM-5.3-Flash (49.9). Among the following, Muse Spark 1.3 from Meta (44.3), Kimi K3 (41.3), GLM-5.3 (36.4), GPT-6 Astra (33.4) and Claude Opus 5.5 (28.7).
The detail of the CyberGym-E2E test, where one must reproduce a real vulnerability and then correct it, explains the gap. Large 4 succeeds in 81.7% of the tasks, the best score of the 18 models. GPT-6 Astra is at zero with a refusal rate of 100%, Claude Opus 5.5 at 0.8% with 98.5% refusals. The refusal stems from a publisher policy, since Grok 4.7 and GPT-6 Luna, also closed, handle the task. Mistral draws its argument from this. "Provider-level refusals can block legitimate vulnerability research and incident response," the company writes, and losing access to a capability in the middle of an incident "can itself become a critical security risk".
The argument has a documented precedent. In its account of the intrusion carried out by an autonomous agent in July, Hugging Face explains that it first submitted its forensic analysis to commercial APIs, before being blocked by "the safety guardrails of providers, which cannot distinguish an incident responder from an attacker". The company ended up running the analysis on GLM-5.2, an open-weight model, on its own infrastructure, while specifying that this was "not an argument against the safety measures of hosted models". Guillaume Lample, co-founder and chief scientist of Mistral, goes further in the statement transmitted to CNBC. Cyber defence capabilities "will enable companies and governments to defend themselves against malicious actors who jailbreak closed models to conduct cyberattacks".
Mistral simultaneously maintains a discourse of prudence. Before the publication of the weights, the model is tested in real conditions by cybersecurity managers, approved partners and state authorities, who access a version "with reduced moderation and extended cyber capabilities". The boundary between the public version and the "extended" version will however amount to little once the weights are downloadable, since an open model can be retrained.
Sovereignty, licence and timeline, the three questions for businesses
That same morning, before the announcement, Arthur Mensch took part in a discussion entitled "Sovereign, Open, Safe: Pick Two?" at the AI Everything show in Abu Dhabi. The CEO of Mistral estimated there that "99% of use cases" handled by the company can be handled with open models and that there is "no reason to pay the API margin of closed models". "The narrative that Europe cannot compete is simply not true," he added. The post promises a European region operated end-to-end by Mistral "under European law", then a deployment on private cloud or on-premises with the weights.
Three points remain to be settled for a buyer. The licence, absent from the sheet on 6 October, will decide whether Large 4 can be reused as freely as Large 3. The operating cost remains unknown, Mistral not having published a reference hardware configuration for a model of 1,050 billion parameters. The timeline, for its part, is shifting. Mistral indicates that the computing capacity financed by the September fundraising "comes online in the coming months" and announces "significant and rapid improvements in the weeks and months to come".
With the weights, Mistral promises details on the architecture, complementary benchmarks and its post-training method. The licence does not appear in this list.
Our articles will then appear first in Google Top Stories.
