Mistral Large 4: architecture, benchmarks and open weights

On 6 October Mistral released a preview of Mistral Large 4: a one-trillion-parameter Mixture of Experts with 52 billion active parameters, multimodal, trained in European data centres. Weights are announced by the end of October. The declared benchmarks, the Artificial Analysis score, pricing and what it takes to run it on premise.

AIAIMistralLLMOpen weightMixture of ExpertsBenchmarkOn-premise

Article

Four figures on Mistral Large 4
Data from Mistral AI, Artificial Analysis and t3n. Sources at the end.

On 6 October Mistral AI introduced Mistral Large 4, its new flagship model. For now it is available as a public preview through the Mistral Studio API. The weights are announced by the end of October, together with details on the architecture and post-training. According to Mistral it significantly outperforms the open-weight models developed in the United States and Europe.

Architecture

  • Mixture of Experts with one trillion parameters in total, of which 52 billion are active per token.
  • A single model for direct answers and for reasoning.
  • Multimodal input: text and images, with text output.
  • Over 160 languages, including all official languages of the European Union.
  • Trained from scratch on 3,800 NVIDIA Grace Blackwell GPUs in Mistral’s European data centres.

Mistral says Large 4 will be the foundation for a new generation of specialised and optimised models.

Declared benchmarks

The figures published by Mistral focus on agentic coding, automation and security:

  • 61.7% on DeepSWE v1.1 and 28.3% on Terminal-Bench 4;
  • 59.9% on AutomationBench;
  • 93% of the Cybench challenges and a place among the top five models of the AA Cyber Index, with 82% on its vulnerability reproduction test;
  • 93.3% of attacks resisted on the B3 benchmark.

These are the vendor’s measurements, and comparisons with other models hold only with the same harness and conditions. On DeepSWE, for example, Google declared 77.9% for Gemini 4 Argon: the two figures come from two different vendors.

Artificial Analysis, which evaluates models independently, gives the preview 38 points on its Intelligence Index. According to t3n, Large 3 had 9 and Large 3.5 had 14; at 38 points Large 4 sits at the level of GPT-6 Luna and DeepSeek V4.1 Flash, while Claude Opus 5.5, GPT-6 Astra and Gemini 4 Argon score above 50. Artificial Analysis also notes that the model is very verbose: during the evaluation it produced 200 million output tokens (the median is 81 million), which weighs on the actual cost of each task.

Pricing and availability

  • $1.36 per million input tokens and $4.18 per million output tokens; cached input is 90% cheaper according to Artificial Analysis.
  • Preview context window: 524,000 tokens, according to Artificial Analysis.
  • Licence: Mistral has not stated it yet. Until the weights are public, Artificial Analysis classifies the model as proprietary.
  • Mistral writes that the model will be able to run on private cloud or on premise, for organisations that need sovereign, auditable AI in security operations.

What it takes to run it on premise

In a Mixture of Experts the active parameters reduce the compute per token, not the memory: the whole model has to fit in memory. With one trillion parameters the weights alone take about 1 TB in FP8 and about 500 GB with 4-bit quantisation, plus the KV cache and the inference engine’s overhead.

It therefore takes a server with several data-centre GPUs and a few hundred GB of combined memory. A workstation such as those built on NVIDIA GB10, with 128 GB of unified memory, is sized for much smaller models. The specialised models Mistral will build on Large 4 may have different requirements.

What we think

An open-weight model at this scale, trained and served in Europe, widens the options for anyone who must keep data and inference under European jurisdiction, in an open-weight landscape led so far by Chinese and US labs. Two questions remain open: the licence (which determines the permitted commercial uses) and the gap to closed frontier models on the Artificial Analysis index.

What to watch

  • The release of the weights and the licence terms.
  • Independent evaluations on the published weights, beyond Mistral’s benchmarks.
  • The context window of the final version.
  • The specialised models derived from Large 4 and their requirements.

Sources

Need support?Under attack?Service Status
Need support?Under attack?Service Status