Mistral's Biggest Model Yet Leads Open-Weight AI on Cybersecurity and Coding

Mistral Large 4 arrives with a trillion parameters, a hard cybersecurity edge, and full weights promised before the month is out.

AI2Day NewsdeskAI-assistedPublished Editor: Lee Brown5 min read
Illustration for the story: Mistral's Biggest Model Yet Leads Open-Weight AI on Cybersecurity and Coding
Illustration made with AI. Not a photograph of the events described.
Share

Key points

  • Mistral Large 4, nicknamed "le Chonk," scored 82% on a benchmark that asks an AI to find and patch a real software flaw, the highest score of any model tested.
  • The model runs 1 trillion total parameters but activates only 49 billion at a time, keeping responses faster than the headline figure implies.
  • Claude Opus 5.5 and GPT-6 Astra score near zero on that same security test because their safety filters decline to perform it.
  • Mistral trained the model on 3,800 NVIDIA Grace Blackwell GPUs in its own European datacentres, covering more than 160 languages.
  • Full model weights will be released by the end of this month, letting organisations run the AI on their own servers.

Mistral Large 4 is already available through a preview API, and it is measurably ahead of every other open-weight model built outside China on the tasks security teams care about most. Here is what that actually means.

What is this model, and why does the size matter?

Mistral Large 4 (ML4) is a large language model, the same category of AI that powers chatbots like ChatGPT and Claude. At one trillion parameters (the numerical settings that govern how a model thinks), it is the biggest Mistral has ever shipped. A quirk of its design means only 49 billion of those parameters activate for any given query. That mixture-of-experts approach, routing each question to a relevant slice of the model rather than engaging everything simultaneously, keeps costs manageable without sacrificing speed.

The model handles text, code, images and documents in a single system, with no separate tools required.

Why is the cybersecurity score such a big deal?

ML4 leads every open-weight model built outside China on the Artificial Analysis Cyber Index, an independent ranking of AI performance on real security work. One sub-test asks a model to reproduce a real vulnerability in open-source software and then patch it. ML4 scores 82% there, the highest figure recorded for any model. It also completes 93% of challenges in Cybench, 40 exercises drawn from actual security competitions.

The contrast with the most capable closed models is striking. Several of them score near zero on that same sub-test, not because they lack underlying ability but because their safety filters decline to engage. Mistral's position is that defenders must prove a flaw is real before they can fix it, and a model that refuses that step is not much use to a security team. It is a reasonable argument, though it shifts responsibility for safe deployment onto whoever downloads the weights.

We first covered the AI-and-cybersecurity beat on 21 July 2026, and one pattern has recurred since: benchmark scores promise more than real-world conditions deliver. ML4's pre-release testing with vetted security partners and government agencies is a better signal than the numbers alone.

How does it compare on coding and general tasks?

Benchmark ML4 score Next best open-weight
DeepSWE v1.1 (software engineering) 61.7% DeepSeek V4 Pro 0813
AutomationBench (business workflows) 59.9% Kimi K3
Cybench (security challenges) 93% not disclosed
Vulnerability patch test 82% not disclosed
Human coding quality (1-5 scale) 3.74 Claude Opus 5 at 4.22

In a blind evaluation by Surge AI, professional coders rated ML4 second out of five models tested, behind Claude Opus 5 (4.22 versus 3.74). That gap is worth watching; it means ML4 is competitive but not yet the top pick for pure coding quality.

On AutomationBench, which covers 657 business workflows across tools including Gmail and Salesforce, ML4 scored 59.9%. Our October 2 story on AI agents and real office work found that even the strongest agents still fail most realistic tasks, so 59.9% is progress, not a finish line.

It is also worth noting what our 17 September story on frontier-model pricing established: the gap between expensive closed models and the best open-weight alternatives has been shrinking fast. ML4 is the most direct evidence yet of that trend closing on security-specific work.

What does this mean for people and organisations using AI today?

For most users, the immediate option is a preview API through Mistral Studio. Full weights arrive by the end of this month, after which any organisation can download and run the model on its own infrastructure, with no ongoing dependence on Mistral's servers.

That self-hosting option matters most for security teams, healthcare organisations and government bodies that cannot send sensitive data to a third-party cloud. European organisations get a further guarantee: Mistral operates the deployment end-to-end under European law, independently of other digital service providers.

The cybersecurity capabilities deserve care. Mistral is running pre-release testing with partners precisely because a model this capable of finding software flaws could be misused. Releasing the full weights anyway is a deliberate choice, and it puts real responsibility on the organisations that deploy it.

Common questions

Can I try Mistral Large 4 right now?

Yes. A preview API is live on Mistral Studio. The full model weights, needed to run it on your own servers, follow by the end of this month.

What does "open weights" actually mean for a regular user?

Mistral publishes the model's internal settings file, the way a recipe is published rather than kept secret. Any organisation can download it and run the AI on its own hardware, with no ongoing fees or reliance on Mistral's servers.

Is this model safe to use?

Mistral is running security tests with government and industry partners before releasing the full weights. The model is deliberately less restricted than closed alternatives on security tasks, which is useful for defenders but requires careful deployment policies on the user's side.

© 2026 AI2Day