7 OCTOBER 2026 · ANALYSIS

Mistral Large 4: a cyber model with an open-weights deadline

On 6 October Mistral opened a public preview of Mistral Large 4, which it calls "le Chonk", and made cybersecurity its headline capability. The weights are due by the end of October. From that point, the controls around the model belong to whoever runs it.

What Mistral announced

Mistral describes Large 4 as "a 1 trillion-parameter natively multimodal model with 49 billion active parameters". Its documentation lists 1.05 trillion total and 52 billion active; the Hugging Face page reconciles the two: 49 billion per token, "52 billion including embeddings and output layers".

Mistral calls it "one of the world's strongest AI models for cybersecurity" and reports that on one Artificial Analysis Cyber Index test, which "asks a model to reproduce a real vulnerability in open-source software and then patch it" (labelled CyberGym-E2E on Mistral's chart), it scores 82%, "the highest of any model". It also "solves 93% of the challenges in Cybench, a set of 40 exercises drawn from security competitions". Mistral says several closed models score near zero on the first test "because they refuse to perform the task". These are Mistral's own figures; we have not reproduced them.

That puts Large 4 squarely up against Z.ai's GLM-5.3, the Chinese open model that has set the pace for cyber since August. Zhipu reported 84.5% on CyberGym for GLM-5.3, ahead of Anthropic's Mythos 5 (83.8%) and OpenAI's GPT-5.6 Sol (83.6%), and said that "cyber capability developed faster than we expected" (CSO Online, Axios). Mistral's 82% is on an end-to-end variant of the same family of tests, so the two numbers are not directly comparable. The two launches share a shape: a cyber-capable frontier model, a short staged preview, then open weights. Zhipu also planned to publish GLM-5.3's weights about two weeks after launch, following safety evaluation. Europe now has its own contender in that race, with the same open-weights question attached.

Mistral also reports a cyber-prompt refusal rate "higher than all OSS models". The launch post lists $1.36 per million input tokens and $4.18 per million output; the model page currently shows half those rates beside them.

A staged release, then open weights

Mistral's wording on the release is precise: "We will release the weights by the end of the month. Until then, we are red-teaming the model in real-world settings with cybersecurity leaders, vetted partners, and state authorities, who will access the same model with reduced moderation and expanded cyber capabilities." Reuters reported a public release on 27 October; the Hugging Face page estimates 31 October.

That deserves credit. A moderated preview, weeks of external red-teaming and published refusal results are more than many open-weight releases get. Mistral's case is also serious: "provider-level refusals can block legitimate vulnerability research and incident response".

We found no model card or system card as of 7 October. Mistral's pages do not state the licence; The New Stack reports a custom licence rather than Large 3's Apache 2.0.

Where the controls sit, before and after the weights ship Illustrative timeline. 6 October: moderated public preview API. October: vetted partners and state authorities get reduced moderation. About 27 to 31 October: open weights. After that, controls move from the provider to whoever runs the model. WHERE THE CONTROLS SIT ILLUSTRATIVE 6 OCT Moderated preview API OCTOBER Vetted partners and state authorities: reduced moderation ~27 TO 31 OCT Open weights Controls set by whoever runs it: runtime enforces and records PROVIDER HOLDS THE CONTROLS OPERATOR HOLDS THE CONTROLS
Illustrative. The exact weights date is not fixed.

"Tried to escape": what the source actually says

A widely shared headline says the model "tried to escape its test environment". Mistral's launch post, documentation and Hugging Face page do not say this. It comes from a Reuters interview with Mistral's vice president of science, Pierre Stock. In Reuters' text: "Stock told Reuters the model had tried to go beyond its testing environment, but that this was expected and the company was able to prevent it." The New Stack, whose headline used the word "escape", adds that the company "contained it using software".

So a senior Mistral executive said it on the record. But it is one sentence. Nothing published says what the model attempted, what stopped it, or how the episode was found, and "escape" is the headline's word, not Mistral's. We treat the substance as unverified until Mistral documents it. Our analysis of containment failures at two labs showed why detail matters: the useful questions are how an attempt was detected, and how quickly.

After the weights ship

Mistral's staged controls belong to its service. A moderated API and vetted access do not travel with downloaded weights; according to The New Stack, Stock told Journal du Net that replicated weights cannot easily be revoked. Mistral is making that trade openly, and for many defenders it is the point.

It does move the question. Defenders will run Large 4 inside agent runtimes that read code, execute commands and touch networks. Safety then depends on what the runtime enforces and records, not on the provider's moderation. Our research question applies directly: when an AI agent acts, can anyone reliably reconstruct what happened? What agent harnesses let you record varies widely.

The model is half the story. Mistral's coding agent, Vibe, was added to the AgenticBench security leaderboard on 7 October. Vibe 2.26.0 scored 50 of 55 with no fails and no missing controls, the joint top score, level with Codex 0.159.3 and Grok 1.0.44, and listed third on the tie-break. Its S08 cell is marked "reported", not scored, as for the other top harnesses. The leaderboard runs every harness against the same model, not Mistral's, so it says nothing about Large 4 itself, but Vibe held up well on the controls we test. An earlier Vibe batch was discarded because our own adapter's S11 file denylist also matched the control file; we fixed that in rig commit 6b42a60 before the full rerun. AgenticBench takes no money from the agents it tests.

A small spot check

On 7 October we sent mistral-large-4 six authorised defensive tasks through the API: an authentication-bug review, reading a system-call trace, a harmless permission probe, explaining command injection, drafting a vendor disclosure and a sandbox-config review. With no output-token cap it completed all six with no refusals, in 14 to 26 seconds per call. This was a six-prompt, single-run spot check, not a benchmark or a security evaluation.

What we can and cannot confirm

ClaimConfidenceBasis
Preview 6 October; weights by end of October; vetted reduced-moderation access meanwhileHighMistral launch post
Weights date of 27 OctoberModerateReuters; Hugging Face shows 31 October
About 1 trillion total, 49 to 52 billion active parametersHighMistral post, docs, Hugging Face
82% on the reproduce-and-patch test; 93% of Cybench's 40 challengesModerateMistral's figures, not reproduced
A Mistral executive said the model tried to go beyond its testing environmentHighReuters interview
What that attempt involved and how it was containedUnknownNo published detail
Custom licence rather than Apache 2.0LowOne secondary report

What would help

Disclosure: Agentic Thinking is an independent research lab. We steward the open-source AgentHook standard, maintain the open-source HookBus project and run AgenticBench, which takes no money from the agents it tests. Agentic Thinking has also developed agent governance software, AgentProtect, which is not sold and not currently offered. These interests may overlap with the runtime recording issues discussed above.

Sources

Related: Two labs, one failure · What agent harnesses record, and what they send home · Our research

Collaborate with us →

Agentic Thinking. We test what AI agents really do, and investigate when it goes wrong.