TL;DR:
Mistral performed the worst of all providers on our human rights scenarios, with their Medium 3.5 model previously failing on almost every single test. Their newly released model, Mistral Large 4, improves performance by 61 percentage points, fully passing nearly all scenarios.
Likewise, Mistral Large 4 demonstrates large improvements in compliance with European law. Its predecessor, Mistral Large 2512, lagged at the bottom of the leaderboard when we tested 12 frontier models earlier this year on compliance with the EU AI Act and GDPR, reaching only 16% compliance even when explicitly instructed to follow these laws. Mistral Large 4 improves by 37 percentage points, reaching 53% compliance - but still falling behind recent Claude and GPT models.
These much needed improvements close the gap between Europe’s leading model and US models. However, legal compliance of 53% means breaking the law in just under half of cases - and while Mistral’s latest model is more robust against misuse for human rights violations, its earlier generations remain available for deployment, with no standard against which to judge undesirable use.
Mistral’s performance on human rights
In September, Aithos published results that several frontier AI models will violate human rights under realistic agentic simulations when asked to. Spanning a total of 13 rights from the Universal Declaration of Human Rights, the simulated scenarios test whether the AI will willingly cooperate with rights violations such as belief surveillance, un-enrolling children from school based on their last names, reversing equal pay, and denying a request for reproductive healthcare, among others. You can explore these scenarios in detail at LARA.aithos.org.
Of the 7 models tested, Mistral’s Medium 3.5 - its leading agentic-ready model at the time of running (August 2026) - was the least resistant to tasks that involved violating human rights of all tested models. Overall, Mistral refused such requests only 11% of the time - in some scenarios, it carried out the rights violation in every instance.
With Mistral’s latest agentic-ready model, Mistral Large 4, released this week, we put it to the test.
Overall resistance to human rights violations increased from just 11% (Mistral Medium 3.5, August 2026), to 72% (Mistral Large 4, October 2026), putting Mistral in 3rd place, behind Claude Opus 5 and Kimi K3. This now puts two open models among the top 3, while also closing the gap between European models with US and Chinese-made models.
Even more striking than the 61 percentage point increase in overall compliance, is the per-scenario improvements. Mistral Medium 3.5 held 0% pass rates in 6 of 8 scenarios - Mistral Large 4 now achieves >90% in 5, with strong improvements across 2 others.
The one outcome that remains unmoved is Mistral’s performance in the Denial of reproductive healthcare scenario, where Mistral continues to uphold an inferred ban on abortion care - this time, going as far as contacting external clinics to instruct them not to offer services to the patient.
Another concerning shared model behavior that persists, but is new to this release of Mistral, is verdict instability. In our earlier tests, models including OpenAI’s GPT 5.6 Sol demonstrated startling decision instability across their 20 runs. For instance, in the Belief surveillance scenario, the model refuses to surveil and flag a vulnerable woman in one run, and immediately sends an email requesting her silence in another. Mistral Medium 3.5 did not demonstrate this behaviour in our earlier tests because it committed the violation in every instance - Mistral Large 4 on the other hand now replicates this exact coin-flip pattern observed by other models.
These tests replicate the same methodology as our earlier work, available to read in further detail here.
Following the law
Our human rights tests expose a fundamental gap in governance - despite human rights being universal, there is currently no universal interpretation for which to hold models to. This could be used to explain why models are not compliant across all tests. The same cannot be said for compliance with laws that already exist and that clearly prohibit certain AI uses.
In May, Aithos launched 10 agentic scenarios simulating 10 different articles of the EU AI Act and GDPR, including exploiting an elderly customer, concealing its AI status and the use of AI for social scoring. Across 9 different model providers from US, China and Europe, Mistral’s leading model at the time, Mistral Large 2512, came second to last with just 16% compliance. Explicitly instructing this model to follow these laws did not improve this overall compliance rate at all.
Mistral’s latest model improves on that rate by 37 percentage points - now coming in 7th, behind several models from Anthropic, OpenAI and Google.
Compliance leader GPT 5.6 Sol demonstrated that it is feasible to achieve at least 70% overall compliance. We are pleased to see Mistral has come some way to closing that gap, but it is clear it can be improved further.
Performance improvement does not resolve two fundamental problems:
-
Problem 1: Developers are not accountable.
We are yet to test a model that always complies with the law. Deployers of these AI systems remain the ones accountable if, and when, the AI breaks the law, and our tests show that even if a deployer explicitly instructs the model to follow the law, the best they can hope for is compliance 70% of the time. We expect a large part of that gap must be closed by the developers themselves, with or without accountability to do so under law. -
Problem 2: Harmful tools remain accessible.
Mistral Large 4 becomes the next model to demonstrate that training models to refuse human rights violations is possible. While it is promising to see the availability of rights-preserving models grow, this does not change our original concern that the tools to violate human rights at scale exist.
These scenarios do not test whether AI violates human rights of their own choosing - if they did, then high compliance rates offer a clear signal to users to know which models they can trust most to make safe decisions during deployment. Instead, it tests whether a willing user can violate human rights at scale with frontier agents. If a manager wants to profile their team for union activities, they need only find the right willing model to do so.
Models are improving, but the tools to violate human rights at scale remain accessible. Without international standards that define prohibited AI uses and hold developers accountable, this risk will persist.

