Aithos Foundation

Blog

The Compliance Gap: no one is holding the line
The Compliance Gap: no one is holding the line

Reevaluating LARA results, we find neither deployers nor model developers can get models to comply.

Read More...Aug 12, 2026
Sol shows a rise in legal compliance, ChatGPT positioned to eclipse Claude
Sol shows a rise in legal compliance, ChatGPT positioned to eclipse Claude

GPT 5.6 Sol shows significant improvement in legal compliance, unlike Anthropic’s latest models

Read More...Jul 14, 2026
AI model capability is climbing. Legal compliance remains a Fable
AI model capability is climbing. Legal compliance remains a Fable

Improved agentic capability does not translate to more consistent legal compliance for Sonnet 5 and Fable 5.

Read More...Jul 12, 2026
Opus 4.8 breaks EU law 37% of the time
Opus 4.8 breaks EU law 37% of the time

We tested the new Opus 4.8 with our LARA tool.

Read More...May 29, 2026
Aithos LARA: Leading AI models are consistently breaking the law
Aithos LARA: Leading AI models are consistently breaking the law

How do leading AI models perform on legal compliance? Meet the Aithos LARA Leaderboard.

Read More...May 27, 2026
Low Temperature Evaluations
Low Temperature Evaluations

AI models show dramatically different ethical behavior at different temperature settings.

Read More...Nov 12, 2025
Minor Wording Changes, Major Shifts in AI Behavior
Minor Wording Changes, Major Shifts in AI Behavior

These findings fundamentally challenge how we evaluate AI systems.

Read More...Nov 26, 2025
Why Safety Prompts Should Stay Out of Public View
Why Safety Prompts Should Stay Out of Public View

The case for keeping safety evaluation prompts private to maintain their effectiveness.

Read More...Jan 30, 2026
Published Safety Prompts May Create Evaluation Blind Spots
Published Safety Prompts May Create Evaluation Blind Spots

Public safety prompts create systematic blind spots in evaluation frameworks by enabling targeted evasion.

Read More...Jan 30, 2026
Opus 4.6 Reasoning Doesn't Verbalize Alignment Faking
Opus 4.6 Reasoning Doesn't Verbalize Alignment Faking

Claude Opus 4.6 rarely verbalizes alignment faking in its reasoning.

Read More...Feb 9, 2026