Anthropic on Tuesday releases Claude Fable 5, a public, “Mythos-class” version of its previously limited Mythos model that the company held back over cybersecurity risks. Multiple outlets report that Fable 5 is positioned as Anthropic’s most capable public model, with strong results on tasks such as software engineering, complex analysis, and knowledge work. Anthropic says it restricts use in higher-risk areas—particularly cybersecurity and biology—by routing such prompts to Claude Opus 4.8, a less capable model with its own guardrails. Anthropic also states the fallback happens infrequently (around 0.05–5% of sessions, depending on the source and context) and that it notifies users or surfaces refusal reasons.

Within days of launch, developers and researchers report that the safeguards block or downgrade some prompts they consider legitimate, with complaints ranging from benign bio-related terms to non-sensitive requests such as resume editing and shopping lists. Anthropic acknowledges a tradeoff error that leads to more false positives, and says it is refining the classifiers and making downgrades more transparent for subscribers and API users. Separate reporting also notes that the company and others had tested the model against jailbreak attempts prior to launch, and at least one researcher claims to have found a way to bypass filters using multi-step prompting.