Up Close
OpenAI puts a frontier model under stronger controls as Anthropic loosens a narrow safety bottleneck
OpenAI said internal evaluations of its upcoming Astra model indicate potential critical-level cyber capabilities under its Preparedness Framework. The company said it is scaling up robustness testing, imposing stricter security controls for high-capability models, pausing Astra work that does not yet meet those controls, and introducing universal monitoring for risky actions across agentic uses of Astra.
This appears to move frontier-model safety from a policy layer into product operations: OpenAI is not only describing a threshold, it is changing how one model may be developed and used. That is likely to matter beyond OpenAI, because a published pause tied to internal capability testing gives other suppliers a clearer precedent for slowing deployment when an evaluation crosses their own line.
Anthropic said updates to Claude Fable 5’s biology safeguards reduce biology-related fallbacks to less capable models by about 85% across its products. It said the change expands the model’s ability to answer health and educational biology questions and some clinical tasks, while continuing to fall back to Claude Opus 5 on dual-use topics such as virology, toxicology and molecular design.
This suggests the current competitive work is not simply to add more refusals, but to narrow them so safer handling covers more ordinary use without opening the most sensitive areas. Taken together, OpenAI’s tighter cyber controls and Anthropic’s narrower biology fallbacks are likely to show the same shift from broad blocking towards finer-grained governance inside the model layer itself.
Evidence
- 2026-08-07 OpenAI reported that internal evaluations of its upcoming Astra model indicate potential critical-level cyber capabilities under its Preparedness Framework and announced scaled-up robustness testing, stricter security controls for high-capability models, a pause on Astra activities that do not yet meet those controls, and universal monitoring for risky actions across agentic uses of Astra. openai.com
- 2026-08-07 Anthropic announced updates to Claude Fable 5’s biology safeguards that reduce biology-related fallbacks to less capable models by about 85% across its products, expanding the model’s ability to handle health and educational biology questions and some clinical tasks while continuing to fall back to Claude Opus 5 on dual-use topics such as virology, toxicology and molecular design. anthropic.com