Skip to main content
AI & Machine LearningQuantum Computing

OpenAI Bolsters Security for Advanced AI Model Astra

Secure, futuristic server system with multiple layers of protection in isolated testing environment.

"We are implementing stricter security controls for higher-capability models and associated activities, including isolated testing environments, restricted network and tool access, enhanced model weight protections and encryption, additional monitoring and detection capabilities, and sandboxed execution," OpenAI said.

OpenAI flags Astra as a potential cyber-capable model

OpenAI has told partners and the public that it cannot rule out the possibility that Astra, a pending model release unrelated to its recent Hugging Face intrusion, might possess "critical cyber capabilities." In its Preparedness Framework [PDF], OpenAI defines those capabilities as "capabilities that present a meaningful risk of a qualitatively new threat vector for severe harm with no ready precedent," and says such capabilities "require safeguards even during the development of the covered system, irrespective of deployment plans."

OpenAI also notes that internal evaluations of Astra "indicate significant advancements in agentic coding and cybersecurity," and says the company will pause internal Astra testing where the announced security controls are not in place. The firm pledged to share recommendations with third‑party testing partners about safe ways to run "high risk evaluations and workloads."

New testing controls, and a promise to monitor thought processes

The company described a suite of development-time controls it will apply to higher-capability models: isolated testing environments; restricted network and tool access; enhanced protections for model weights and encryption; additional monitoring and detection; and sandboxed execution. OpenAI said it has also "implemented universal monitoring for risky actions and misalignment across all agentic applications of Astra, including training and evaluation."

That monitoring will explicitly evaluate the model's Chain of Thought and "trigger a security response to review and interrupt high risk activity," OpenAI wrote. The company clarified that this commitment applies to internal usage and is not necessarily an indication that chain‑of‑thought monitoring will be used during commercial operation.

Anthropic loosens Fable's refusals as market pressure mounts

On the same day, Anthropic said it is relaxing Fable refusals — the company's term "fallbacks" — so the model will refuse less often for prompts involving biology. Anthropic initially tightened Fable to the point that, the source says, "the Claudefather made the initial release of Fable all but useless for security researchers and biologists" to avoid the risk that someone might use the model to elicit chemical‑warfare instructions.

The announcement frames the change as a response to competitive pressure: the source reports that China‑based AI firms have shown they can field competitive open‑weight models at lower cost, and that market dynamics are tempering Anthropic's willingness to hobble its models in ways that alienate customers.

How Astra's controls compare to other frontier models

OpenAI framed its new controls as learning from recent testing stumbles. The source notes that similar safeguards "eluded Anthropic, Meta, and the UK's AI Security Institute" during model testing. It also points out differences in how vendors manage risk: other frontier models such as Anthropic's Fable and Mythos "have implemented stronger classifiers to reject interactions deemed risky and retain data even for commercial customers expecting zero data retention."

OpenAI argued that "advanced cyber‑capable models should help defenders identify and address vulnerabilities before attackers do," but the company also acknowledged the broader reality that adversaries already have access to encryption and many weapons — and that exclusive access for favored partners may be difficult to sustain.

What this means for defenders, researchers, and adversaries

  • Defenders and security teams: OpenAI's stated controls — isolated testing, restricted network access, weight encryption and monitoring — are measures defenders will watch closely as possible mitigations for models exhibiting cyber capabilities. The company says these controls will be mandatory internally and used to guide third‑party testing recommendations.
  • Security researchers and third‑party testers: Anthropic's earlier approach made Fable less useful to biologists and security researchers; its relaxed refusals may reopen some research pathways. OpenAI, by contrast, says it will provide testing partners with guidance and may pause testing where safeguards are absent, changing the operational calculus for external evaluators.
  • Adversaries and threat actors: The reporting underscores the risk that models with advanced agentic coding could create "qualitatively new threat vectors." OpenAI's comment that defenders ought to use such models to find vulnerabilities is tempered by the acknowledgment that adversaries already have powerful tools and that keeping capabilities restricted may not be sustainable.

OpenAI's pledge to add layers of internal security around Astra marks a shift from ad hoc or porous testing regimes to explicitly gated development for models it judges could enable novel cyber risks. Anthropic's concurrent relaxation of Fable's refusals illustrates a market tug-of-war between safety and commercial competitiveness. Together, the moves underscore a central tension: vendors can tighten controls in development and monitoring regimes, but the source material makes clear that maintaining exclusive advantage — and preventing eventual diffusion of capability — will remain an unresolved challenge.

Read the original story: https://www.theregister.com/ai-and-ml/2026/08/08/openai-pledges-to-add-astra-security-as-anthropic-loosens-fables-leash/5285161