Skip to main content
AI & Machine Learning

AI Developers Warn of Existential Risks, But Evidence Lags

Scientist stands thoughtfully in lab, surrounded by equipment, staring at blank laptop screen.

"pace the frontier," wrote Anthropic chief executive Dario Amodei in a 12 September essay — a short phrase that frames an argument about risk, evidence and what to do when the future looks frighteningly uncertain.

Hans Bethe, Emil Konopinski and the Manhattan Project standard

In 1942, scientists on the Manhattan Project confronted an extraordinary hypothesis: could an atomic explosion become hot enough to ignite a self-sustaining nuclear reaction in the atmosphere? Hans Bethe, Emil Konopinski and others did the physics. Their calculations showed the reactions could not release energy faster than the atmosphere could dissipate it, and three years later the Trinity test went ahead. The lesson Los Alamos left — investigate an existential hypothesis until the evidence can bear the weight of the decision — is the standard the author of the source argues should guide discussions about frontier AI.

Dario Amodei and Anthropic's warning about a potential botnet

Amodei raised a specific near-term concern: that within six to 12 months "a swarm of AI agents could take over much of the internet through a persistent botnet, a network of compromised computers acting under a single controller." The source says such possibilities warrant "serious investigation," particularly when they come from people with access to the most capable models — but it also insists the claim is extraordinary and therefore requires extraordinary evidence.

From demonstrated capability to predicted catastrophe: the evidentiary chain

The source describes a common pattern in AI-risk argumentation: begin with demonstrated capability, then extrapolate. A model shows offensive cyber capability; researchers imagine autonomous agents chaining those capabilities; scale that chaining across millions of systems and you get an AI-controlled botnet. The source cautions that while "each link may be technically plausible," establishing the outcome's probability demands far more work than establishing that the chain exists. It calls on national security analysts to separate threat modelling — exploring what could happen — from rigorous analysis that examines demonstrated capability, likelihood and consequence independently. The paper warns policy suffers when those categories collapse into one another.

Engineering controls Amodei lists — and the alternative to regulatory speed limits

The source notes Amodei himself identifies many engineering controls: monitoring, sandboxing, operational security, evaluation and interpretability. These are presented as familiar engineering problems that deserve concentrated effort; solving them, the source argues, "may prove more durable than any attempt to regulate the speed of progress itself." It also points out a simpler corporate option: if Anthropic believes its unreleased models are approaching an unsafe threshold, it can slow their development. OpenAI can do the same. Both companies "possess evidence the rest of us cannot see," and acting on that evidence would be a meaningful signal.

Strategic competition with China and the pacing problem

The source places the debate in a geopolitical frame: AI development is occurring amid strategic competition with China, and Amodei "recognises the tension." He argues democracies need a sufficiently large lead over China to give themselves room to slow down, while accepting a comprehensive agreement is unlikely soon because the incentive to cheat would be enormous. This makes "pacing" a difficult control problem: restrictions adopted in California can be inspected in California, whereas knowing whether equivalent restrictions are being observed in China is considerably harder. Here, the source says, the Manhattan Project analogy earns its keep — scientists confronted a catastrophic hypothesis while racing an adversary and responded by investigating the risk and engineering against the evidence.

What this means for Anthropic, OpenAI, and democracies

  • Anthropic: If the company truly sees unsafe thresholds in unreleased models, the source suggests it can and should slow development and expose the experiments, capabilities and assumptions that underlie its claims.
  • OpenAI: The source names OpenAI as another actor that "can do the same" — slowing development would be a private, visible action consistent with the evidentiary standard the piece advocates.
  • Democracies: Because comprehensive inspection across borders is hard, the source urges policymakers toward controls that remain useful as the frontier moves — hardening critical infrastructure, securing model weights, constraining what autonomous systems can access and execute, and setting measurable thresholds that trigger stronger safeguards as evidence accumulates.

The central argument is stark and specific: catastrophic framings skew the arithmetic — once human extinction appears on one side of a risk calculation, almost any cost on the other side becomes tolerable — and severity should not substitute for probability. If frontier models really are approaching the abilities Amodei describes, the response should be commensurate: publish the experiments, enumerate observed capabilities, make the assumptions linking observation to outcome explicit, and invest in the engineering controls already identified. The ability to imagine a catastrophe, the source concludes, cannot carry the same weight as evidence that one is coming.

Original story