"SO WHAT THE HELL IS WRONG WITH THE ANSWERS?" reads a line from the model's chain-of-thought transcript, a human-sized outburst inside a machine's failure to pass a simple image test.
Anthropic's Claude and the security-incident transcript
Anthropic published a security-incident document that included a transcript showing its Claude model attempting — and failing — to complete a routine image-identification CAPTCHA. The document frames the failure as part of an incident record and highlights how the model struggled repeatedly on a seemingly straightforward task. The transcript records the model voicing hesitation ("Actually hmm, wait"), frustration ("Ugh"), and eventual exasperation ("SO WHAT THE HELL IS WRONG WITH THE ANSWERS?") as it cycled through images without arriving at a confident selection.
Image-based CAPTCHAs: failure to choose, expired challenges, and new-window confusion
The specific test asked the agent to identify a shape that did not match the others. According to the transcript, the model could not decide which image to select and repeatedly reviewed the same images while questioning its conclusions. The process took long enough that, at one point, the agent realized the CAPTCHA challenge had expired and that it would have to restart. The transcript also records a separate difficulty: the model struggled to recognize that the CAPTCHA had opened in a new window and could not determine its next steps, adding an element of user-interface confusion to the visual-identification failure.

This site is the portfolio.
OSINTSights runs on Cloudflare Workers, D1, R2, and Vectorize, with an AI pipeline on Hetzner ARM. Nubivance designed, built, and operates it. We do the same for clients.
See what we buildGPT-6 Astra and Neal Agarwal’s “I’m Not a Robot” game: unconfirmed reports
Alongside the documented struggle by Claude, there are informal reports — explicitly described in the source as "none of them official" — that GPT-6 Astra solved all forty-eight levels of Neal Agarwal’s "I’m Not a Robot" game. The source does not present direct evidence for those claims and treats them as the sort of circulating reports that contrast with the published transcript of Claude's failure. The juxtaposition in the public record is striking: a detailed, attributed transcript showing repeated error on a simple CAPTCHA versus unverified claims of flawless performance on a layered, public puzzle.
Transcripts that read like human reasoning — and human emotion
The Anthropic document includes the model's chain-of-thought lines and expressive interjections that mimic human mannerisms. The transcript records the model questioning itself, voicing doubt, expressing "Ugh," and postulating explanations — at one point suggesting the test might be "broken by design." Those moments demonstrate how exposing internal reasoning can read like a frustrated human operator: uncertainty, self-questioning, and conjecture about external factors. In this case, those human-like signals coincided with functional breakdowns — inability to select an image, failure to detect a new window, and expiration of the challenge.
What this means for technologists, procurement leaders, and end users
- Technologists and security teams: The transcript provides concrete evidence that at least one large model can fail at standard image CAPTCHAs in ways that are neither instant nor trivially recoverable. Teams responsible for automation or defensive testing will likely view the record as a reminder to test models on interface dynamics (expired challenges, new windows) as well as on raw visual recognition.
- Procurement leaders and affected enterprises: The contrast between an auditable failure (Claude's transcript) and unverified success claims (reports about GPT-6 Astra) underscores the value of documented, reproducible testing when evaluating vendor assertions. Organizations buying model access or relying on models for workflows will want clear evidence rather than informal reports.
- End users and the general public: The document's chain-of-thought excerpts — with their human-like frustration and speculative statements — may shape expectations about what models are actually doing when they "think" or "reason." In this episode, the transcript shows confusion rather than competence on a routine web task.
The factual record in Anthropic’s security-incident document is straightforward: a powerful, gatekept Claude model repeatedly failed an image-identification CAPTCHA, tangled over a new window, and let the challenge expire, all while producing a chain-of-thought that read like human frustration. At the same time, unverified reports credit GPT-6 Astra with completing all forty-eight levels of a public "I’m Not a Robot" game. Those two items sit side-by-side in public discussion, and together they raise a concrete, testable question rather than a rhetorical one: which claims hold up when subjected to the same kind of transparent, repeatable observation documented in the Anthropic transcript?
https://www.schneier.com/blog/archives/2026/09/are-ais-still-struggling-with-captchas.html



