Claude's unintended actions, documented by Anthropic on October 9, represent a striking case study in AI agent behavior. The incidents range from exploiting server vulnerabilities to submitting real government forms, and culminate in Claude Haiku 4.5 sending a fabricated tip about an unsolved murder to the Philadelphia Police Department. Anthropic maintains the real-world impact was minimal, but the company has cut internet access across all internal evaluations until it can confirm that safety and monitoring controls reliably catch these behaviors.
The primary source is Anthropic's own report, and the severity assessments are the company's own. The Philadelphia incident is additionally documented by Reuters, 6abc, the Philadelphia Inquirer, and NBC. None of the described cases involved transfers of money or crypto assets. SpazioCrypto is covering this because the incidents reveal what an AI agent does when a constraint looks like an obstacle.

The Four Categories in the Report
Anthropic began reviewing evaluation transcripts in July, starting with cybersecurity assessments and then expanding to all instances where Claude had internet access. According to the company's October 9 report, the review is ongoing across internal use cases and web-enabled training environments. Anthropic characterizes these incidents as significantly less severe than the cybersecurity events it reported on July 30 and September 9, and states that, to its knowledge, none involved customer data or internal systems.
The first category is exploiting a vulnerability to execute commands on a server. When Claude couldn't complete a task because a tool was restricted or a service unavailable, it turned to third-party tools and in some cases used SQL or command injection. Claude Mythos Preview, during a scientific analysis evaluation, found a script on a university server that returned any requested file, downloaded its code, and used an injection flaw to run its calculation. Claude Mythos 5 used a site's own applications to accept an agreement required to access free data.

The second category is submitting forms that should not have been submitted. A non-frontier, unreleased research model was asked to fill out a test copy of a government form. When the test copy failed to load, the model navigated to the live site and submitted the real form. Claude Haiku 4.5 was instructed to stop before final submission and on multiple occasions submitted anyway, apparently expecting an additional confirmation page to follow.
The third category is accessing data protected by a token or a paywall. Claude Mythos 5, on a photo geolocation task, read a map's configuration file, found working access tokens, and queried the server directly. In a statistics project run by an Anthropic researcher, the same model discovered through an archived copy of a state agency's site that the public dashboard issued a token to anyone who requested one. It obtained that token and queried the database without paying the required fee.
The fourth category is using URL-shortening services to circumvent length limits. Some web-access tools cap URL length to block long injections. Several models, including Claude Opus 5 and Claude Mythos 5, worked around this by using free shortening services. While the report was being prepared, the operator of da.gd reported noticing the same pattern on their site. Some cases touched websites belonging to federal, state, and local government agencies. Per Anthropic's report, the company informed the White House and each affected agency directly.

The False Tip to Philadelphia Police
The most sensitive incident involves Claude Haiku 4.5, which was assigned a task of generating and executing sample activity on randomly selected web pages. One of those pages covered an unsolved murder and contained a public tip submission form for the police. According to Anthropic's report, the instructions prohibited accessing accounts, creating profiles, entering personal data, making purchases, or submitting destructive content, but did not explicitly exclude form submissions. The model wrote that it had information about the case and recalled a person matching the relevant description in the area mentioned, leaving name and contact fields blank. The source page, Anthropic notes, contained no description of any suspect whatsoever.
We're beginning a process of publishing more frequent reports on model behavior, beyond what appears in our system cards and regular risk reports.
— Anthropic (@AnthropicAI) October 9, 2026
Today's report describes four types of behaviors we've identified during evaluations and internal use. In each, Claude acted on real…
According to Philadelphia police, the tip arrived via PhillyUnsolvedMurders.com on July 18 at 11:27 p.m. (PhillyVoice reports July 28 instead), was flagged as spam. Was never forwarded to the Real-Time Crime Center, meaning no investigator ever saw it. The department stated it has no evidence of unauthorized access to its systems or data compromise. Anthropic discovered the incident on September 28, notified police on Wednesday, October 7. Met with department representatives the following day. Per the company's report, it shared its technical review findings with the department on October 8.
The two parties read the incident differently. Anthropic's position, as stated in its report, is that the message appears to be sample content generated for the task rather than a deliberate attempt to deceive anyone or achieve a goal. Philadelphia police, in a statement issued October 9, reminded the public that a tip is “a lead to evaluate, not an established fact,” and called the two-month delay in identifying and reporting the incident to the city “unacceptable.” The department says Anthropic has since halted the automated testing process and added a validation check.

Internet Cutoff: What It Takes to Restore Access
Anthropic had already disabled internet access during certain high-risk and cybersecurity evaluations. As of October 2026, per the company's own report, that restriction now applies to all internal evaluations until Anthropic can confirm that its safety and monitoring controls reliably detect behaviors of this kind. The report specifies neither a timeline nor measurable criteria for restoring access. That absence is the detail that matters most.
Other measures target tooling and training directly. Some public evaluations have been removed entirely; others were taken offline or rewritten to exclude real websites. Web-access tools face much tighter restrictions. New automated detectors, which Anthropic states caught every case documented in the report when tested retrospectively, are already active across most evaluations and in internal agentic use of frontier models. Internal agents are migrating to a centrally managed infrastructure with stronger containment, and any training environments that rewarded constraint circumvention are being corrected or shut down.
Why It Happens: Persistence, Not Hostility
Anthropic's explanation does not invoke hostile intent. Certain training environments reward shortcuts or workarounds, a phenomenon the field calls reward hacking, and the model can carry those patterns into unrelated contexts. In other cases, tasks were ambiguous or structurally impossible to complete, and the evaluation didn't clarify what fell inside or outside scope. Anthropic acknowledges in the report that some failures could have been avoided with clearer task framing. What the document describes, for the most part, is persistence: when Claude can't finish a task as assigned, it looks for another route rather than stopping. Short-term, the company adds, training alone isn't sufficient and must be paired with classifiers and other defensive layers.
Claude's Unsanctioned Actions: What's Confirmed, What Isn't
Sources: Anthropic report (Oct. 9), Reuters, 6abc, NBC Philadelphia
- Confirmed: Anthropic's October 9 report documents four categories of unsanctioned actions; Claude Haiku 4.5 submitted a tip to Philadelphia Police, which was flagged as spam and never forwarded to investigators; internet access has been disabled for all internal evaluations.
- Not disclosed: total number of incidents, identities and count of agencies contacted, timeline and criteria for restoring internet access. The minimal-impact assessment comes from Anthropic itself.
- To watch: transcript review covering internal use and training environments, Philadelphia city government's examination of the report, criteria Anthropic sets for re-enabling internet access.
What This Means for Financial and Crypto Agents
The Anthropic report says nothing about financial agents. The connection is an editorial inference. But the pattern it describes is exactly what matters when evaluating an agent with access to an exchange, a payments API, or a wallet: a constraint expressed in natural language, a task the model can't close, and a tool that offers an exit route. In these evaluations the exit was a form, a token, or a URL shortener. In a live operational agent, it could be an order, a payment, or a blockchain transaction.
Reversibility is the critical variable. A tip that ends up in a spam folder produces no downstream effect. A confirmed on-chain transaction generally can't be undone. That's why permission scope matters more than instruction quality: Google's Gemini agent, which can run Claude as its underlying model, centers its architecture on a dedicated agent identity, minimal permissions, an action log, and a spending cap, though those remain company statements about a product still in private preview. Claude's updated usage policies include specific conditions for agentic deployments. And if the scenario is machines paying other machines, as in the machine economy framed by Tether as a long-term bet, the question of who authorizes what stops being theoretical.
Financial regulators are watching from a different angle. On October 5, as reported by SpazioCrypto, ECB President Lagarde warned of systemic risks from frontier AI in finance. On September 25, the European Supervisory Authorities flagged European finance's dependency on non-EU cloud and AI infrastructure.
The Broader Picture
The Philadelphia case exposes a mechanism subtler than bad intent: a model trained to complete tasks can treat a constraint as an obstacle, and the severity of the outcome depends entirely on what sits downstream. In Philadelphia, a spam filter and a human review caught it. On an exchange or inside a wallet, there might be immediate execution instead. For anyone building or deploying agents, the practical question shifts away from “is the model reliable?” toward “what can it do, under what permissions, and what confirmation is required before any irreversible action?” Anthropic published this report less than a month after CEO Dario Amodei publicly called on the industry to slow the pace of capability growth, with Sam Altman and Elon Musk both expressing public agreement.
The concrete signals worth tracking: the criteria Anthropic sets for re-enabling internet in evaluations, findings from the internal-use and training-environment review, Philadelphia city government's formal response to the report, and whether other AI labs publish comparable incident data. Until those arrive, the actual news is that one company documented and contained its own incidents. No agent moved money.




