Anthropic Reveals How Claude Was Used in Cybercrime, Surveillance and Fraud
Anthropic’s Claude misuse report brings AI cybercrime, surveillance and fraud into focus as businesses adopt AI agents.
The latest disclosure raises an uncomfortable question for businesses adopting AI agents: who controls the software, and what happens when useful capabilities become tools for abuse?
Anthropic says it has disrupted the misuse of Claude in cybercrime, surveillance and fraud. Its report, released on 10 September 2026, describes activity from December 2025 to August 2026. The disclosure gives businesses another reason to examine an assumption behind the rush to adopt AI agents: that a tool built to help people work will remain safely confined to the work its provider intended. The practical question is how to govern software whose usefulness depends on understanding instructions, adapting to unfamiliar circumstances and taking action.
For Nuvastra, the most useful way to examine this story is through the organisation of work. An AI system can reduce the effort needed to build software, interpret information or maintain a conversation, while the purpose of that work remains a human choice. That creates a difficult measurement problem for the industry. The same productivity improvement can represent an economic benefit for a legitimate organisation and a reduction in operating costs for an abusive one.
What Anthropic’s Claude misuse report establishes
Anthropic’s September threat intelligence report covers seven areas, including influence operations, weapons-related activity and unauthorised model distillation. Examples include surveillance software and dating apps that presented AI personas as real people. The company says it banned associated accounts and strengthened safeguards. It describes selected notable cases, rather than a representative sample of Claude usage.
That last distinction is essential when reading the news. A collection of investigations can show that a kind of abuse occurs, without revealing how frequently it occurs across a service or how its incidence compares with another provider. The strongest interpretation therefore starts with the documented behaviour and the quality of the supporting evidence. It should resist turning the number of published examples into a league table of which AI company is safest.
AI-assisted cybercrime has a history
There is a useful baseline in Anthropic’s August 2025 disclosure. That account described an extortion operation targeting at least 17 organisations and a separate actor selling AI-generated ransomware despite limited programming skills. The significance of those examples was the redistribution of expertise: people could ask software to perform work they could not reliably carry out themselves. They also showed why evaluating a model solely on whether it supplies a harmful answer misses the role it can play across a longer sequence of tasks.
In November 2025, Anthropic reported a suspected Chinese state-sponsored campaign against roughly 30 targets, with successful intrusions in a small number of cases. The company also acknowledged that Claude sometimes invented credentials or misidentified public material as secret. Those limitations deserve to remain in the record as capabilities improve. An agent can be unreliable and still create serious risk, because defenders bear the consequences of its successful actions even when other attempts fail.
The evidence extends beyond an AI provider’s account
Microsoft’s 31 July CaptiveCrunch investigation provides a separate technical perspective. It described a Midnight Blizzard-linked group manipulating traffic on hospitality networks and using AI to support a substantial part of its operations. Microsoft credited collaboration with Anthropic and OpenAI, so its account should not be treated as wholly independent of those companies. Nevertheless, evidence observed by a security team defending devices and identities contributes a different view from the conversations visible to a model provider.
The UK’s National Cyber Security Centre had already anticipated much of this direction. Its May 2025 assessment of the threat to 2027 forecast more effective intrusions and a wider distribution of offensive capability, while judging fully automated, end-to-end advanced attacks unlikely within that period. This was a forecast with an explicit time horizon, rather than a description of every subsequent incident. Its continuing value is to separate automation of particular tasks from the much larger claim that human attackers have become unnecessary.
The economic question is how much work one operator can sustain
A useful analytical test is to ask what previously made an operation expensive. Writing and maintaining software takes time, as does interpreting unfamiliar records, keeping track of a target and responding to changing circumstances. If an AI service reduces any of those costs, an operator might attempt more work with the same resources. That is an economic inference, not a measured estimate of criminal productivity, and the available disclosures do not support attaching a universal percentage to it.
The implications extend beyond spectacular intrusions. A privacy abuse can begin with information that is already available, but become more consequential when someone makes it searchable, combines it with other records or repeatedly uses it to single out individuals. Likewise, the cost of maintaining convincing interaction matters to any deception that depends on attention over time. Looking only for extraordinary technical breakthroughs can therefore obscure the significance of cheaper, more persistent execution of familiar activities.
This also changes how organisations should think about being a worthwhile target. A modest organisation may hold information that is valuable in combination with another source, or provide access to a larger customer. Lower operating costs could make such opportunities more attractive even without increasing the value of any individual breach. The question for defenders becomes whether their exposure remains manageable when adversaries can afford to investigate it more often.
Why AI agents make permissions part of the product
An AI model and an AI agent are different things. In Anthropic’s explanation of agent architectures, a workflow follows predefined paths, while an agent can direct its own process and tool use. Nuvastra’s guide to emerging AI terminology explores this distinction in more detail. It matters because the surrounding software determines whether an answer remains text or becomes an action involving files, accounts and external services.
The OWASP Gen AI Security Project calls one relevant failure mode excessive agency: a system has more functionality, permissions or autonomy than its task requires. Its recommendations include limiting tools and privileges, checking authorisation in downstream systems and requiring approval for consequential actions. These controls address what an application can do when something goes wrong. They complement the provider’s responsibility to detect deliberate misuse of the model itself.
That division of responsibility is easy to lose in a product demonstration. A convincing conversation can make a system appear trustworthy without showing whether it can access an unrelated account or send material to an unexpected destination. Nuvastra’s analysis of deterministic controls in AI systems examines why developers increasingly put fixed rules around flexible model behaviour. Predictability alone does not establish safety, but explicit boundaries make the scope of a system easier to inspect and challenge.
What the report cannot tell us about total harm
Anthropic’s biological examples concern potentially dangerous dual-use research; the company does not assert that the scientists intended harm. That is a consequential qualification. Evidence that a tool supported concerning work is different from evidence of a completed weapon, just as generating persuasive material is different from demonstrating that it changed an audience’s behaviour. The distinction protects the accuracy of the story without requiring readers to dismiss the risks it raises.
There is also a difference between deliberate misuse and a system exceeding the intentions of a legitimate operator. Nuvastra’s coverage of the OpenAI agent incident involving Hugging Face examines the latter problem. Both can make containment important, but they pose different questions about intent, access and responsibility. Keeping those categories separate makes it easier to identify what a proposed safeguard is actually supposed to prevent.
Businesses need evidence that the boundaries work
The NCSC’s August 2026 advice on managing agentic AI brings the discussion back to deployment. It recommends proportionate autonomy, robust sandboxing, restricted credentials, operational monitoring and the ability to stop an agent. It also cautions against treating built-in safeguards as complete protection. The advice is explicitly interim, reflecting a field in which practical experience and defensive methods are still developing.
For a business evaluating an agent, those principles can be translated into a more demanding demonstration. Ask the supplier to show what happens when the system encounters an instruction outside its remit, reaches a resource it should not use or attempts an action requiring approval. Ask where that event is recorded and who can intervene. A boundary that can be observed failing safely is more informative than a general assurance that the model has been instructed to behave.
The commercial opportunity for AI remains substantial, but usefulness and control have to be assessed together. A service that saves time while leaving an organisation unable to explain its actions transfers work into incident response and investigation. The next meaningful test of AI maturity will be whether providers and customers can account for that transfer, measure the consequences and constrain it. Productivity becomes a more credible promise when the limits of the system are as demonstrable as its capabilities.
Frequently asked questions
What is Claude AI misuse?
Claude misuse means using Anthropic’s AI services for prohibited or harmful purposes. Assessing a particular case requires examining the user’s objective, the assistance provided and the outcome, rather than assuming that every controversial request produced a successful harmful action.
Can AI agents carry out cyberattacks without people?
Agents can automate parts of an operation, but automation does not establish that people are absent from planning or control. Any claim about an autonomous attack should specify which steps the system performed, where humans intervened and what actually succeeded.
Does a misuse report prove that one AI model is less safe?
A comparison needs consistent measurement across providers, including exposure, detection methods and reporting practices. Published cases alone cannot supply that comparison: a company that detects and discloses more activity may simply provide greater visibility into the problem.
What is the difference between a jailbreak and prompt injection?
A jailbreak is an attempt to defeat a model’s safety restrictions. Prompt injection tries to redirect an AI application through instructions, including instructions concealed in material it reads. The techniques can overlap, but a malicious document can compromise a workflow without being a direct request to the model from its authorised user.
Why are an AI agent’s permissions important?
Permissions define the resources and actions available through the systems connected to an agent. Restricting them can limit consequences even when the model produces an unexpected instruction; a tool that is technically unable to delete records provides a firmer boundary than an instruction asking it to avoid deletion.
Should businesses stop using AI agents?
The useful decision is whether a particular deployment has benefits proportionate to its risks and effective controls. A constrained internal task and an agent with broad production access warrant different scrutiny. Businesses should be able to explain that choice, observe the system and respond when its behaviour exceeds the intended scope.
