AI Safety Team Halts Multiple Attempts to Weaponise Large Language Models
Anthropic's first public transparency report reveals how researchers attempted to use Claude for gain-of-function virology and toxin design - and why detection remains harder than the industry admits.

When Research Crosses the Line
Anthropic terminated access for multiple researchers in 2026 after its automated safety systems flagged attempts to use Claude for what the company describes as high-risk biological work, including drafting grant proposals for gain-of-function studies on viruses with pandemic potential. The disclosures, published 10 September in the company's first transparency report on misuse, offer a rare window into how frontier AI labs police dual-use capabilities - and how murky that policing can be.
The company documented five separate incidents involving biological research, alongside cases in which users attempted to harness Claude for surveillance software, exploit development, propaganda generation, and weapons-system design. Every account implicated was banned. Yet Anthropic declined to name the individuals or their institutions, citing both the ambiguity of intent and the risk of retaliation against working scientists.
At Opentechwire, we have tracked the proliferation of reasoning-capable models across Asia and the West over the past eighteen months, and the policy question has shifted from whether large language models can accelerate sensitive research to how labs distinguish between legitimate and dangerous use at scale. Anthropic's report suggests the answer is still uncomfortably subjective.
The Chikungunya Grant That Set Off Alarms
In May, Claude's biological safety classifier - a separate screening layer trained to identify prompts related to pathogens, toxins, or dual-use techniques - flagged a request to draft a research grant. The proposed study centred on chikungunya virus, a mosquito-borne pathogen for which no licensed treatment exists and which can cause joint pain lasting months. The grant outlined methods to enhance the virus's transmissibility and its ability to evade immune defences, hallmarks of gain-of-function research.
Gain-of-function work modifies organisms to study traits that do not exist in nature, often to anticipate evolutionary paths or test countermeasures. The technique has been central to virology for decades, but it also sits at the heart of biosecurity debates: the same methods that help predict a pandemic can, in theory, trigger one.
According to Anthropic, two factors elevated this particular prompt from routine to high-risk. First, the research design combined two enhancement goals - transmission and immune evasion - in a single pathogen. Second, metadata indicated the user's affiliation with a military research institute. The company did not disclose which country or institution, but the combination was enough to trigger an immediate review and account suspension.
Jacob Klein, Anthropic's head of threat intelligence, told journalists that the company rarely encounters users who announce malicious intent outright. "You are not seeing someone in a comic book kind of way say, 'Hey, I want to build a biological weapon to kill everybody,'" Klein said. "It is an incredibly nuanced situation."
Bird Flu, Venom Peptides, and the Grey Zone
Two other biological cases followed similar patterns. In one, a researcher used Claude to design experiments that would increase the pathogenicity of avian influenza, a virus already under intense scrutiny because of its sporadic human transmission and high fatality rate among those infected. In another, a user asked the model to construct an atlas of venom toxin peptides and to build a generative pipeline optimising toxin characteristics for undisclosed purposes.
Anthropic classified all three as violations of its acceptable use policy, which prohibits "development or use of weapons, including… biological agents." Yet the company acknowledged the difficulty of distinguishing between a scientist drafting a legitimate vaccine-development proposal and someone pursuing weaponisation. Both might query the same viral sequences, request the same molecular simulations, or explore the same gain-of-function pathways.
The company's solution has been to err on the side of caution. When the biological safety classifier flags a prompt, human reviewers assess the user's history, affiliations, and the specificity of the request. If the context suggests dual-use risk - particularly military ties or requests that combine multiple enhancement techniques - the account is suspended and the case logged for model retraining.
This approach raises questions about false positives. Anthropic's report does not quantify how many legitimate researchers have been locked out, nor does it describe an appeals process. The company stated only that it prioritised "possible consequences" over user convenience, a stance that reflects the asymmetry of risk in biological misuse: a single engineered pathogen, if released, could dwarf the harms of over-enforcement.
Surveillance, Exploits, and Weapons Systems
Biological misuse formed the most sensitive portion of the report, but it was not the only category. Anthropic documented attempts to use Claude for building surveillance tools, generating software exploits, producing disinformation at scale, and designing unspecified weapons systems. The company provided less detail on these cases, noting that some involved state-affiliated actors and that disclosure could compromise ongoing investigations or reveal detection methods.
The surveillance cases are particularly relevant to the policy environment in Asia, where export controls on AI inference chips and model weights have tightened over the past year. Regulators in Tokyo, Seoul, and Singapore have begun treating high-capability models as dual-use technologies subject to the same licensing regimes that govern semiconductor manufacturing equipment and precision machine tools. Anthropic's report will likely inform those frameworks, especially as governments weigh whether to require transparency reports as a condition of market access.
In the exploit-development cases, users attempted to generate code for known vulnerabilities or to automate reconnaissance for zero-day discovery. Anthropic's acceptable use policy explicitly prohibits "malware, ransomware, or tools designed to compromise systems," but enforcement depends on the model's ability to recognise malicious intent embedded in ostensibly neutral requests. The company did not disclose whether the flagged exploits targeted specific organisations or were part of broader security research.
Detection Is Harder Than the Brochure Suggests
One of the more candid admissions in Anthropic's report is that its safety classifiers operate in a fog of ambiguity. The biological safety layer, for instance, scans for keywords, structural patterns, and domain-specific jargon associated with pathogen research. But a virologist drafting a grant to develop a chikungunya vaccine might use identical language to someone designing a bioweapon. The model cannot read intent; it can only flag patterns and defer to human judgment.
This limitation is not unique to Anthropic. Every frontier lab - OpenAI, Google DeepMind, Alibaba Cloud, and the growing cohort of open-weight projects - faces the same challenge: large language models are general-purpose tools, and general-purpose tools are, by definition, dual-use. A model trained to assist with protein folding will also assist with toxin design. A model that can debug benign code can debug malware.
The industry's preferred solution has been a layered defence


