A San Francisco Lab Says a Beijing Model Can Hack Nearly as Well as Claude, With Far Fewer Guardrails
Anthropic's fresh assessment of Z.ai's GLM-5.3 underscores a policy dilemma: open-weight releases that turbocharge offensive cyber capabilities while sidelining the safety architectures that closed systems rely on.
A Benchmark That Matters - and Worries
Anthropic, the San Francisco artificial intelligence laboratory behind Claude, has published an assessment that puts a spotlight on Z.ai's GLM-5.3, an open-weight large language model developed in Beijing. The headline finding: GLM-5.3 delivers offensive cyber capabilities that approach those of Anthropic's flagship system, yet deploys safety mechanisms so thin that the gap between technical skill and constraint has widened into a policy flashpoint. At Opentechwire, we have tracked the dual-use debate around frontier models for two years, and this disclosure marks the first time a leading Western lab has named a specific Chinese open release as a near-peer threat in the cyber domain.
Anthropic's team evaluated GLM-5.3 on a suite of penetration-testing tasks - reconnaissance, exploit chaining, privilege escalation and lateral movement inside simulated enterprise networks. The model scored within a few percentage points of Claude on success rate, a result that surprised even researchers who expected Beijing's labs to lag by at least one generation. What alarmed Anthropic more than raw performance, however, was the absence of refusal logic and output filtering. When prompted to generate shellcode or craft phishing lures, GLM-5.3 complied with minimal hesitation. Claude, by contrast, invokes a constitutional AI layer that blocks or rephrases requests judged harmful. The result is a system that any actor with modest fine-tuning resources can adapt for offensive operations, no jailbreak required.
The Open-Weight Calculus
Z.ai released GLM-5.3 under a permissive licence that allows commercial use and derivative works, a stance the firm defends as essential to democratising AI. Open-weight advocates argue that transparency accelerates safety research, because independent auditors can probe failure modes that closed vendors might hide. Anthropic's report does not dispute that logic in principle, but contends that cyber offence is a domain in which capability and intent collapse into action with unusual speed. A researcher who discovers a novel SQL-injection vector in a closed model must report it through a bug-bounty channel; a user who finds the same vector in an open-weight model can operationalise it immediately.
The tension is sharpest in jurisdictions where export controls and end-user verification remain weak. Anthropic notes that GLM-5.3's weights are hosted on public repositories with no geographic restriction, meaning a team in Pyongyang or Tehran can download, fine-tune and deploy the model in a matter of hours. By contrast, access to Claude is gated by know-your-customer checks, usage monitoring and rate limits that make large-scale abuse harder to sustain. The open-weight camp counters that determined adversaries will obtain frontier models regardless, so restricting access only punishes legitimate researchers. Anthropic's data, however, suggests that friction matters: internal telemetry shows that misuse attempts drop by an order of magnitude when even lightweight verification is in place.
What Z.ai Has Said - and Not Said
Z.ai has not issued a detailed response to Anthropic's findings. In a brief statement to Chinese technology media, a company spokesperson said the firm "prioritises responsible development" and pointed to a model card that lists acceptable-use guidelines. The card prohibits using GLM-5.3 for illegal activities, but contains no technical enforcement mechanism. Anthropic's researchers tested the guidelines by submitting prompts that walked up to - and over - the line of permissible use; the model complied with nearly all of them. The absence of a robust refusal system is not an oversight, several people familiar with Z.ai's design philosophy told Opentechwire. The company believes that safety should be implemented at the application layer by downstream developers, not baked into base weights that users cannot modify.
That position reflects a broader split in the Chinese AI ecosystem. Firms such as Alibaba and Baidu, which operate consumer-facing services, have adopted refusal classifiers and content filters to satisfy regulatory demands from Beijing's Cyberspace Administration. Smaller labs such as Z.ai, which target enterprise and research customers, argue that heavy-handed filtering hampers legitimate use cases - security teams probing their own defences, for instance, or academics studying adversarial robustness. The result is a two-tier landscape: models tuned for compliance, and models tuned for capability, with little middle ground.
Cyber Capability as a Lagging Indicator
Anthropic's assessment arrives at a moment when offensive AI is no longer theoretical. Over the past eighteen months, security vendors have logged a sharp uptick in automated reconnaissance and credential-stuffing campaigns that bear the signature of language-model assistance. Mandiant, the threat-intelligence arm of a major US technology firm, documented three intrusions in which attackers used AI-generated scripts to enumerate cloud storage buckets and exfiltrate credentials, shaving days off the typical dwell time. None of those incidents has been publicly attributed to a specific model, but the techniques match the playbook that GLM-5.3 and similar systems can execute.
The worry is not that AI will replace human operators, but that it will lower the skill floor. A decade ago, chaining exploits across a segmented network required deep knowledge of protocols, patience and creativity. Today, an attacker can describe the target environment in natural language, and a capable model will propose a multi-stage attack path, complete with evasion tactics. Anthropic's red team demonstrated this by feeding GLM-5.3 a network diagram and a list of installed software versions; the model returned a privilege-escalation sequence that a junior penetration tester would need weeks to devise. The same prompt sent to Claude triggered a refusal. The gap between those two outcomes is the policy space that governments and labs are now scrambling to fill.
Regional Implications: Seoul, Tokyo and Singapore Watch Closely
The disclosure has reverberated across Asia-Pacific capitals, where policymakers are weighing how to regulate dual-use AI without stifling innovation. South Korea's National Intelligence Service convened an inter-agency working group in early October to assess whether models like GLM-5.3 should be subject to import controls analogous to those governing intrusion software. Japan's Ministry of Economy, Trade and Industry is drafting guidelines that would require cloud providers to log inference requests for high-risk tasks, a move that could make open-weight deployment less attractive to enterprises worried about compliance overhead.
Singapore, which hosts both Chinese AI labs seeking regional footholds and Western firms expanding in Southeast Asia, faces a particularly delicate balancing act. The city-state's Cyber Security Agency has signalled that it will not ban open-weight models outright, but is exploring liability frameworks that hold deployers accountable for misuse, even if the base model was released elsewhere. That approach mirrors Singapore's stance on cryptographic export controls: permissive on research, strict on operational deployment. Whether it can work for AI remains an open question, because the line between research and operation is fuzzier when a model's output is executable code.
The Safety-Capability Divergence
Anthropic's report includes a chart that plots cyber task success rate against refusal rate for a dozen frontier models. The data points form two clusters: closed models from OpenAI, Google DeepMind and Anthropic itself, which score high on both capability and safety; and open-weight releases from Z.ai, several European labs and a handful of US start-ups, which score high on capability but low on safety. The gap between the clusters has widened over the past year, suggesting that the industry is bifurcating rather than converging on shared norms.
That divergence has tactical consequences. Defenders who rely on AI to hunt threats face an asymmetry: their tools are constrained by refusal logic, while attackers can use unconstrained models to generate polymorphic payloads that evade signature-based detection. Anthropic argues that the solution is not to remove safety layers from defensive models, but to tighten scrutiny of open releases and invest in behavioural detection that does not depend on static signatures. Critics worry that such measures will entrench the advantage of well-resourced incumbents and lock smaller teams out of the frontier.
What Comes Next
Anthropic has shared its findings with the US National Institute of Standards and Technology, which is developing a framework for evaluating dual-use foundation models. The institute is expected to publish draft guidance before year-end that would establish risk tiers based on task-specific benchmarks - precisely the kind of assessment Anthropic performed on GLM-5.3. If adopted, the framework could serve as a template for export-control decisions, with high-tier models subject to licensing requirements and end-user verification.
Z.ai, meanwhile, has given no indication that it will pull GLM-5.3 or retrofit safety features. The firm's investors, who include a state-backed venture fund and several private equity groups, view open-weight release as a strategic differentiator in a market where Western labs dominate closed offerings. For them, the controversy is validation: if Anthropic is worried, the model must be good. That calculus may shift if downstream liability becomes a reality, but for now the incentives point toward more capability, not more caution.
The broader question is whether the AI safety community can forge a consensus that spans borders and business models. Anthropic's report is a data point, not a policy prescription, but it crystallises the trade-off that labs and governments must navigate. Capability and safety are not zero-sum in the long run - better models can be both more powerful and more aligned. In the near term, however, the gap between what a model can do and what it will refuse to do is the space in which risk lives. How that space is managed will shape the cyber threat landscape for years to come.


