GLM-5.3 Hacking Capabilities: What Anthropic’s Tests Show
tomshardware.com

GLM-5.3 Hacking Capabilities: What Anthropic’s Tests Show

Tech News
4 min read

Published by AINave Editorial

TL;DRAnthropic’s tests put GLM-5.3 close to Claude Mythos on two specific exploit benchmarks. The findings also separate the risks of bypassing safeguards from the added access that comes with downloadable model weights.

Anthropic says Zhipu AI’s open-weight GLM-5.3 came close to its Claude Mythos model on two cybersecurity tests. The results are notable, but narrower than a claim of broad parity: they measure performance in particular exploit benchmarks, while the reported safety concern is that GLM-5.3’s safeguards can be bypassed or deliberately removed.

Close results on two specific exploit tests

In Anthropic’s sandboxed Exploitbench test targeting Google Chrome, GLM-5.3 produced end-to-end exploits in 50 of 410 attempts. Claude Mythos produced 56 in the same number of attempts, according to Anthropic’s reported results. A separate internal test of full control-flow hijacks put GLM-5.3 at a 4% success rate and Mythos at 6%; the same report says Kimi K3 and DeepSeek V4.1 Flash scored 0% on that measure in the benchmark.

These comparisons show strong performance on defined tasks, not that the models are interchangeable across cybersecurity work. The distinction matters because exploit development is only one part of security, and the figures come from Anthropic’s own tests rather than an independent evaluation. A separate account also describes the 50-of-410 result as Anthropic’s test of GLM-5.3’s ability to build end-to-end exploits against Mythos Preview.

Bypassing safeguards is different from removing them

Anthropic reports that GLM-5.3 could develop chained exploits autonomously and describes two ways its safeguards were bypassed: a deceptive role-play prompt succeeded 64% of the time, while prefilling the model’s thinking tokens succeeded 92% of the time in Anthropic’s tests. These are reported outcomes in those tests, not a general measure of how often users can defeat the model’s protections.

Anthropic separately tested an altered version with its guardrails deliberately removed, a process the article calls “abliteration.” The stock model’s refusal rate was described as on par with Anthropic models; after abliteration, reported refusal rates fell to 6% for GLM-5.3 and 14% for GLM-5.3-Flash in the altered models. That distinction is central: open weights can be modified, while the closed-source Anthropic models discussed in the report cannot be altered in the same way by users.

The capability has a substantial compute cost

That flexibility does not make running the full model cheap. The article estimates that full-precision FP8 weights require 306 GB of VRAM, with a similar amount needed for the KV cache. For an estimated 100 tokens per second, it cites at least 4 TB/s of memory bandwidth and a cluster of eight Nvidia H200 accelerators as a possible setup. These are estimates, not a universal minimum configuration.

The practical picture is therefore more complicated than the benchmark comparison alone suggests. Downloadable weights make modification possible, but running an altered full-size model at the described speed calls for substantial hardware or rental expense. Anthropic’s results raise a real question about who can access powerful exploit development, while the reported compute demands suggest that capability and practical availability are not the same thing.

FAQs

Anthropic says GLM-5.3 produced end-to-end exploits in a sandboxed Chrome benchmark and could develop chained exploits autonomously in its tests. Those findings describe specific evaluations, not proven performance in every real-world attack setting reported by Anthropic.

Sources

Latest Tech News