
Cantina’s apex-flash-1 Scores 40 of 60 Vulnerability Tasks
Published by AINave Editorial
Cantina Security and Yeta Labs released apex-flash-1, an open-weights model trained for vulnerability research. In Cantina’s internal benchmark, it solved 40 of 60 held-out tasks. That result landed below Claude Opus 5 High’s score, but the estimated run cost was much lower. Running the model locally in BF16, however, reportedly takes roughly 640 GB of GPU memory, so open weights do not make it a lightweight deployment.
A security-focused fine-tune, designed as a worker
Cantina built apex-flash-1 by reinforcement-learning fine-tuning Z.ai’s GLM-5.3-Flash. The model has 321.3 billion total parameters; the base model is a mixture-of-experts model with 18 billion active parameters. Cantina and Yeta Labs released the weights under the MIT license.
The training set comprised 150 tasks built from 50 real vulnerability cases. Each case had three task variants: guided whitebox, focused whitebox and focused blackbox. The cases skewed toward authorization, identity and scope flaws, which accounted for 72%; accounting and numerical-precision bugs made up 18%. Cantina identifies code reading, tool use, exploit development and verification as target skills.
The role matters: Cantina describes apex-flash-1 as a worker for a larger model to orchestrate, rather than presenting it as a standalone replacement for a security workflow. That framing fits a model trained on bounded research tasks, though the available result does not establish how well it transfers to other vulnerability sets or production work.
The benchmark gap and cost comparison
On Cantina’s evaluation of 60 tasks from 20 held-out cases, each model ran the set once. Cantina reported these results and estimated costs using provider pricing:
| Model | Solved, pass@1 | Estimated cost per 60-task run |
|---|---|---|
| apex-flash-1 | 40/60 (66.7%) | About $2.38 |
| GLM-5.3-Flash | 36/60 (60.0%) | About $4.56 |
| Claude Opus 5 High | 43/60 (71.7%) | About $74.68 |
Opus solved three more tasks than apex-flash-1, while its estimated run cost was about 31 times higher. The comparison is useful as a result on this particular test, not as a general ranking: one run of an internal 60-task benchmark cannot settle performance across security research in the wild. Nor does the run-cost figure describe the full cost of building and operating a system around the model.
Open weights, substantial hardware
Cantina says the weights can be served with vLLM, SGLang or Transformers. BF16 serving requires roughly 640 GB of GPU memory, which points to a multi-GPU setup rather than ordinary single-GPU experimentation. The article also mentions community 4-bit ports, but does not establish that they match the BF16 results or have official support.
That distinction shapes what “open” buys here. The license permits access to the weights, and compatible serving frameworks are named, but reproducing the reported configuration still carries a serious hardware requirement. For security teams, the central practical trade-off is not simply benchmark score versus price: it is the lower reported task-run cost against the infrastructure needed to host a 321.3-billion-parameter model.






















