Skip to content
AI.info

The Pulse

Cantina Releases Apex Flash-1, a 321B Open-Weight Security Model

Cantina says Apex Flash-1 solved 40 of 60 tasks across 20 held-out vulnerability cases, at a lower estimated run cost than the comparison models. The company released the 321-billion-parameter model with Yeta Labs on October 2, 2026, descri

Cantina Releases Apex Flash-1, a 321B Open-Weight Security Model

AI.info Team ·

Apex Flash-1 solved 40 of 60 security tasks in Cantina’s evaluation, a 66.7% pass-at-one score for the company’s new open-weight model. Cantina released the model on October 2, 2026, with Yeta Labs, describing it as a worker for focused vulnerability investigations rather than a stand-alone security agent. The 321-billion-parameter checkpoint is available on Hugging Face under an MIT license.

Cantina built 150 training tasks from 50 vulnerability cases

Cantina says it converted 50 vulnerability cases into 150 training tasks, giving each case three versions with different levels of access and direction. In guided white-box tasks, the model gets source code and detailed guidance; focused white-box tasks provide code but less direction; focused black-box tasks offer limited guidance and access to a running target without its source. Cantina used GLM-5.3-Flash as the base model and applied reinforcement learning with GRPO.

The cases draw on vulnerabilities Cantina says it found through security work. Its release says 72% of the 50 training cases involved authorization, identity or scope binding; the rest covered accounting and precision, time validation and signature replay, business rules and payment validation, and server-side request forgery. The company builds functioning, production-like environments and checks whether an agent achieves a defined result against a running target.

One example recreates an artifact-isolation flaw in Forgejo, an open-source code-hosting platform. An agent with access to one workflow had to trace how a signed download URL and a later artifact lookup handled workflow permissions differently, then retrieve an artifact from a separate, protected workflow. Cantina says its verifier checked whether the model recovered the protected artifact’s contents.

The 60-task result comes with a narrow comparison

On 60 tasks drawn from 20 held-out vulnerability cases, Apex Flash-1 solved 40, compared with 36 for base GLM-5.3-Flash and 43 for Claude Opus 5 High. Those results correspond to pass-at-one scores of 66.7%, 60% and 71.7%, respectively. Cantina estimated the cost of running the full set at $2.38 for Apex Flash-1, $4.56 for the base model and $74.68 for Opus 5 High, using provider pricing.

The figures come from Cantina’s own evaluation: each model ran the task set once, and the cases came from environments the company constructed. They offer a comparison on this particular test, not an independent assessment across other software, security teams or agent setups. Cantina’s model card describes the model as intended for text-based security tasks and says its image and video performance have not been evaluated.

An abliterated version changes the safety tradeoff

Cantina also released an experimental “abliterated” derivative with modified refusal behavior. The company frames local, controllable models as useful to defenders working within their own systems, arguing that attackers can adapt open models regardless of corporate usage policies. Cantina’s release states:

The problem is that cybersecurity is adversarial.

The modified-refusal version is not limited to security requests, according to its model card, and it has not received a separate full-suite evaluation. The 40-of-60 result belongs to the standard Apex Flash-1 checkpoint, not the derivative. Cantina says the model is intended for use under a larger agent’s direction; the downloadable weights, however, are not restricted to that arrangement.

Transfer beyond Cantina’s test remains unmeasured

The release packages security cases into a model that developers can download and run, while its benchmark remains a small company-run test built around Cantina’s own environments. The company says it plans to expand the data mix, add harder multi-step cases and train across different agent harnesses. Although Cantina reports results on 20 held-out vulnerability cases, how performance transfers to unfamiliar codebases or other orchestration setups is not established by this evaluation.

Sources

Explore

More articles