Can an Open Model Do Security Research? Cantina’s apex-flash-1 Solves 40 of 60 Held-Out Bug Tasks


Cantina Security, with Yeta Labs, has released apex-flash-1, an open-weights model trained specifically for vulnerability research. It is a reinforcement learning fine-tune of Z.ai’s GLM-5.3-Flash, released on Hugging Face under the MIT license.

Is it deployable? Yes, the MIT weights serve on vLLM, SGLang or Transformers, but BF16 needs roughly 640 GB of GPU memory.

What Cantina Built

apex-flash-1 has 321.3B total parameters, per its Hugging Face safetensors metadata. The GLM-5.3-Flash base is a Mixture-of-Experts model with 18B active parameters.

Cantina trained it with GRPO using a rank-256 LoRA plus selective full-parameter training. The data covers 150 tasks built from 50 real vulnerability cases.

Each case appears in 3 variants: guided whitebox, focused whitebox and focused blackbox.

Authorization, identity and scope flaws make up 72% of cases. Accounting and numerical precision bugs add 18%. Time validation, business rules and SSRF cover the rest.

As per the model card on HF, RL rollouts ran inside the Codex agent harness on production-like software and protocol environments.

Benchmark Results

Cantina evaluated 60 tasks from 20 held-out vulnerability cases. Each model ran the set once, with costs estimated from provider pricing.

  • apex-flash-1: 40/60 solved (66.7% pass@1), about $2.38
  • GLM-5.3-Flash (base): 36/60 solved (60.0%), about $4.56
  • Claude Opus 5 High: 43/60 solved (71.7%), about $74.68

Opus solved 3 more tasks but cost about 31x more per run. That is roughly $0.06 per solved task for apex-flash-1 versus $1.74 for Opus. These are company-reported numbers on an internal benchmark.

A Worker Model, Not an Orchestrator

Cantina positions apex-flash-1 as a worker orchestrated by a larger model. The card lists code reading, tool use, exploit development and verification as target skills.

An experimental apex-flash-1-abliterated variant ships with modified refusal behavior. It was not separately evaluated.

Cantina’s rationale is that defenders need capable models they can run and control locally.

Interactive Explainer



Source link

  • Related Posts

    GPT-6 Astra vs GPT-6.1 Sol vs Gemini 4 Argon vs Claude Fable 5.1: Which Frontier Model Fits Which Job

    Anthropic, OpenAI and Google DeepMind shipped 4 frontier-class models within 30 days. Claude Fable 5.1 arrived on September 1. GPT-6 Astra followed on September 3. GPT-6.1 Sol and Gemini 4…

    Aleph Alpha Releases Kolibri: A 78.1B Open-Weight English-German MoE Model With Only 3.46B Active Parameters

    Aleph Alpha has released Kolibri, an open-weight Mixture-of-Experts (MoE) language model built for German and English. Kolibri has 78.1B total parameters but activates only 3.46B, or 4.4%, per token. It…

    Leave a Reply

    Your email address will not be published. Required fields are marked *