Cantina Security and Yeta Labs released apex-flash-1 (MIT), their first open-weights security research model: a GRPO post-train of GLM-5.3-Flash meant to work as a focused worker under a larger orchestrating agent, reading code, using tools, building exploits and verifying them against a running target. On 60 tasks from 20 held-out vulnerability cases it scores 66.7% pass@1 (40/60) vs 60.0% for base GLM-5.3-Flash and 71.7% for Claude Opus 5 High, at an estimated $2.38 per run vs $74.68 for Opus 5 High. An experimental abliterated variant ships alongside.
Key Takeaways
- ✓Score: 66.7% pass@1 (40/60) on held-out tasks vs 60.0% for base GLM-5.3-Flash and 71.7% for Claude Opus 5 High
- ✓Cost: ~$2.38 for the 60-task run vs $4.56 (GLM-5.3-Flash) and $74.68 (Opus 5 High) at provider pricing
- ✓Training: 150 tasks from 50 real vulnerability cases in three views; rank-256 LoRA on all experts and routers plus full updates to 16 activation-selected experts, GRPO
- ✓Data mix: 72% authorization/identity/scope-binding bugs, 18% numerical precision, remainder signature replay, payment rules and SSRF
- ✓License: MIT open weights on Hugging Face plus an experimental abliterated variant; Codex harness recommended
Developer 3s Key Decision Metrics
Turn your technical choice into a development budget
Compare 29+ dev plans & simulate token costs vs $20/mo subscriptions
Project Links & Resources
Direct AccessIn-Depth Technical Analysis
Cantina Security and Yeta Labs released apex-flash-1 (release post, weights), arguing that gating cyber capability in proprietary models mostly constrains defenders, since attackers can already run and modify open models. The model is a GRPO post-train of GLM-5.3-Flash on production-like environments built from real vulnerabilities Cantina found and disclosed: 50 cases, each in guided whitebox, focused whitebox and focused blackbox views (150 tasks), with verifiers checking final target state. Updates used rank-256 LoRA across all experts and routers plus full-parameter updates to 16 activation-selected experts; full updates to all experts collapsed the model.
On Cantina's internal held-out set (60 tasks from 20 unseen cases, first-draw pass@1) it solves 40/60 (66.7%) vs 36/60 for base GLM-5.3-Flash and 43/60 for Claude Opus 5 High, at an estimated $2.38 per run vs $4.56 and $74.68. This is not a public benchmark, and image/video performance is unevaluated. It is designed as a focused worker orchestrated by a larger model, with the Codex harness recommended. Weights are MIT-licensed, with an experimental abliterated variant for researchers. Use only in authorized, isolated environments.
Benchmark side-by-side against alternatives, or calculate monthly token cost vs subscription break-even.
Discussion & Comments
0Sign in to join the discussion
Connect with AI developers to exchange benchmark insights.