Developer 3s Key Decision Metrics
On 2026-09-30 InclusionAI (Ant Group) announced Ling-3.1-flash: ~560B total / ~25B active MoE hybrid-reasoning model, designed for up to 1M context (trial/gateway routes often expose 256K–262K). Weights are not on Hugging Face yet. Vercel AI Gateway and OpenCode Zen list limited-time free access (~through 2026-10-13). Vendor-reported scores include SWE-Pro 65.39, Terminal-Bench 4.0 40.40, CyberGym 87.90—not independently reproduced.
Key Takeaways
- ✓Scale: ~560B total / 25B active (vs Ling-3.0-flash 124B/5.1B); design target 1M context; trial/gateway often 256K–262K
- ✓Architecture (vendor summary): 7×KDA : 1×Gated MLA; 512 routed experts pick 8 + 1 shared
- ✓Vendor benches (unreproduced): SWE-Pro 65.39, TB4.0 40.40, CyberGym 87.90, BrowseComp 91.67
- ✓Access: Vercel AI Gateway inclusionai/ling-3.1-flash(-free) via Novita free ~to Oct 13; OpenCode Zen ling-3.1-flash-free
- ✓Caveats: no public weights yet; free routes may use data to improve the model; some scores are environment-qualified
Finished reading? Explore benchmark rankings & pricing
Real-world SWE-bench scores & $20/mo vs API cost break-even calculator
Project Links & Resources
Direct AccessIn-Depth Technical Analysis
Core Background & Industry Pain Points
Flash-tier models are getting larger while staying sparse. InclusionAI (Ant Group) announced Ling-3.1-flash on 2026-09-30: ~560B total / ~25B active MoE for agents, search, office, and coding. Reality check: callable via gateways, weights not public; Ant’s docs page still lists the 3.0 family. Use Vercel AI Gateway or OpenCode Zen free promos for self-tests—don’t plan capacity on imaginary HF checkpoints.
Architecture Highlights & Internals
Vendor summary: hybrid-linear stack with higher linear-attention ratio (~7 KDA : 1 Gated MLA); 512 routed experts pick 8 + 1 shared. Context design up to 1M; trial/gateway routes often expose 256K–262K (Vercel: 262,144 / 32,768 out). Hybrid reasoning targets tool use and long histories. Do not reuse Ling-3.0-flash MIT deploy runbooks for 3.1.
Authoritative Benchmarks & Measured Scores
All headline scores are vendor-reported and largely unreproduced. Common cites: SWE-Pro 65.39, Terminal-Bench 4.0 40.40, CyberGym 87.90, BrowseComp 91.67; DeepSWE / TB2.1 trail DeepSeek-V4.1-Flash. HealthBench Professional 65.35 is AQ-environment-qualified. Treat charts as directional.
Developer Hands-on Guide
Try inclusionai/ling-3.1-flash(-free) on Vercel (free ~through Oct 13) or OpenCode ling-3.1-flash-free; wire via npx vercel ai-gateway setup. Keep proprietary code off free routes. Benchmark your own repo tasks vs current flash models. Plan as hosted until InclusionAI publishes weights + license.

Discussion & Comments
0Sign in to join the discussion
Connect with AI developers to exchange benchmark insights.