On 2026-09-30 InclusionAI (Ant Group) announced Ling-3.1-flash: ~560B total / ~25B active MoE hybrid-reasoning model, designed for up to 1M context (trial/gateway routes often expose 256K–262K). Weights are not on Hugging Face yet. Vercel AI Gateway and OpenCode Zen list limited-time free access (~through 2026-10-13). Vendor-reported scores include SWE-Pro 65.39, Terminal-Bench 4.0 40.40, CyberGym 87.90—not independently reproduced.

Key Takeaways

  • ✓Scale: ~560B total / 25B active (vs Ling-3.0-flash 124B/5.1B); design target 1M context; trial/gateway often 256K–262K
  • ✓Architecture (vendor summary): 7×KDA : 1×Gated MLA; 512 routed experts pick 8 + 1 shared
  • ✓Vendor benches (unreproduced): SWE-Pro 65.39, TB4.0 40.40, CyberGym 87.90, BrowseComp 91.67
  • ✓Access: Vercel AI Gateway inclusionai/ling-3.1-flash(-free) via Novita free ~to Oct 13; OpenCode Zen ling-3.1-flash-free
  • ✓Caveats: no public weights yet; free routes may use data to improve the model; some scores are environment-qualified
🧭

Finished reading? Explore benchmark rankings & pricing

Real-world SWE-bench scores & $20/mo vs API cost break-even calculator

🔬

In-Depth Technical Analysis

Core Background & Industry Pain Points

Flash-tier models are getting larger while staying sparse. InclusionAI (Ant Group) announced Ling-3.1-flash on 2026-09-30: ~560B total / ~25B active MoE for agents, search, office, and coding. Reality check: callable via gateways, weights not public; Ant’s docs page still lists the 3.0 family. Use Vercel AI Gateway or OpenCode Zen free promos for self-tests—don’t plan capacity on imaginary HF checkpoints.

Architecture Highlights & Internals

Vendor summary: hybrid-linear stack with higher linear-attention ratio (~7 KDA : 1 Gated MLA); 512 routed experts pick 8 + 1 shared. Context design up to 1M; trial/gateway routes often expose 256K–262K (Vercel: 262,144 / 32,768 out). Hybrid reasoning targets tool use and long histories. Do not reuse Ling-3.0-flash MIT deploy runbooks for 3.1.

Authoritative Benchmarks & Measured Scores

All headline scores are vendor-reported and largely unreproduced. Common cites: SWE-Pro 65.39, Terminal-Bench 4.0 40.40, CyberGym 87.90, BrowseComp 91.67; DeepSWE / TB2.1 trail DeepSeek-V4.1-Flash. HealthBench Professional 65.35 is AQ-environment-qualified. Treat charts as directional.

Developer Hands-on Guide

Try inclusionai/ling-3.1-flash(-free) on Vercel (free ~through Oct 13) or OpenCode ling-3.1-flash-free; wire via npx vercel ai-gateway setup. Keep proprietary code off free routes. Benchmark your own repo tasks vs current flash models. Plan as hosted until InclusionAI publishes weights + license.