Ling-3.0-Flash
Ant Group's efficient open-weight reasoning model built for production AI agents
Ling-3.0-Flash is the latest entry in Ant Group’s efficiency-focused Ling line, built by its inclusionAI lab and unveiled July 23, 2026 as a hybrid-reasoning foundation model for production AI agent workloads. It’s a 124 billion parameter mixture-of-experts model that activates only about 5.1 billion parameters per token, natively supports a 256,000 token context window that can scale up to 1 million, and Ant says it matches or beats models two to three times its parameter count on reasoning, instruction following, and long-context benchmarks. A cluster-level hierarchical caching system cuts time-to-first-token on long conversations by 60 to 80 percent. Ant priced the first-party API at $0.075 per million input tokens and $0.22 per million output tokens, among the cheapest rates measured for a model scoring 38 or above on the Artificial Analysis Intelligence Index, and released the weights on Hugging Face under the MIT license.