Ant Group surprises the AI world with Ling-3.0-Flash, an open-weight model with 124 billion parameters that activates only 5.1 billion per token and surpasses benchmarks of its one trillion predecessor. A feat that redefines computational efficiency and the cost of implementing AI agents.
Ant Group, the fintech giant behind Alipay, launched Ling-3.0-Flash on July 23, 2026. The model was developed by its inclusionAI lab and challenges the notion that in AI, bigger is always better.
Ling-3.0-Flash has a total of 124 billion parameters but activates only around 5.1 billion per token during inference. This Mixture of Experts (MoE) architecture allows only a subset of the neural network to engage for each input, details Cryptobriefing.
In the AI Intelligence Analysis Index, the model scored 38, matching or surpassing its predecessor Ling-2.6-1T, which had a trillion complete parameters. This makes it approximately eight times smaller in total parameter count but competitive in the metrics that matter for production deployments.
The main achievement lies in core reasoning and instruction-following tasks. The architecture combines hybrid reasoning with MoE, achieving efficiency without sacrificing quality.
Context length also advances: the native window is 262,000 tokens, with plans to expand to one million. The attention mechanism combines Kimi’s Delta Attention layers and Multi-Head Latent Attention, known as a hybrid-linear approach.
This design keeps memory and computation requirements manageable as context grows. Ant Group describes this technique as fundamental to supporting complex agent tasks.
Ling-3.0-Flash was specifically designed for production-grade AI agents running at high frequency. These workflows demand rapid token generation, reliable adherence to instructions, and handling of long context without degradation.
A massive library where the librarian only needs to pull five books at a time, regardless of how complex the question is. That’s the metaphor developers use to explain the approach.
Inference efficiency is key to reducing costs. By activating only 5.1B parameters, the computation required per query is directly reduced, making operation cheaper at scale.
Ant Group has already tested the model in its own Alipay services. Although specific figures were not disclosed, the company emphasizes that the model is optimized to handle complex production agent tasks.
The model was published on Hugging Face under the MIT license, the most permissive within open source. This allows for use, modification, and commercial redistribution without royalties.
Free access to the API was also offered through OpenRouter and Kilo until August 3, 2026. Additionally, it is accessible through Ant Group’s channels and Vercel’s AI gateway.
The MIT license removes barriers for developers and companies wanting to integrate it into their products. This contrasts with other open-source models that impose restrictions on commercial use.
Ant Group’s decision to open the model aims to position the company as a leader in efficiency. According to the original report, this could accelerate AI adoption in sectors where inference costs are prohibitive.
If a 124B MoE model with 5.1B active can match or surpass the baseline of a trillion parameters, the cost structure for implementing capable AI decreases significantly. Inference is where AI companies spend money after training.
This advancement could pressure other labs to rethink their scaling strategies. Efficiency becomes a key competitive factor, not just the number of parameters.
The AI agent market directly benefits, as response times are critical. Models like Ling-3.0-Flash allow for faster interactions and predictable costs.
CryptoBriefing highlights that this is a notable step in the AI efficiency race, with implications for the entire industry. The combination of speed, low cost, and open license could change the landscape of artificial intelligence.
Ant Group demonstrates that innovation does not always mean building the largest model, but the smartest for real-world use. The development community can already experiment with this new generation of efficient AI.
This content is provided for general informational purposes only and doesn't constitute financial, investment, legal, or tax advice. Any events, rewards, online promotions, or related information mentioned herein should not be considered a recommendation, solicitation, or invitation to purchase, sell, trade, or otherwise deal in any crypto assets. Crypto assets are highly volatile and may result in loss. The availability of WEEX services, products, and related events may vary by region. You are responsible for ensuring that your participation is in accordance with applicable local laws and regulations.





![[Editorial] In a Market Where Uncertainty is the Norm, 'Resilience' Ultimately Determines Success](/public-static/33_70806c0ee0.png?format=avif)























