Ant Group's Ling 3.0 Flash Beats 1T Parameter Model
Ant Group's inclusionAI released Ling 3.0 Flash, a 124B-parameter model that matches massive systems using just 5.1B active parameters to dramatically lower the cost of AI reasoning.

Ant Group's AI lab, inclusionAI, has released Ling 3.0 Flash, a 124-billion-parameter mixture-of-experts reasoning model. Despite its large total size, the model operates with only 5.1 billion active parameters per token. This architectural efficiency allows it to outperform a previous one-trillion-parameter flagship model on most benchmarks while using just an eighth of the total parameters and a twelfth of the active ones.
The model achieves these efficiency gains through a novel hybrid architecture that alternates Kimi Delta Attention, a linear attention mechanism, with Multi-head Latent Attention layers at a 5:1 ratio. It also halves the mixture-of-experts activation ratio to 1/64, down from 1/32, during pretraining. On the Artificial Analysis Intelligence Index, Ling 3.0 Flash scored 38, representing a 24-point jump over its predecessor and matching the performance of MiMo-V2.5 and Qwen3 27B with far fewer active parameters.
For practitioners, the model offers strong agentic capabilities, ranking second among flash-tier open-weights on the τ³-Banking benchmark. However, it is highly verbose, generating roughly four times more output tokens than its peers, which can increase the overall cost per task. Additionally, its hallucination rate fell from 97% to 44%, though this improvement stems primarily from the model learning to abstain rather than gaining factual knowledge, as its raw accuracy only ticked up from 16% to 18%.
The weights for Ling 3.0 Flash are available under an MIT license on Hugging Face in a 255GB BF16 format and a 128GB FP8 format. Developers can also access the model via APIs from inclusionAI and DeepInfra, with pricing set between $0.075 and $0.22 per one million tokens.
This is our own summary of reporting by AlphaSignal



