DeepSeek Releases V4.1 Flash AI Model Using Mixture-of-Experts Architecture
The new model activates only a fraction of its 552 billion parameters while outperforming several competitors on benchmarks.
Quick Look
Chinese AI developer DeepSeek launched V4.1 Flash, a 552 billion-parameter model using a Mixture-of-Experts design that activates fewer parameters to reduce computing costs while outperforming rivals on key benchmarks.
AI-generated summary
Why It Matters
Chinese AI companies face rising hardware costs and foreign chip export curbs that tighten compute constraints.
The Chinese artificial intelligence developer said on Thursday that V4.1 Flash used a new “Causal-Encoder-Decoder” architecture. While built on a massive 552 billion-parameter framework, it relies on a Mixture-of-Experts (MoE) design.
In traditional AI, every query runs through the entire system. An MoE system routes tasks only to the specific subnetworks best suited for the job. By activating just 8 billion parameters to process inputs and 16 billion to generate responses, DeepSeek said it significantly reduced the computing power needed per request.
DeepSeek described V4.1 Flash as the smallest model in its new series, with native multimodal visual understanding. The model outperformed V4 Pro on benchmarks evaluating coding, cybersecurity and autonomous agent tasks, the company said.
With rising hardware costs and foreign chip export curbs tightening compute constraints, Chinese players are racing to offer efficient models that deliver high-end reasoning at fraction-of-a-cent operational costs.
On Terminal-Bench 2.1, which tests AI on real-world compute tasks, V4.1 Flash scored 90.6, surpassing OpenAI’s GPT-5.6 Sol at 88.8, Moonshot AI’s Kimi K3 at 88.3, and DeepSeek’s own V4 Pro at 87.9.
Open Questions
- When will V4.1 Flash be widely available to commercial users?
- How do operating costs scale under heavy enterprise loads?




