Breaking
ITItaly wins 3-0 against Sweden at the debut of the European volleyball championship in NaplesBRMother believes she won a truck in a draw and family laughs at the confusion on WhatsAppBRStorm with strong winds and hail causes tree falls and disruption in Campo GrandeKRTwo residents inhaled smoke in a house fire in Incheon... Property damage worth 33 million wonARIsraeli occupation forces kidnap a man and his son in Quneitra and take them to an unknown destinationTRIsrael's air and ground bombardment of Gaza hit civilian targets, killing and wounding many Palestinians.RUEmergency power outage in Petropavlovsk-Kamchatsky: consumers with a load of 17 MW were left without electricityRUBodø-Glimt goalkeeper Nikita Khaikin suffered a groin injury in the match against BayernCNTensions in the Middle East escalate: The game between the United States and Iran continues, Yemen’s Houthi armed forces occupy key areas in the Red Sea, and international oil prices exceed $100AUFirst-ever NFL game in Australia draws global fans to MCG for Rams-49ers clashITItaly wins 3-0 against Sweden at the debut of the European volleyball championship in NaplesBRMother believes she won a truck in a draw and family laughs at the confusion on WhatsAppBRStorm with strong winds and hail causes tree falls and disruption in Campo GrandeKRTwo residents inhaled smoke in a house fire in Incheon... Property damage worth 33 million wonARIsraeli occupation forces kidnap a man and his son in Quneitra and take them to an unknown destinationTRIsrael's air and ground bombardment of Gaza hit civilian targets, killing and wounding many Palestinians.RUEmergency power outage in Petropavlovsk-Kamchatsky: consumers with a load of 17 MW were left without electricityRUBodø-Glimt goalkeeper Nikita Khaikin suffered a groin injury in the match against BayernCNTensions in the Middle East escalate: The game between the United States and Iran continues, Yemen’s Houthi armed forces occupy key areas in the Red Sea, and international oil prices exceed $100AUFirst-ever NFL game in Australia draws global fans to MCG for Rams-49ers clash
BackDeepSeek V4.1 Flash nearly matches GPT-6 Astra in design benchmark at 1.4% of the cost
DeepSeek V4.1 Flash nearly matches GPT-6 Astra in design benchmark at 1.4% of the cost
Developing
Decrypt1 hour agoTech2 min read

DeepSeek V4.1 Flash nearly matches GPT-6 Astra in design benchmark at 1.4% of the cost

Quick Look

  • OpenDesign Arena benchmarked 13 AI models on real-world design tasks, with GPT-6 Astra scoring highest at 82.7 points and DeepSeek V4.1 Flash close behind at 81.2 points while costing only $0.023 per design versus $1.61 for the leader.
  • DeepSeek's model uses a Causal Encoder-Decoder architecture activating only 8-16 billion of its 552 billion parameters, enabling low cost and fast delivery.

AI-generated summary

Why It Matters

OpenDesign Arena evaluates AI models on practical design tasks like building web apps and landing pages, scoring both brief compliance and design quality. The benchmark focuses on real-world usability for working designers rather than general AI capabilities.

Font size

OpenDesign, the company behind the benchmark site OpenDesign Arena, ran 13 AI models through the same batch of design tasks this week. The top scorer was OpenAI's GPT-6 Astra. But DeepSeek's newest model, V4.1 Flash, reached 98% of that top score while charging about 1.4% of the top price.

OpenDesign Arena scores models on everyday design work—building web apps, dashboards, mobile screens, and landing pages—out of 100 points. Thirty of those points check whether the output actually meets the brief; the other 70 grade design quality on layout, hierarchy, color, and style fit.

It's built to answer a narrower question than most AI leaderboards ask: Which model should a working web designer actually use tomorrow.

On that scale, GPT-6 Astra averaged 82.7 points, taking 11.1 minutes and $1.61 per finished design. DeepSeek V4.1 Flash scored 81.2, finished the job in 5.3 minutes, and cost $0.023. Claude Fable 5.1 came in at 80.3, took 12.8 minutes, and cost $3.66.

Every other model OpenDesign tested—Grok 4.6, Qwen 3.8-Max, Kimi K3, GLM-5.3 Flash, and Gemini 3.8 Flash among them—scored lower than DeepSeek V4.1 Flash and cost more to run. That's 11 of the 13 models tested. Only GPT-6 Astra beat it outright, and only by a point and a half.

DeepSeek's technical report for V4.1 Flash explains where the savings come from. The model carries 552 billion parameters total—the internal settings a model tunes during training to store what it has learned—but wakes up only 8 billion of them to read an incoming prompt and 16 billion to write the response. DeepSeek calls this a Causal Encoder-Decoder design, and it's the same trick behind the model's fast completion times.

This isn't DeepSeek's first pass at closing a capability gap on the cheap. Weeks earlier, the company's V4 Pro model landed within 5% of Claude Fable 5 on a separate benchmark comparison while charging a fraction of Fable's rate. DeepSeek has also been recruiting engineers in Beijing to build its own Code Harness, aiming to own the full agentic stack instead of just supplying the model underneath it.

OpenDesign's testing setup narrows what these numbers can prove. A model's output only gets scored if it renders as a working webpage in the first place; anything blank, broken, or cut off scores zero and doesn't get retested. That means the benchmark measures reliable, everyday design output, not general reasoning or coding skill.

GPT-6 Astra, which OpenAI released on September 3, already carries a reputation for doing a bit of everything—laying out a circuit board, drafting a tax return, building a 3D scene—but early testers flagged it as a weaker writer than the model it replaced. Its price and pace on OpenDesign's chart fit that same generalist design: slower and pricier than DeepSeek's cheaper entry, but still the highest scorer in the field.

DeepSeek V4.1 Flash's delivery rate—the share of outputs OpenDesign judged ready to hand off without revision—came in at 57.7%. GPT-6 Astra's delivery rate was 60%. Claude Fable 5.1's was 56.7%.

What to Watch

AI outlook — possibilities, not facts

  • DeepSeek will gain increased adoption in cost-sensitive markets for AI-assisted design

    Likely · Within months

Open Questions

  • How does DeepSeek V4.1 Flash perform on non-design tasks like reasoning or coding?
  • What is the long-term sustainability of DeepSeek's low-cost model given its parameter activation strategy?
  • Will other AI companies adopt similar encoder-decoder architectures to reduce costs?

Related Topics

This article was originally published by Decrypt.

Related Stories

OpenAI's AI agents solve Navier-Stokes problem, with implications for smart-contract security
Developing·

OpenAI's AI agents solve Navier-Stokes problem, with implications for smart-contract security

OpenAI reported that 10,000 concurrent AI agents solved a Navier-Stokes fluid-motion problem in 88 hours, with formal verification in Lean taking another 17 hours using GPT-6 Astra. The breakthrough establishes cases C and D of the Millennium Prize formulation and could automate theorem proving for smart-contract security, reducing labor-intensive verification while increasing the importance of accurate specification design for DeFi protocols and tokenized assets.

CryptoSlate
2 min read
Researchers reduce quantum attack resource benchmark for Bitcoin and Ethereum cryptography by 86%
Developing·

Researchers reduce quantum attack resource benchmark for Bitcoin and Ethereum cryptography by 86%

Researchers using AI coding agents reduced the computational resource benchmark for a key step in a potential quantum attack on Bitcoin and Ethereum's cryptography by 86%, lowering the score from 10.75 billion to 1.496 billion in the ECDSA.Fail open competition. The work involved designing and testing a quantum circuit for cracking secp256k1 elliptic curve signatures, though no actual private key was cracked. Researchers note migration to post-quantum cryptography is underway, with NIST proposing to deprecate vulnerable algorithms after 2030.

Decrypt
2 min read
Superfluid Bug Drains Over $100,000 from GoodDollar Reserves via Excess G$ Creation on Celo
Developing·

Superfluid Bug Drains Over $100,000 from GoodDollar Reserves via Excess G$ Creation on Celo

A vulnerability in Superfluid's Celo deployment allowed an attacker to bypass liquidation safeguards, create excess G$ tokens, and drain over $100,000 from GoodDollar's reserves on Celo and XDC networks. GoodDollar reported 86,588 cUSD and $20,857 withdrawn, with external liquidity pools also affected. Both projects paused operations and are preparing incident reports to explain the cross-network impact.

CryptoSlate
2 min read
More on this topicopendesign arena