DeepSeek releases V4.1-Flash with new YOCO decoder-decoder architecture

DeepSeek's V4.1-Flash carries 552B total parameters but activates only 8B per prompt, with native image input and up to a 1M-token context window. On OpenDesign Arena it scored 81.2/100, 98% of GPT-6 Astra's 82.7, at $0.023 per finished design versus Astra's $1.61.
How this story unfolded
6 days · 10 reports · 40 community posts · 50 of 52 shown
- Sep 8
- Sep 9
- Sep 10
DeepSeek formally launches V4.1 Flash, routes V4 Pro requests to Flashtechnode.com
DeepSeek V4.1 Flash now available on AI Gatewayvercel.com
DeepSeek AI Released DeepSeek-V4.1-Flash with 1M Context, FP4 KV Cache, and Cross-Layer Attention Reusemarktechpost.com
DeepSeek’s New Low-Cost Model Deals a Fresh Blow to OpenAI, Z.aibloomberg.com
deepseek-ai/DeepSeek-V4.1-Flash · Hugging Facehuggingface.co
DeepSeek's New Model Nearly Matches GPT-6 Astra on Design—at 1.4% of the Costdecrypt.co
- Sep 11
- Sep 12
- Sep 13
- Sep 14
More stories today
Superhuman acquires notetaker Fathom
Fathom has over 400,000 monthly active users and 1 million+ meetings recorded; it raised over $30M and was valued at $94M in 2024. Superhuman tested its own notetaker internally before buying a finished product.
TechCrunch·1 hour ago

AI coding spend yields 25% more output as code duplication rises 81%
AI coding tools delivered 25% more output per dollar of spend, but code duplication rose 81%, per an analysis of industry data. Rippling has added an anti-tokenmaxxing AI spend console as scrutiny of AI coding costs grows.
The New Stack·1 hour ago

Suno users report poor results from V6 update
Reddit users in r/SunoAI say the V6 update produces worse output, with one poster reporting that every suggested prompting method failed to restore prior quality.
r/SunoAI·1 hour agoReddit user runs Qwen 3.8 Next on 16GB VRAM, 32GB RAM
A r/LocalLLaMA user set up a REAP version of Qwen 3.8 Next on a 16GB VRAM / 32GB RAM system, calling it a challenge that might help others with limited hardware. No benchmarks yet against Qwen 3.8 27B QK4.
r/LocalLLaMA·1 hour ago
Laurie Voss: as coding costs collapse, product work becomes the job
Laurie Voss argues the cost of writing code has collapsed and the cost of reviewing, fixing and operating it is following. What remains is finding out what people want, defining it precisely, and making it pleasant to use — a per-software cost that doesn't transfer.
Simon Willison's Weblog·1 hour ago
TechCrunch Disrupt 2026 panel to debate AI-driven de-extinction
Colossal Biosciences co-founder and CEO Ben Lamm joins a Real World AI Stage fireside chat, "Can We Engineer Nature's Comeback?", at Disrupt 2026, held October 13–15 at Moscone West in San Francisco. The session covers AI's role in genetic analysis and biological modeling for de-extinction.
TechCrunch·1 hour ago

Reddit post claims China open sourced an RSI roadmap
A r/Singularity post titled "While USA 'Slows Down', China Just Open Sourced RSI Roadmap BTW" links to an r/accelerate thread. No article, official announcement, or corroborating source accompanies the claim.
r/Singularity·1 hour ago
Atria Dawn Preview research agent verifies its own output
Hasan Toor·2 hours ago