Alibaba releases Qwen3.8-Flash-Next, previewing Qwen4 architecture
Qwen3.8-Flash-Next is a multimodal MoE with 125B parameters plus 51B N-gram embeddings, activating only 6B per token. It has a 262K-token native context, extensible to 1M with YaRN, and beats Claude Opus 4.6 Max on 8 of 9 comparable benchmarks. QwenCloud API pricing: $0.16/1M input and $0.47/1M output tokens.
How this story unfolded
9 days · 15 reports · 51 community posts · 66 of 70 shown
- Aug 25
- Aug 26
Alibaba’s Qwen to open-source Qwen3.8-Flash-Next, previewing Qwen4 architecturetechnode.com
unsloth/Qwen3.8-Flash-Next-GGUFhuggingface.co
Qwen/Qwen3.8-Flash-Next · Hugging Facehuggingface.co
Alibaba’s Qwen Team Releases Qwen3.8-Flash-Next: A 125B Multimodal MoE With 6B Active Parameters Previewing the Qwen4 Architecturemarktechpost.com
Qwen/Qwen3.8-Flash-Next-FP8huggingface.co
Experiment with Qwen3.8-Flash-Next 176B Model on NVIDIA GB300 NVL72 for Agentic Codingdeveloper.nvidia.com
- Aug 27
- Aug 28
- Aug 29
- Aug 30
- Aug 31
- Sep 1
- Sep 2
- Sep 3
More stories today
Martian launches AI Frontier dashboard comparing 44 LLMs
TestingCatalog News·37 minutes ago
Reddit user shares skill to prevent AI agent errors
A Reddit user created a skill that fires whenever an AI agent makes a mistake, reducing repetitive apologies. The proof-of-concept is available on GitHub.
r/LocalLLaMA·45 minutes ago
Cut GPU inference cold start from 8 minutes to under a minute
Instrumenting the full path from pod creation to first inference response on a GPU node running a 70B-class model revealed six sequential phases, not one bottleneck. For a 64 GB model, 65% of startup time is spent in one phase.
The New Stack·46 minutes ago

Qwen by email
Get an email when Qwen has news
No news that day, no email.
Guide to production token optimization for scaling AI apps
The New Stack's guide explains how enterprise AI applications hit a token-consumption wall as they scale, often misdiagnosed as a billing issue. It covers strategies for optimizing token usage in production.
The New Stack·46 minutes ago

Cerebras unveils CS-4 with WSE-3 Turbo, claims 30x faster inference
Cerebras announced the CS-4 rack-scale AI system, powered by three WSE-3 Turbo chips with 4 trillion transistors and 900,000 AI cores per wafer. The company claims up to 30x faster inference than conventional GPUs, targeting frontier AI and real-time agentic workloads.
YouTube·50 minutes ago
OpenAI's Brockman: AGI may have been created around this time
Kimmonismus·54 minutes ago
Google study: transfer learning for genomic prediction in underrepresented populations
Google Research evaluated transfer learning from European cohorts to improve polygenic risk score prediction in underrepresented populations, finding it helps small cohorts but degrades accuracy as target sample sizes grow, especially for traits with population-specific genetic architectures.
Google Research Blog·56 minutes ago

Meta offers 95% discount on Muse Spark for users sharing data
Meta's Muse Spark model offers an explicit discount averaging about 95% for users who share prompts and outputs. Standard pricing is $1.25 per million input tokens and $4.25 per million output tokens; contributor pricing drops these to 10 cents and 20 cents respectively.
TechCrunch·56 minutes ago
