AnalysisAI ModelsOctober 10, 2026

Custom CUDA megakernel speeds Qwen3.8-27B on a single RTX 3090

Read original source →reddit.com

A user reports a CUDA megakernel for Qwen3.8-27B running 1.4-1.9x faster than llama.cpp with MTP on an RTX 3090, hitting 140 tok/s on code generation. The kernel was written with Claude Opus 5.5.

2 sources

More stories today

Open the live feed