Llama.cpp pull request moves sampling to GPU for inference speed gains

A new pull request moves sampling to the GPU, yielding an 8% increase in tokens per second for Qwen3.6:35b on an NVIDIA 5090. Testing on an NVIDIA P40 showed a 4% inference speed boost.
1 source
Daily brief
Get tomorrow's AI brief in your inbox
More stories today
- NVIDIA releases Alpamayo 2 Super, open-source 34B model for robotaxis
- Omnigent: open source AI meta-harness unifies coding agents
- Rezolve Ai argues commerce AI has a measurement problem
- User compares Claude Code vs Codex iOS apps, says OpenAI wins on UI/UX
- Russell Kaplan shares how Cognition builds proactive agents