AnalysisAI ModelsJuly 24, 2026
Qwen3.5 35B A3B runs at 55 tok/s on RTX 5060 Ti with Garlic

A Reddit user achieved 55 tok/s running Qwen3.5 35B A3B in float8 on an RTX 5060 Ti by extending Garlic inference kernels. The work builds on prior optimization for Qwen3 30B A3B.
1 source
More stories today
- Tool automates complex projects with AI agent swarms
- Airtap launches AI agent for iMessage and RCS
- Sam Altman says 'we are now in the singularity'
- AI Mania Is Eviscerating Global Decision-Making
- Rauch: Software factory is product, agents are key