AnalysisDevelopersSeptember 15, 2026

Custom llama.cpp fork targets RTX 30-series with 90+ TPS

A Reddit user released a llama.cpp fork tuned for Ampere GPUs, claiming 90+ tokens per second at temperature 1 for agentic and coding workloads up to 100K context. Some optimizations also carry over to Blackwell and Lovelace cards.

1 source

More stories today

Open the live feed