LaunchDevelopersJuly 21, 2026

Weka launches storage platform that caches 100% of AI model's pre-computed tokens

The platform eliminates GPU memory recomputation by caching all pre-computed tokens for long contexts and multi-turn conversations. This reduces inference costs and GPU load. Weka claims it frees up significant GPU capacity.

1 source