AnalysisAI ModelsJune 26, 2026

Nemotron-3-Super-120B-A12B achieves full 504K token retrieval

The hybrid Mamba+MoE model holds perfect needle retrieval up to 504K tokens on 4×3090 GPUs using ~71GB. Mamba/SSM layers keep a constant-size recurrent state, making long context nearly free.

1 source