NVIDIAAnalysisDevelopersSeptember 16, 2026

TensorRT Edge-LLM runs MLPerf Edge Agentic 6.4x faster on Jetson AGX Thor

NVIDIA's TensorRT Edge-LLM ran Qwen3.6-27B on a single Jetson AGX Thor Developer Kit at 52.33 tokens/sec, finishing all 1,007 turns in 24 minutes 36 seconds versus llama.cpp's 2 hours 37 minutes. It uses NVFP4 quantization, tree-based multi-token prediction, and KV cache reuse.

1 source

More stories today

Open the live feed