AnalysisAI ModelsSeptember 26, 2026

GLiNER2.5-Decide converted to LiteRT for on-device Android GPU inference

Read original source →huggingface.co

The 340M DeBERTa-v3-large model runs at 66-70 ms per request on a Galaxy S26 GPU at 128 tokens, 175 ms at 256. Its decisions matched the official fp32 model on all 361 desktop test requests and all 126 Android (request, window) pairs.

2 sources

More stories today

Open the live feed