AnalysisAI ModelsJuly 19, 2026
Nick Ung on building evaluations that actually matter

Offline evals often pass at 90% but fail in production due to synthetic test sets that don't match real users. Nick Ung discusses how to build more representative evaluations.