AnalysisPolicyAugust 6, 2026

New papers show AI fairness and explainer audits can be fooled

Four new arXiv papers probe audit integrity: a dual-penalty framework fools white-box explainers (LIME, SHAP, Integrated Gradients), and new lower bounds quantify how much companies can manipulate black-box fairness audits. One proposal counters this with manipulation-proof "oblivious" audits against deceptive model providers.

1 source

More stories today

Open the live feed
New papers show AI fairness and explainer audits can be fooled