AnalysisPolicyAugust 6, 2026

New papers show AI fairness and explainer audits can be fooled

Four new arXiv papers probe audit integrity: a dual-penalty framework fools white-box explainers (LIME, SHAP, Integrated Gradients), and new lower bounds quantify how much companies can manipulate black-box fairness audits. One proposal counters this with manipulation-proof "oblivious" audits against deceptive model providers.

How this story unfolded

2 days · 4 reports · from Aug 4

  1. Aug 4
  2. Aug 5
  3. Aug 6

Daily brief

Get tomorrow's AI brief in your inbox

More stories today

Open the live feed