New papers show AI fairness and explainer audits can be fooled
Four new arXiv papers probe audit integrity: a dual-penalty framework fools white-box explainers (LIME, SHAP, Integrated Gradients), and new lower bounds quantify how much companies can manipulate black-box fairness audits. One proposal counters this with manipulation-proof "oblivious" audits against deceptive model providers.
How this story unfolded
2 days · 4 reports · from Aug 4
- Aug 4
- Aug 5
- Aug 6
Daily brief
Get tomorrow's AI brief in your inbox
More stories today
- Qwen 3.6 runs locally on Radeon R9700 and multi-Mac agent farms
- Tool compiles knowledge bases and runs parallel research across Claude, Codex, Pi
- ESP32-S3 board turned into $5 personal AI assistant for Telegram
- Filmmaker releases 'Memories' short film made with H3 and Seedance 2.5
- WorkOS argues REST and MCP are complementary, not competing, for agents