New research exposes LLM unlearning gaps and relearning risks
Three arXiv papers reveal that unlearned LLMs often fail to stay unlearned: brief fine-tuning can revive removed knowledge, and existing robustness predictors relying on weight-space distance are insufficient.
How this story unfolded
3 days · 4 reports · from Aug 25
- Aug 25
- Aug 27
- Aug 28
Daily brief
Get tomorrow's AI brief in your inbox
More stories today
- Skeptic shares first impressions of ChatGPT Plus
- Hard sci-fi authors largely oppose LLMs, survey finds
- Anthropic tests folderless Claude Code sessions on Desktop and iOS
- Ethan Mollick: Using weaker AI for human-facing content may soon be disrespectful
- Pocket TTS training stack open-sourced