Scaling Laws for Mixture Pretraining Under Data Constraints

Across more than 2,000 language-model training runs, Apple ML Research finds scarce target corpora can be reused 15–20 times in mixture pretraining, with repetition the central driver of target-domain performance. The proposed repetition-aware scaling law covers multilingual, domain-specific, and quality-filtered data mixtures.
1 source
Apple by email
Get an email when Apple has news
No news that day, no email.
More stories today
- LangSmith adds Preview Builds to test agent changes before production
- Exa plugin gives ChatGPT Work and Codex access to 100B+ websites
- Ramp launches its own AI model router, called Router
- Apple’s AirPods Should Avoid Meta’s Mistakes
- Docker's Tushar Jain on AI-native runtime for agent autonomy