AppleAnalysisAI ModelsAugust 20, 2026

Scaling Laws for Mixture Pretraining Under Data Constraints

Read original source →machinelearning.apple.com

Across more than 2,000 language-model training runs, Apple ML Research finds scarce target corpora can be reused 15–20 times in mixture pretraining, with repetition the central driver of target-domain performance. The proposed repetition-aware scaling law covers multilingual, domain-specific, and quality-filtered data mixtures.

1 source

More stories today

Open the live feed