Categories: FAANG

Scaling Laws for Optimal Data Mixtures

Large foundation models are typically trained on data from multiple domains, with the data mixture—the proportion of each domain used—playing a critical role in model performance. The standard approach to selecting this mixture relies on trial and error, which becomes impractical for large-scale pretraining. We propose a systematic method to determine the optimal data mixture for any target domain using scaling laws. Our approach accurately predicts the loss of a model of size N trained with D tokens and a specific domain weight vector h. We validate the universality of these scaling laws by…
AI Generated Robotic Content

Recent Posts

LLMs Are Not (Consistently) Bayesian: Quantifying Internal (In)consistencies of LLMs’ Probabilistic Beliefs

Modern AI systems are being deployed in complex domains such as medicine, science, and law,…

33 mins ago

MAPS: Netflix’s Multimodal Asset Personalization at Scale

By Emma Yanyang Kong, Aditya Deshpande, Asad Abbasi, Bowei Yan, David Fagnan, Ashish Rastogi, Dhaval…

33 mins ago

Batch write and discover records in Amazon SageMaker Feature Store

Amazon SageMaker Feature Store is a fully managed, purpose-built repository to store, share, and manage…

33 mins ago

Nvidia CEO Jensen Huang Took a Call From Donald Trump in the Middle of an All-Hands

The unexpected interruption came hours before the president wrote a congratulatory post on Truth Social…

2 hours ago

Shared-memory AI system lets microscope components coordinate in real time

Arco Bast studies how neurons communicate. Earlier this year, the Janelia postdoc encountered a more…

2 hours ago