Categories: FAANG

How to Scale Your EMA

*=Equal Contributors
Preserving training dynamics across batch sizes is an important tool for practical machine learning as it enables the trade-off between batch size and wall-clock time. This trade-off is typically enabled by a scaling rule; for example, in stochastic gradient descent, one should scale the learning rate linearly with the batch size. Another important machine learning tool is the model EMA, a functional copy of a target model whose parameters move towards those of its target model according to an Exponential Moving Average (EMA) at a rate parameterized by a momentum…
AI Generated Robotic Content

Recent Posts

Whatever AI Safety Is, It’s Not This

Asking AI companies to self-regulate is a great way to pretend like you’ve accomplished something.

43 mins ago

This new qubit could be 100 times less error-prone in superfluid quantum computer breakthrough

A proposed qubit made with superfluid helium could cut quantum computing error rates by around…

43 mins ago

Can a machine or AI agent be surprised? Helping autonomous systems respond to the unexpected

Let's say you ask ChatGPT a question that stumps it, or a Waymo vehicle encounters…

43 mins ago

New Model Ideogram 4.5 (with edit) (open source soon)

submitted by /u/NewEconomy55 [link] [comments]

24 hours ago

On the Effectiveness-Fluency Trade-Off in LLM Conditioning: A Systematic Study

Controlling the output of Large Language Models (LLMs) is a central challenge for their reliable…

24 hours ago

Query claims in natural language with Amazon Bedrock Knowledge Bases

Claim answers are scattered across adjuster diary entries, repair estimates, police reports, payment ledgers, and…

24 hours ago