Categories: FAANG

The Slingshot Mechanism: An Empirical Study of Adaptive Optimizers and the Grokking Phenomenon

This paper was accepted to the “Has it Trained Yet?” (HITY) workshop at NeurIPS 2022.
The grokking phenomenon as reported by Power et al., refers to a regime where a long period of overfitting is followed by a seemingly sudden transition to perfect generalization. In this paper, we attempt to reveal the underpinnings of Grokking via a series of empirical studies. Specifically, we uncover an optimization anomaly plaguing adaptive optimizers at extremely late stages of training, referred to as the Slingshot Mechanism. A prominent artifact of the Slingshot Mechanism can be measured by the cyclic…
AI Generated Robotic Content

Recent Posts

This sub has had a distinct lack of dancing 1girls lately

So many posts with actual new model releases and technical progression, why can't we go…

8 hours ago

10 Common Misconceptions About Large Language Models

Large language models (LLMs) have rapidly integrated into our daily workflows.

8 hours ago

TII Falcon-H1 models now available on Amazon Bedrock Marketplace and Amazon SageMaker JumpStart

This post was co-authored with Jingwei Zuo from TII. We are excited to announce the…

8 hours ago

Scaling high-performance inference cost-effectively

At Google Cloud Next 2025, we announced new inference capabilities with GKE Inference Gateway, including…

8 hours ago

‘War Is Here’: The Far-Right Responds to Charlie Kirk Shooting With Calls for Violence

Prominent far-right figures and elected officials have called for vengeance following the death of conservative…

9 hours ago

Just tried HunyuanImage 2.1

Hey guys, I just tested out the new HunyuanImage 2.1 model on HF and… wow.…

1 day ago