Categories: FAANG

Neural Transducer Training: Reduced Memory Consumption with Sample-wise Computation

The neural transducer is an end-to-end model for automatic speech recognition (ASR). While the model is well-suited for streaming ASR, the training process remains challenging. During training, the memory requirements may quickly exceed the capacity of state-of-the-art GPUs, limiting batch size and sequence lengths. In this work, we analyze the time and space complexity of a typical transducer training setup. We propose a memory-efficient training method that computes the transducer loss and gradients sample by sample. We present optimizations to increase the efficiency and parallelism of the…
AI Generated Robotic Content

Recent Posts

Do I look like I know what a VAE is!?

submitted by /u/YajuShinki [link] [comments]

7 hours ago

Agent or Workflow? A Practical Test for Knowing When You Actually Need an AI Agent

In this article, you will learn the key differences between AI workflows and agents, and…

7 hours ago

Compressing Streaming Neural Audio Encoders via Latent-Space Distillation

System-wide Dictation on Apple devices runs entirely on-device, and the speech it transcribes reaches the…

7 hours ago

Speaker-labeled transcription with WhisperX on SageMaker AI

Any team working with spoken audio hits the same wall with generic speech-to-text. Think contact-center…

8 hours ago

Power your agents: Gemini 3.8 Live with Live Avatar is now generally available

Following our announcement of Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking last week,…

8 hours ago

Anonymous Men Have Turned Cyberharassment Into a Group Sport—Here’s One Woman’s Side of the Story

This week on Uncanny Valley, we take you behind our feature on the women struggling…

8 hours ago