Categories: FAANG

Compressing Streaming Neural Audio Encoders via Latent-Space Distillation

System-wide Dictation on Apple devices runs entirely on-device, and the speech it transcribes reaches the foundation model through a tokenizer: an encoder that maps short windows of waveform onto the representation the language model reads. Because that model is sparsely activated under Instruction-Following Pruning, only a small subset of its experts occupies DRAM at any time, so the always-on tokenizer competes for the same memory, and its parameter count bears directly on power and latency. In this work we study how to compress such a tokenizer by distillation, taking as the supervision…
AI Generated Robotic Content

Recent Posts

Do I look like I know what a VAE is!?

submitted by /u/YajuShinki [link] [comments]

7 mins ago

Agent or Workflow? A Practical Test for Knowing When You Actually Need an AI Agent

In this article, you will learn the key differences between AI workflows and agents, and…

7 mins ago

Speaker-labeled transcription with WhisperX on SageMaker AI

Any team working with spoken audio hits the same wall with generic speech-to-text. Think contact-center…

7 mins ago

Power your agents: Gemini 3.8 Live with Live Avatar is now generally available

Following our announcement of Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking last week,…

7 mins ago

Anonymous Men Have Turned Cyberharassment Into a Group Sport—Here’s One Woman’s Side of the Story

This week on Uncanny Valley, we take you behind our feature on the women struggling…

1 hour ago

Making AI more trustworthy by making it red-flag its own doubtful answers

Artificial intelligence models can give users the wrong answer and do so with great confidence.…

1 hour ago