Categories: FAANG

Improvements to Embedding-Matching Acoustic-to-Word ASR Using Multiple-Hypothesis Pronunciation-Based Embeddings

In embedding-matching acoustic-to-word (A2W) ASR, every word in the vocabulary is represented by a fixed-dimension embedding vector that can be added or removed independently of the rest of the system. The approach is potentially an elegant solution for the dynamic out-of-vocabulary (OOV) words problem, where speaker- and context-dependent named entities like contact names must be incorporated into the ASR on-the-fly for every speech utterance at testing time. Challenges still remain, however, in improving the overall accuracy of embedding-matching A2W. In this paper, we contribute two methods…
AI Generated Robotic Content

Recent Posts

A Tale of Two Flink Autoscalers

Samuel Yeboah, Francesco Di Chiara and Mingliang LiuToday, Netflix runs two Flink autoscalers. That is…

4 hours ago

Agentic Data Operations Platform (ADOP): Data engineering into hours

Data engineering teams routinely spend weeks standing up a single new data source: writing ETL,…

4 hours ago

Cloud CISO Perspectives: Sticking to security fundamentals in the AI era

Welcome to the first Cloud CISO Perspectives for August 2026. Today, Chris Betz explains why…

4 hours ago

The Unlikely Place at the Center of China’s AI Boom

Cheap energy, abundant land, and proximity to Beijing have turned a city in Inner Mongolia…

5 hours ago

AI could help design cities, but planners need safeguards

AI is showing up in nearly every aspect of daily life—from internet searches to visits…

5 hours ago

Sparse attention for H3 minimax, enjoy up to 2.5x speed up.

Added to my node pack, sparse attention SLA node for H3 Minimax. speed increase of…

1 day ago