Categories: FAANG

Personalization of CTC-based End-to-End Speech Recognition Using Pronunciation-Driven Subword Tokenization

Recent advances in deep learning and automatic speech recognition have boosted the accuracy of end-to-end speech recognition to a new level. However, recognition of personal content such as contact names remains a challenge. In this work, we present a personalization solution for an end-to-end system based on connectionist temporal classification. Our solution uses class-based language model, in which a general language model provides modeling of the context for named entity classes, and personal named entities are compiled in a separate finite state transducer. We further introduce a…
AI Generated Robotic Content

Recent Posts

as far as is know it does t2i and i2i

submitted by /u/dev_ne [link] [comments]

12 hours ago

Build And Understand a Vector Database From Scratch in 10 Easy Steps

In this article, you will learn how a vector database works under the hood by…

12 hours ago

What’s Actually Inside 24,723 Tokens of a Search Result? We Broke It Down, Field by Field

Sponsored Content       It's no secret that AI agents burn massive amounts of…

12 hours ago

Dynamically Scaled Activation Steering

Activation steering has emerged as a powerful method for guiding the behavior of generative models…

12 hours ago

Leave the Class Path in the Rearview Mirror

Introducing composable, module system native and agent friendly command line tools for modern Java developmentBy…

12 hours ago

Amazon SageMaker Inference: 2026 year-to-date launches in review

Generative AI inference is uniquely hard: models are tens to hundreds of gigabytes, latency requirements…

12 hours ago