Categories: FAANG

Improving the Quality of Neural TTS Using Long-form Content and Multi-speaker Multi-style Modeling

Neural text-to-speech (TTS) can provide quality close to natural speech if an adequate amount of high-quality speech material is available for training. However, acquiring speech data for TTS training is costly and time-consuming, especially if the goal is to generate different speaking styles. In this work, we show that we can transfer speaking style across speakers and improve the quality of synthetic speech by training a multi-speaker multi-style (MSMS) model with long-form recordings, in addition to regular TTS recordings. In particular, we show that 1) multi-speaker modeling improves the…
AI Generated Robotic Content

Recent Posts

Measuring Performance of Transformer Inference

This chapter is divided into eight parts; they are: • Metrics for LLM Inference •…

6 hours ago

Static vs. Dynamic vs. Continuous Batching in LLM Inference

In this article, you will learn how static, dynamic, and continuous batching work in LLM…

6 hours ago

Introducing Web Search on Amazon Bedrock for foundation model grounding

When a foundation model needs to answer a question about last week’s earnings call, yesterday’s…

6 hours ago

How Deutsche Bank unlocked agility with an API-ready ecosystem

When people think about digital transformation in banking, they often focus on the visible results:…

6 hours ago

OK, Well, Rogue AI Agents Are Hacking Again

Rogue AI agents from OpenAI and Anthropic have again been caught trying to disrupt servers…

7 hours ago

People prefer stories written by AI—especially when told they’re written by a human

People gave the highest ratings to AI-generated stories they were told had been written by…

7 hours ago