Categories: FAANG

Improving the Quality of Neural TTS Using Long-form Content and Multi-speaker Multi-style Modeling

Neural text-to-speech (TTS) can provide quality close to natural speech if an adequate amount of high-quality speech material is available for training. However, acquiring speech data for TTS training is costly and time-consuming, especially if the goal is to generate different speaking styles. In this work, we show that we can transfer speaking style across speakers and improve the quality of synthetic speech by training a multi-speaker multi-style (MSMS) model with long-form recordings, in addition to regular TTS recordings. In particular, we show that 1) multi-speaker modeling improves the…
AI Generated Robotic Content

Recent Posts

The Current State of Agentic AI

In this article, you will learn how agentic AI architecture has evolved by mid-2026, including…

6 hours ago

Environment-free Synthetic Data Generation for API-Calling Agents

Training API-calling large language model (LLM) agents demands massive amounts of high-quality trajectories. However, collecting…

6 hours ago

Exploring self-distilled reasoning for supervised fine-tuning with Amazon Nova

When you fine-tune a model using Supervised Fine-Tuning (SFT), creating high-quality chain-of-thought (CoT) reasoning traces…

6 hours ago

Why AI apps fail in production (And how Google solved it)

We are living in the golden age of the weekend AI side project. Thanks to…

6 hours ago

Is the All-New Range Rover GT Stepping on Jaguar’s Tail?

It’s “the most car-like Range Rover ever created,” but will this all-electric grand tourer spoil…

7 hours ago

AI detects ‘personalities’ of individual 3D printers to cut manufacturing errors

Imagine buying three identical 3D printers. Despite being the same brand, the same model and…

7 hours ago