Categories: FAANG

Improving the Quality of Neural TTS Using Long-form Content and Multi-speaker Multi-style Modeling

Neural text-to-speech (TTS) can provide quality close to natural speech if an adequate amount of high-quality speech material is available for training. However, acquiring speech data for TTS training is costly and time-consuming, especially if the goal is to generate different speaking styles. In this work, we show that we can transfer speaking style across speakers and improve the quality of synthetic speech by training a multi-speaker multi-style (MSMS) model with long-form recordings, in addition to regular TTS recordings. In particular, we show that 1) multi-speaker modeling improves the…
AI Generated Robotic Content

Recent Posts

Introducing Claude Fable 5.1 on AWS

Today, we’re excited to announce the availability of Claude Fable 5.1 on Amazon Bedrock and…

8 hours ago

The Range Rover Electric: Specs, Price, Availability

After long delays, JLR’s biggest gamble with its Range Rover brand is here with huge…

8 hours ago

A new kind of AI that does its thinking cheaply without words

There may soon be a new kind of artificial intelligence in town, one that uses…

8 hours ago

Connect an AgentCore Runtime hosted MCP server to Amazon Quick

Model Context Protocol (MCP) servers allow foundation models to access external data and tools, supporting…

1 day ago

The Best Labor Day Mattress Deals on Beds We’ve Tried in Our Homes

It’s one of the best times of the year to buy a mattress, and our…

1 day ago

A “quantum bath” puts quantum entanglement on autopilot

Physicists have demonstrated a new way to entangle distant quantum bits without the constant measurements…

1 day ago