Categories: FAANG

Joint Speech Transcription and Translation: Pseudo-Labeling with Out-of-Distribution Data

Self-training has been shown to be helpful in addressing data scarcity for many domains, including vision, speech, and language. Specifically, self-training, or pseudo-labeling, labels unsupervised data and adds that to the training pool. In this work, we investigate and use pseudo-labeling for a recently proposed novel setup: joint transcription and translation of speech, which suffers from an absence of sufficient parallel data resources. We show that under such data-deficient circumstances, the unlabeled data can significantly vary in domain from the supervised data, which results in…
AI Generated Robotic Content

Recent Posts

Meet Dyson’s New Robot Vacuum Line: The Dyson Nurovi Line (2026)

The Nurovi line includes three lidar-powered robovacs. But to get the model I’m most intrigued…

5 mins ago

One material, two transistor types: ‘Universal charge injector’ points toward densely stacked AI chips

A new approach could help make future AI chips smaller and more energy-efficient. A KAIST-led…

5 mins ago

GoT cast as Lebanese families

submitted by /u/Rokkit_man [link] [comments]

23 hours ago

Optimizing cost and latency with Amazon Bedrock prompt caching

Prompt caching in Amazon Bedrock can reduce your input token costs by up to 90…

23 hours ago

AI ‘Actor’ Tilly Norwood Told Me That ‘All Lives Matter’

The virtual character, which is promoting its upcoming movie Misaligned, tries to evade politics by…

1 day ago

The shape behind the Einstein problem just revealed strange new physics

A mathematical shape famous for covering a surface without ever repeating has revealed an unexpected…

1 day ago