Categories: FAANG

Multimodal Large Language Models with Fusion Low Rank Adaptation for Device Directed Speech Detection

Although Large Language Models (LLMs) have shown promise for human-like conversations, they are primarily pre-trained on text data. Incorporating audio or video improves performance, but collecting large-scale multimodal data and pre-training multimodal LLMs is challenging. To this end, we propose a Fusion Low Rank Adaptation (FLoRA) technique that efficiently adapts a pre-trained unimodal LLM to consume new, previously unseen modalities via low rank adaptation. For device-directed speech detection, using FLoRA, the multimodal LLM achieves 22% relative reduction in equal error rate (EER) over…
AI Generated Robotic Content

Recent Posts

Do I look like I know what a VAE is!?

submitted by /u/YajuShinki [link] [comments]

8 hours ago

Agent or Workflow? A Practical Test for Knowing When You Actually Need an AI Agent

In this article, you will learn the key differences between AI workflows and agents, and…

8 hours ago

Compressing Streaming Neural Audio Encoders via Latent-Space Distillation

System-wide Dictation on Apple devices runs entirely on-device, and the speech it transcribes reaches the…

8 hours ago

Speaker-labeled transcription with WhisperX on SageMaker AI

Any team working with spoken audio hits the same wall with generic speech-to-text. Think contact-center…

8 hours ago

Power your agents: Gemini 3.8 Live with Live Avatar is now generally available

Following our announcement of Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking last week,…

8 hours ago

Anonymous Men Have Turned Cyberharassment Into a Group Sport—Here’s One Woman’s Side of the Story

This week on Uncanny Valley, we take you behind our feature on the women struggling…

9 hours ago