Categories: FAANG

Multimodal Large Language Models with Fusion Low Rank Adaptation for Device Directed Speech Detection

Although Large Language Models (LLMs) have shown promise for human-like conversations, they are primarily pre-trained on text data. Incorporating audio or video improves performance, but collecting large-scale multimodal data and pre-training multimodal LLMs is challenging. To this end, we propose a Fusion Low Rank Adaptation (FLoRA) technique that efficiently adapts a pre-trained unimodal LLM to consume new, previously unseen modalities via low rank adaptation. For device-directed speech detection, using FLoRA, the multimodal LLM achieves 22% relative reduction in equal error rate (EER) over…
AI Generated Robotic Content

Recent Posts

GoT cast as Lebanese families

submitted by /u/Rokkit_man [link] [comments]

6 hours ago

Optimizing cost and latency with Amazon Bedrock prompt caching

Prompt caching in Amazon Bedrock can reduce your input token costs by up to 90…

6 hours ago

AI ‘Actor’ Tilly Norwood Told Me That ‘All Lives Matter’

The virtual character, which is promoting its upcoming movie Misaligned, tries to evade politics by…

7 hours ago

The shape behind the Einstein problem just revealed strange new physics

A mathematical shape famous for covering a surface without ever repeating has revealed an unexpected…

7 hours ago

AI can sound empathetic and human—but not at the same time

AI-generated texts are increasingly perceived as human, but people can still recognize human writing as…

7 hours ago

I trained the missing encoder for YuE2, so we can all bring our own music into it

YuE2 is an impressive open music model. Give it a style prompt and lyrics, and…

1 day ago