Categories: FAANG

Multimodal Large Language Models with Fusion Low Rank Adaptation for Device Directed Speech Detection

Although Large Language Models (LLMs) have shown promise for human-like conversations, they are primarily pre-trained on text data. Incorporating audio or video improves performance, but collecting large-scale multimodal data and pre-training multimodal LLMs is challenging. To this end, we propose a Fusion Low Rank Adaptation (FLoRA) technique that efficiently adapts a pre-trained unimodal LLM to consume new, previously unseen modalities via low rank adaptation. For device-directed speech detection, using FLoRA, the multimodal LLM achieves 22% relative reduction in equal error rate (EER) over…
AI Generated Robotic Content

Recent Posts

Pushing MiniMax H3 quality on an RTX 3070 8GB — movie screenshots, voice refs + 0.5MP workflow

Wanted to see how far I could push the quality using what I already have.…

13 mins ago

Agents, Graphs, Loops & More: A Look Inside How Game of Life Is Actually Architected

I’ve spent close to a decade watching this industry build conversational AI, first through Chatbots…

18 mins ago

AI-driven development lifecycle using Amazon Bedrock AgentCore

Engineering teams adopting the AI-Driven Development Lifecycle (AI-DLC) with Amazon Bedrock AgentCore and coding agents…

18 mins ago

Wikipedia Workers Unionize for the First Time

More than 200 people in roles such as engineering, finance, and communications will now be…

1 hour ago

Why organic chemistry may help build AI that can explain its answers

While most believe artificial intelligence (AI) is changing science, researchers at the University of Notre…

1 hour ago

REFACTOR-VLA: Unsupervised Library Learning of Typed Motor Programs

Most current vision-language-action (VLA) models—such as OpenVLA, π0, RT-2, and RDT-1B—are “monolithic.” This means they…

1 day ago