Categories: FAANG

Multimodal Large Language Models with Fusion Low Rank Adaptation for Device Directed Speech Detection

Although Large Language Models (LLMs) have shown promise for human-like conversations, they are primarily pre-trained on text data. Incorporating audio or video improves performance, but collecting large-scale multimodal data and pre-training multimodal LLMs is challenging. To this end, we propose a Fusion Low Rank Adaptation (FLoRA) technique that efficiently adapts a pre-trained unimodal LLM to consume new, previously unseen modalities via low rank adaptation. For device-directed speech detection, using FLoRA, the multimodal LLM achieves 22% relative reduction in equal error rate (EER) over…
AI Generated Robotic Content

Recent Posts

Hidden prompts can plant false memories in AI agents, researchers warn

Large language models (LLMs), the computational algorithms underpinning ChatGPT, Gemini and other artificial intelligence (AI)-powered…

11 hours ago

4 Best Walking Pads for Small Spaces and Standing Desks (2026)

Our remote team clocked serious hours walking, working, and sometimes jogging to find the best…

1 day ago

Chinese AI model takes US tech industry by surprise with abilities rivaling Claude and ChatGPT

Another powerful new artificial intelligence model from China took the U.S. tech industry by surprise…

1 day ago

Agentic AI Security: Defending Against Prompt Injection and Tool Misuse

In this article, you will learn what prompt injection and tool misuse are in the…

2 days ago

When Unlearning Is Free: Leveraging Low Influence Points to Reduce Computational Costs

As concerns around data privacy in machine learning grow, the ability to unlearn—or remove—specific data…

2 days ago

In-House LLM Serving at Netflix

By AI Platform’s Model Runtime team and Inference teamIntroductionMost organizations consume LLMs through hosted APIs.…

2 days ago