Categories: FAANG

SlowFast-LLaVA-1.5: A Family of Token-Efficient Video Large Language Models for Long-Form Video Understanding

We introduce SlowFast-LLaVA-1.5 (abbreviated as SF-LLaVA-1.5), a family of video large language models (LLMs) offering a token-efficient solution for long-form video understanding. We incorporate the two-stream SlowFast mechanism into a streamlined training pipeline, and perform joint video-image training on a carefully curated data mixture of only publicly available datasets. Our primary focus is on highly efficient model scales (1B and 3B), demonstrating that even relatively small Video LLMs can achieve state-of-the-art performance on video understanding, meeting the demand for…
AI Generated Robotic Content

Recent Posts

Pushing MiniMax H3 quality on an RTX 3070 8GB — movie screenshots, voice refs + 0.5MP workflow

Wanted to see how far I could push the quality using what I already have.…

56 mins ago

Agents, Graphs, Loops & More: A Look Inside How Game of Life Is Actually Architected

I’ve spent close to a decade watching this industry build conversational AI, first through Chatbots…

1 hour ago

AI-driven development lifecycle using Amazon Bedrock AgentCore

Engineering teams adopting the AI-Driven Development Lifecycle (AI-DLC) with Amazon Bedrock AgentCore and coding agents…

1 hour ago

Wikipedia Workers Unionize for the First Time

More than 200 people in roles such as engineering, finance, and communications will now be…

2 hours ago

Why organic chemistry may help build AI that can explain its answers

While most believe artificial intelligence (AI) is changing science, researchers at the University of Notre…

2 hours ago

REFACTOR-VLA: Unsupervised Library Learning of Typed Motor Programs

Most current vision-language-action (VLA) models—such as OpenVLA, π0, RT-2, and RDT-1B—are “monolithic.” This means they…

1 day ago