Categories: FAANG

Video generation models as world simulators

We explore large-scale training of generative models on video data. Specifically, we train text-conditional diffusion models jointly on videos and images of variable durations, resolutions and aspect ratios. We leverage a transformer architecture that operates on spacetime patches of video and image latent codes. Our largest model, Sora, is capable of generating a minute of high fidelity video. Our results suggest that scaling video generation models is a promising path towards building general purpose simulators of the physical world.
AI Generated Robotic Content

Recent Posts

REFACTOR-VLA: Unsupervised Library Learning of Typed Motor Programs

Most current vision-language-action (VLA) models—such as OpenVLA, π0, RT-2, and RDT-1B—are “monolithic.” This means they…

4 hours ago

Accessing OpenAI models on Amazon Bedrock from Australia with global cross-Region inference

Australian teams working with OpenAI models can now access the latest OpenAI models through Amazon…

4 hours ago

Getting started with Mantis, our open-source bug finding-and-fixing harness

AI models have clearly proven their ability to discover and exploit vulnerabilities without much, if…

4 hours ago

Meta Pushes Its New AI Agent on Employees—but Eases Off on Tokenmaxxing

The company is reducing pressure on workers to use artificial intelligence tools while encouraging them…

5 hours ago

Why did your robotaxi stop? New system helps predict self-driving car mistakes

Self-driving cars are often controlled by deep learning models that sometimes fail in unexpected situations.…

5 hours ago

Introducing Claude Fable 5.1 on AWS

Today, we’re excited to announce the availability of Claude Fable 5.1 on Amazon Bedrock and…

1 day ago