| | Page: https://ernie-research.github.io/NAVA/ NAVA is a 6.3 B-parameter joint audio-video generator that synthesizes synchronized video and audio from a single prompt — including multi-speaker speech with reference-timbre control and image-conditioned continuations. Instead of post-hoc-aligned dual towers or fully unified tri-modal stacks, NAVA uses an Align-then-Fuse MMDiT: a dedicated alignment space first establishes audio-video correspondence, then context (text, speaker embeddings) is fused via cross-attention. On Verse-Bench it sets new SOTA on Sync-C / Sync-D / video quality / audio WER while using 2× to 5× fewer parameters than open-source baselines. submitted by /u/AgeNo5351 |
Hi everyone! When I saw the new trailer for Kirby & The World Beyond I…
Quick note — The Loop’s first session is today, 4:30 PM PDT, live on Zoom.Free, monthly, and genuinely…
When you build an application on top of a large language model (LLM), the prompt…
AI leaders worry antitrust law could stand in the way of what they view as…
Researchers have developed a learning mechanism that uses the natural variability of neural activity—often dismissed…
On August 12, 2026, Alibaba’s Qwen team released Qwen3.8-2.4T-A95B. This is the first time a…