Menu

Categories

Tags

xAI's Grok Imagine 1.5 syncs audio and video, tops Arena leaderboard

June 9, 2026 | Source: x | AI, xAI | 290 views 0 comments

xAI has released a new preview version of its image-to-video model, Grok Imagine Video 1.5, and it's already taking names. The model shot to the top of the Arena.ai video generation leaderboard, scoring an Elo of 1473 — beating out ByteDance's dreamina-seedance-2.0 in initial evaluations.

What sets Grok Imagine 1.5 apart is its ability to generate synchronized audio along with the video in a single inference pass. While other image-to-video tools usually require a separate step to add sound — dialogue, background music, or ambient effects — Grok Imagine 1.5 spits out both the video and matching audio at once. You just feed it a starting image and a natural language description, and it handles the camera moves, scene pacing, and sound design.

Under the hood, Grok Imagine 1.5 is built on xAI's own Aurora engine. Instead of the diffusion Transformer architecture used by many competing models, Aurora is an autoregressive mixture-of-experts (MoE) network that treats text, images, video, and audio as a unified stream of tokens during training. The model can generate clips up to 15 seconds long at a maximum resolution of 720p.

Leave a Reply

Your email address will not be published. Required fields are marked *