Odyssey uses RL to make world models hunt for their own training flaws World model company Odyssey has released PROWL (Prioritized Regret-Driven Optimization for World Model Learning), an adversarial training framework powered by reinforcement learning. The core idea: treat a game environment as a training ground, where a behavior-constrained RL agent actively hunts for the world model's failures in geometry, motion, visual consistency, and action response — then feeds those failure trajectories back into the model for retraining.
The key is creating a scalable feedback loop. PROWL includes a Prioritized Adversarial Trajectory buffer (PAT) that automatically lowers the priority of simple failures once the model learns to handle them, pushing harder unresolved trajectories to the front. The stronger the model gets, the deeper the RL agent has to dig for bugs — a ratcheting spiral. Instead of passively piling up demonstration data, PROWL emphasizes actively generating high-value failure samples.
The team validated PROWL in Minecraft's MineRL environment. According to the paper, on 300 held-out human-played episodes, PROWL reduced action-following error (AFS-EPE) by 12.6% compared to a pretrained baseline, and that improvement expanded to 20.9% on the hardest 10% of episodes. Specific improvements include: fewer cases of predicting the wrong direction or ignoring controls; elimination of visual artifacts like rotation seams and color stripes; a stable crosshair during camera movement; and even stable tracking of extreme out-of-distribution actions like the RL agent's discovered 180-degree whip-turn.
Odyssey was co-founded by former Cruise VP of Product Oliver Cameron (CEO) and former Wayve VP of Technology Jeff Hawke (CTO). In February 2026, the company announced funding from NVIDIA's venture arm NVentures and Samsung Next, joining existing investors GV, EQT, Air Street Capital, and others. The company previously released the Odyssey-2 series of world models; the paper is publicly available.