Forget the alchemy of hyperparameter tuning — OpenAI post-training team member Jiayi Weng has proposed a new reinforcement learning paradigm called "heuristic learning," and he's open-sourced all the code. The trick? Let an AI write and rewrite its own game-playing strategy, without ever training a neural network.
Here's how it works: Weng had Codex (GPT-5.4) write a Python strategy for Atari Breakout, then run a game, review the recording, spot where it missed the ball, and modify the code accordingly. After a few iterations, the strategy score climbed from 387 to a perfect 864. No neural network was trained. The AI just kept tweaking if-else rules, adjusting landing predictions, and adding infinite loop detectors. The final code is a full software system complete with ball path predictor, stuck-ball detector, regression tests, and experiment logs.
The core difference from traditional reinforcement learning is where the knowledge is stored. Traditional RL crams knowledge into neural network parameters — inscrutable to humans and prone to catastrophic forgetting when learning new tasks. Weng's approach flips that: knowledge is code. Humans can read, modify, and put tests around it. Learning something new doesn't overwrite what came before.
Weng didn't just stop at Breakout. His heuristic learning system also scored over 6,000 on MuJoCo Ant (a simulated robot ant walking challenge), and on the full Atari57 benchmark, it came close to the PPO baseline. But he's clear about the limits: pure code can't handle complex perception tasks — you can't write if-else statements to recognize images. His endgame is a hybrid architecture: lightweight neural networks for perception, heuristic learning for real-time logic and safety rules, and a large model at the top that reviews logs and rewrites code, periodically updating itself with high-quality data from the lower layers.
Hand-coded rules fell out of favor not because they didn't work, but because humans couldn't maintain them. Now that AI can write code fast and well, that old path is worth exploring again.