Menu

Categories

Tags

Why Your AI Agent Keeps Bailing After Minutes

May 12, 2026 | alex | Developer, OpenAI | 199 views 0 comments

Codex’s /goal mode lets an agent loop until it finishes a task, but that just magnifies the problem of vague human prompts. OpenAI engineer Chris Hayduk, drawing on internal experience, says fuzzy instructions like “optimize the code” make the model give up early because it doesn’t know when it’s done — or it spirals into aimless edits.

To keep an agent working steadily for days or longer, he lays out three rules:

  • Kill qualitative words; use a checklist instead. Models can’t judge what “better” means, but they can understand “reduce runtime by 20% without breaking any tests.” For subjective tasks like formatting a paper, Hayduk throws a Markdown checklist of 200 formatting requirements at Codex, brute‑forcing the abstract into a quantitative task: “Check every box, and you’re done.”
  • Keep validation time down to minutes. Agents need tests to verify their actions work. Don’t let them run hours on your full production environment — give them a sampled dataset and a lightweight framework. The feedback loop should be as short as possible.
  • Build three files as an external brain. Even the biggest context window will lose memory after days of running. Hayduk suggests creating three local Markdown files: PLAN.md (high‑level plan), EXPERIMENTS.md (record of wins and fails), and EXPERIMENT_NOTES.md (real‑time thought drafts). This forces the model to write its trial‑and‑error process to disk.
Tags: #Codex

Leave a Reply

Your email address will not be published. Required fields are marked *