Menu

Categories

Tags

More memory, worse results: AI over-summarizing backfires in new study

May 12, 2026 | Source: github | AI, Developer | 149 views 0 comments

A set of agent memory experiments by Dylan Zhang, a PhD student in computer science at the University of Illinois, points to a counterintuitive conclusion: making models repeatedly summarize their experiences can actually make them worse.

The most striking result comes from ARC-AGI. The researchers picked 19 problems that GPT-5.4 could solve perfectly without any memory, then fed the model the correct solutions and asked it to write "experience summaries" as it went along. In theory, this is like an open-book review. But after multiple rounds of memory compression, the same model's accuracy dropped from 100% to 54%. The original trajectories weren't the problem — the damage happened when the model rewrote correct trajectories into generic experience.

This memory degradation isn't an isolated incident. In the WebShop online shopping task, using the AWM memory method, the model scored 0.64 after ingesting 8 expert trajectories. When the number of trajectories was increased to 128, the score fell to 0.20 — right back to the no-memory baseline. In other words, the more memory you pile on, the more its benefits cancel out.

The issue isn't "not enough experience" — it's "summarizing too much." When a large language model writes down experience, it's not an objective log; each summary is a regeneration. Eventually, specific conditions get deleted, rules from different tasks get blended together, and the kind of detail that could actually guide behavior turns into useless platitudes like "prioritize the most direct action" or "use the right tool." The original paper shows one extreme example: 50 structured memories were merged into a single summary, collapsing multiple task variations into one generic workflow. The next evaluation round promptly lost 6 to 13 successful samples.

The authors offer a restrained recommendation: don't rush to have your agent write a "mistake notebook" after every round. A safer approach is to keep filtered raw action trajectories and only abstract when necessary. In their experiments, simply preserving original episodes and turning off abstract summarization matched or beat tested compressed memory methods across multiple agent benchmarks. For developers, the lesson is straightforward: showing a model what it actually did is usually more useful than making it memorize a bunch of abstract rules.

Leave a Reply

Your email address will not be published. Required fields are marked *