Menu

Categories

Tags

Online Self-Distillation and Dreaming Could Solve AI's Forgetting Problem

June 28, 2026 | alex | AI | 218 views 0 comments

Large language models have a dirty little secret: once you close the chat, they forget everything they learned during your conversation. Current fixes like bigger context windows or faster retrieval are just band-aids. The real challenge is physically rewriting the model's weights with new knowledge.

Enter online policy self-distillation (OPSD). When an LLM faces a task, its 'teacher state' — which has access to the full conversation history — generates a high-quality answer. Then, using backpropagation in the cloud, the system compares the teacher's token-by-token predictions with the 'student' base model's, providing a dense training signal to nudge the student toward the smarter state. Unlike supervised fine-tuning, which forces the model to memorize every word, self-distillation extracts only the decision-making essence. This sparse update avoids catastrophic forgetting, preserving the model's general knowledge.

Another, more speculative approach is dreaming. When faced with a complex task, the model spends massive compute cycles running internal simulations. It builds a virtual simulator based on patterns observed in its training data, then runs thousands of practice scenarios. If a simulation succeeds, the system records the trajectory and uses it as a lesson to update the base weights. This is no lightweight compression — dreaming is computationally expensive, representing a fourth dimension of LLM scaling.

By 2027 or 2028, AI agents working alongside humans for a week might undergo a performance review. If they pass, the system could use OPSD or dreaming to distill that week's hands-on experience into its permanent weights, making the model smarter the more it's used.

Leave a Reply

Your email address will not be published. Required fields are marked *