Large language models have a persistent problem: once deployed, they can't keep learning new knowledge. Current optimization tricks mostly focus on expanding context windows and speeding up retrieval, but that only lets the model temporarily look things up within a single conversation. Close the chat, and it forgets everything. The real bottleneck for continuous learning isn't retrieval speed — it's how to physically rewrite the experiences from a conversation into the model's underlying weight parameters.
Online policy self-distillation (OPSD) offers a fresh path for updating weights. When an LLM tackles a task, its "teacher state" — which has access to the full long context — generates a high-quality solution. Then, the system uses backpropagation in the cloud to compute the token-level probability difference between the base state (the student) and the teacher state, providing a dense supervisory signal that pushes the base model toward that smarter, high-scoring state.
Compared to supervised fine-tuning (SFT), which forces the model to memorize all dialogue text, self-distillation only extracts the decision-making experience needed to maintain performance. This extremely sparse parameter update avoids catastrophic forgetting, preserving the LLM's original general knowledge.
Another more forward-looking learning path is "dreaming." When facing complex tasks, the LLM consumes massive inference compute to play out internal self-play scenarios. Based on patterns it observes daily, the model automatically builds a virtual simulator environment and runs tens of thousands of task rehearsals inside it. If a rehearsal succeeds, the system records the successful trajectory as a lesson and updates the base model's weights. Unlike lightweight compression that only generates short summaries, dreaming consumes huge cloud compute for repeated rehearsals — it's the fourth dimension of LLM scaling.
By 2027 or 2028, AI agents are expected to get a weekly performance review after working alongside humans. Once approved, the system can use OPSD or dreaming to distill that week's real-world experience into the model's underlying weights, enabling online capability expansion after deployment — making the model smarter the more it's used.