AI interpretability startup Goodfire has released VPD (Virtual Parameter Decomposition), a method that directly decomposes a language model's weight parameters into roughly 10,000 independently understandable and editable subcomponents. Unlike the dominant approach—sparse autoencoders (SAEs), which train an extra decoder to interpret intermediate signals during a model's runtime—VPD works by dissecting the model's own "source code": the computational logic frozen into its parameters.
VPD's core idea is to break each weight matrix into a set of rank-1 matrices (the simplest matrix form) that sum back to the original. Then it trains an auxiliary network to determine, for any given input, which subcomponents are causally necessary and which can be removed without affecting the output. To prevent the auxiliary network from making mistakes, the training process actively searches for counterexamples that would overturn its judgments.
This method cracks a long-standing bottleneck in interpretability: decomposing attention layers. Previous approaches either avoided attention layers entirely or could only analyze within a single attention head. VPD can extract distributed attention algorithms across heads. In experiments, it directly found two interpretable attention patterns from the weights: "previous token" and "syntactic boundary routing."
The team also demonstrated precise editing capabilities: modifying a single subcomponent can change a model's specific behavior with almost no impact on other abilities. Goodfire previously launched a commercial interpretability platform called Silico, and VPD is the underlying technical approach powering that platform.