
NVIDIA is open-sourcing Nemotron 3.5 Lightning, a model designed to handle the execution work for long-running agents. The model has 30 billion total parameters but only activates 3 billion per token, keeping it light and fast for high-frequency chores like tool calls, result checking, and sub-agent scheduling.
The idea is to make agent workloads a sharper division of labor: big models like Nemotron 3 Ultra handle complex planning, while Lightning grinds through the repetitive execution. NVIDIA says it can reach output speeds up to 4x faster than comparable models. PinchBench…