Menu

Categories

Tags

Sakana AI and Nvidia open

May 10, 2026 | Source: sakana | Anthropic, OpenAI | 124 views 0 comments

Sakana AI and Nvidia have open-sourced a new sparse data format called TwELL, along with an acceleration kernel that lets GPUs skip over near-zero calculations when running large language models. The result? H100 inference speeds up by as much as 30%, training gets a 24% boost, and peak memory usage drops — all without sacrificing accuracy.

The feed-forward layers in large language models gobble up most of the parameters and compute. But the dirty secret? More than 80% of those neurons are basically asleep on every token — their activations are near zero, doing nothing for the output. Skip them and you save a ton of compute. The catch is that modern GPUs are built for dense, uniform matrix math. Picking out the sparse, scattered useful data using traditional methods costs more in data movement than you'd save.

TwELL breaks that curse by designing with GPU parallelism in mind. Instead of gathering non-zero data from across the chip, it chops data into small tiles that GPUs handle natively. Each compute core can then pack up useful data locally, eliminating expensive global memory reads and writes. It slides right into the accelerator pipeline.

In tests on a 1.5 billion-parameter model, just a light touch of regularization during training pushed the fraction of neurons that actually need computing below 2% — and performance on seven downstream tasks didn't budge. The data also revealed a pattern: the bigger the model, the more slumbering neurons. A 2 billion-parameter model had 38% fewer non-zero activations than a 500 million-parameter one. That means this hardware-level trick will pay even bigger dividends as models continue to scale.

Leave a Reply

Your email address will not be published. Required fields are marked *