Menu

Categories

Tags

Datadog's Toto 2 proves scaling laws work for time series forecasting

May 16, 2026 | Source: datadoghq | AI | 137 views 0 comments

Datadog just dropped Toto 2, an open-source family of time series models that finally proves scaling laws apply outside language. Ranging from 4 million to 2.5 billion parameters, Toto 2 is the first foundation model family to show that simply making the model bigger consistently improves forecast accuracy — and it hasn't hit a ceiling yet at 2.5B. Until now, the time series world lacked that clean scaling property that makes large language models so powerful.

The family includes five sizes — 4m, 22m, 313m, 1B, and 2.5B — all released under Apache 2.0. They top the leaderboards on three major forecasting benchmarks: BOOM, GIFT-Eval, and TIME. But accuracy isn't the only win. By introducing a continuous patch masking mechanism, Toto 2 swaps autoregressive generation for a single forward pass, slashing inference latency. The 313m version is about as fast as the 120m-parameter Chronos-2.

Cross-domain generalization is another highlight. Toto 2 was pretrained solely on system monitoring metrics and synthetic data — no public time series datasets. Yet it still dominates general-purpose forecasting benchmarks. And it's more parameter-efficient: the 22m model, with just one-seventh the parameters, beats the original Toto 1.0 across every core test.

Leave a Reply

Your email address will not be published. Required fields are marked *