Datadog just dropped Toto 2, an open-source family of time series models that finally proves scaling laws apply outside language. Ranging from 4 million to 2.5 billion parameters, Toto 2 is the first foundation model family to show that simply making the model bigger consistently improves forecast accuracy — and it hasn't hit a ceiling yet at 2.5B. Until now, the time series world lacked that clean scaling property that makes large language models so powerful.
The family includes five sizes — 4m, 22m, 313m, 1B, and 2.5B — all released under Apache 2.0. They top the leaderboards on three major forecasting benchmarks: BOOM, GIFT-Eval, and TIME. But accuracy isn't the only win. By introducing a continuous patch masking mechanism, Toto 2 swaps autoregressive generation for a single forward pass, slashing inference latency. The 313m version is about as fast as the 120m-parameter Chronos-2.
Cross-domain generalization is another highlight. Toto 2 was pretrained solely on system monitoring metrics and synthetic data — no public time series datasets. Yet it still dominates general-purpose forecasting benchmarks. And it's more parameter-efficient: the 22m model, with just one-seventh the parameters, beats the original Toto 1.0 across every core test.