Zyphra has open-sourced the preview version of ZAYA1-74B, a Mixture-of-Experts (MoE) reasoning model that was trained end-to-end on AMD MI300x accelerators. The model boasts 74 billion total parameters, but only 4 billion are activated per forward pass — a hallmark of MoE efficiency. Weights are available on Hugging Face under the Apache 2.0 license.
To optimize long-context performance, the model replaces alternating global attention layers with sliding window attention (SWA) using a 4K window. According to Zyphra's benchmarks, this design cuts KV cache memory usage nearly in half without sacrificing long-text accuracy. The training diet consisted of 15 trillion tokens of pretraining data, followed by a 3 trillion token midtraining phase that extended the context window to 256K.
Because this release hasn't undergone RL or instruction tuning, Zyphra is sharing pass@4 scores — the probability of generating a correct answer in four attempts — as evidence that the base model already has solid reasoning foundations. The fully tuned ZAYA1-74B is expected in the coming weeks.