Menu

Categories

Tags

Zyphra Open-Sources ZAYA1-74B, a 74B MoE Model Trained on AMD MI300x

May 9, 2026 | Source: zyphra | AI, AMD | 154 views 0 comments

Zyphra has open-sourced the preview version of ZAYA1-74B, a Mixture-of-Experts (MoE) reasoning model that was trained end-to-end on AMD MI300x accelerators. The model boasts 74 billion total parameters, but only 4 billion are activated per forward pass — a hallmark of MoE efficiency. Weights are available on Hugging Face under the Apache 2.0 license.

To optimize long-context performance, the model replaces alternating global attention layers with sliding window attention (SWA) using a 4K window. According to Zyphra's benchmarks, this design cuts KV cache memory usage nearly in half without sacrificing long-text accuracy. The training diet consisted of 15 trillion tokens of pretraining data, followed by a 3 trillion token midtraining phase that extended the context window to 256K.

Because this release hasn't undergone RL or instruction tuning, Zyphra is sharing pass@4 scores — the probability of generating a correct answer in four attempts — as evidence that the base model already has solid reasoning foundations. The fully tuned ZAYA1-74B is expected in the coming weeks.

Tags: #Github

Leave a Reply

Your email address will not be published. Required fields are marked *