Menu

Categories

Tags

Cerebras CS-4 claims 30x faster inference with no new chip

August 20, 2026 | Source: t | AI | 72 views 0 comments

Cerebras is back with a new AI inference server — the CS-4. But don't get too excited about the name: underneath, it's still the same WSE-3 silicon, not a next-gen WSE-4. The real upgrades are in the system itself. Each CS-4 packs three WSE-3 Turbo chips, with reworked power delivery, cooling, and chip-to-chip communication.

According to Cerebras, the CS-4 can deliver up to 30x the inference speed of conventional GPU setups, and up to 10x the per-watt throughput of the CS-3. The company doesn't specify which GPU it's comparing against, and real-world performance will vary widely by model and configuration, so treat that 30x as the official best-case number.

On a single chip, Cerebras says inference speed is now up to 2x the previous generation. The company also claims it can push past 1,000 tokens per second even on models with 10 trillion parameters. That result, however, comes from internal testing extrapolation — not an actual 10-trillion-parameter model being put through its paces.

The first CS-4 units ship in Q3 this year, with a new chip and server generation planned for 2027.

Leave a Reply

Your email address will not be published. Required fields are marked *