You'd think a caching system would save you money, but on Zhipu AI's overseas developer platform Z.ai, it did the opposite. Developer Da7em reported on X that when using the platform's GLM model, repeated context wasn't triggering the discounted cached rate — instead, the system charged for fresh tokens every time.
https://twitter.com/Da7_Tech/status/2065785537714098318
Screenshots shared by Da7em show that the actual token usage was around 270,000, but the final bill clocked in at nearly 5 million — a markup of almost 19x. To rule out client-side issues, Da7em tested multiple agent frameworks including Claude Code, Hermes Agent, Zcode, and OpenCode, all of which showed the same anomaly. The conclusion: the server-side caching and billing system is broken.