Zhipu AI has quietly dropped the technical report for GLM-5V-Turbo, its first multimodal programming foundation model. The model has been available on Z.ai API and OpenRouter since early April — this paper is a belated methodology reveal. It’s not open source, but it does support a 200K context window and can plug into agent frameworks like Claude Code and OpenClaw.
What makes GLM-5V-Turbo different from most multimodal models is that it doesn’t treat vision as an accessory bolted onto a language model. Instead, visual perception is baked into the entire pipeline from pre-training — covering reasoning, planning, tool calling, and execution.
The architecture hinges on three key design choices. First, a new visual encoder called CogViT, which uses dual-teacher distillation with SigLIP2 and DINOv3, then aligns via contrastive learning on 8 billion bilingual image-text samples. Second, multimodal multi-token prediction (MMTP), which replaces direct visual embeddings with a shared learnable <|image|> special token, cutting down cross-pipeline communication complexity and stabilizing training. Third, joint reinforcement learning across more than 30 tasks spanning perception, reasoning, and agent execution.
The RL improvements are widespread: 2D image localization +4.8%, video understanding +5.6%, 3D localization +7.7%, OCR +4.2%, chart understanding +7.7%, GUI agent (OSWorld) +4.9%, and multimodal search tool calling +3.5%. The team notes that multi-task RL doesn’t cause the cross-domain interference common in supervised fine-tuning — abilities rise together, and reasoning patterns learned in one domain can transfer to others.
And the benchmark scores are punchy. On Design2Code, GLM-5V-Turbo scores 94.8 — beating Claude Opus’s 4.6 (the benchmark measures code generation from screenshots). It also hits OSWorld 62.3, AndroidWorld 75.7, multimodal search MMSearch 72.9, and BrowseComp-VL 51.9. On text-only programming tasks (CC-Bench-V2), it outperforms its own text-only base model GLM-5-Turbo across backend (22.8), frontend (68.4), and code repository exploration (72.2). MMSearch-Plus scores 30.0 — nearly an 8x improvement over the previous GLM-4.6V — and the custom visual deep search benchmark ImageMining hits 30.7.