OpenRouter: Open-source AI models are now just 3–6 months behind frontier
API aggregator OpenRouter reports that the performance gap between open-source models and the frontier closed-source models has stabilized at just 3 to 6 months. Over the past 18 months, leading closed-source labs haven't managed to widen the gap as expected, while open-source players from China and the US are accelerating the replacement of closed-source models with extremely cost-effective alternatives.
DeepSeek V4 Flash, released just two months ago, has become the go-to alternative. With 284 billion parameters, it scores 79.0% on SWE-bench Verified, approaching GPT-5.5-level performance. Its official first-party pricing is $0.14 per million input tokens and $0.28 per million output tokens — output costs are about 150 times cheaper than GPT-5.5. Even with Western cloud hosting premiums that don't retain training data, the effective cost is roughly 1.3% of closed frontier models.
Beyond price, Zhipu AI's GLM 5.2, released in June 2026, tops the Artificial Analysis open-weight intelligence index and rivals GPT-5.5 in real-world agent evaluations, making it a viable alternative for long-range programming planning. However, GLM 5.2 consumes more tokens during deep reasoning, so enterprises need to balance output costs. The multimodal open-source model MiniMax M3, with its innovative MSA sparse attention architecture, offers native long-context processing for images and video at low token prices, emerging as a strong open-source competitor to Gemini Flash.
Meanwhile, Nvidia's Nemotron 3 Ultra, based on the Mamba-2 hybrid architecture, has become the strongest US-based open-source model, aiming to drive demand for Nvidia hardware and microservices through an open ecosystem.
OpenRouter emphasizes that while frontier closed models will continue to advance, token costs at fixed intelligence levels will keep dropping, offering enterprises significant cost optimization opportunities.