Menu

Categories

Tags

Poetiq’s API harness lifts Kimi 29.9 points, lets Gemini Flash beat Claude Opus

May 16, 2026 | Source: poetiq | AI, Developer | 167 views 0 comments

A six-person startup founded by former Google and DeepMind researchers Shumeet Baluja and Ian Fischer just claimed the top spot on the LiveCodeBench Pro coding benchmark — without touching a single model weight. Poetiq’s Meta-System is an API-only harness that recursively improves by extracting task experience from responses. In tests, it turbocharged off-the-shelf models without any fine-tuning or weight access.

The biggest winner: Kimi K2.6, which saw its accuracy soar from 50.0% to 79.9% — a 29.9 percentage point gain. Even more impressive, the lightweight Gemini 3.0 Flash jumped 10 points, leapfrogging not only its bigger sibling Gemini 3.1 Pro but also — in Poetiq’s words — the “bigger and more expensive” Claude Opus 4.7 and GPT 5.2 High.

At the high end, GPT 5.5 High went from 89.6% to 93.9% with the harness, while the base Gemini 3.1 Pro scored 90.9% — beating Google’s own unreleased strong reasoning model Gemini 3 Deep Think (88.8%). Poetiq argues that traditional fine-tuning locks improvements to a single model, whereas their plug-and-play harness lets companies get better reasoning without the cost of fine-tuning and deploying a full flagship model.

Leave a Reply

Your email address will not be published. Required fields are marked *