Microsoft has open-sourced Phi-Ground, a family of models designed to solve one of AI's most deceptively hard problems: knowing exactly where to click on a screen. Give it a screenshot and a natural language instruction, and it outputs precise click coordinates. The open-source 4-billion-parameter version, paired with a larger model for instruction planning, beat OpenAI Operator and Claude Computer Use in click accuracy on the Showdown benchmark — and took first place across all sub-10-billion-parameter models on five evaluations, including ScreenSpot-Pro.
The team validated their approach on over 40 million data points and discovered that three training tricks commonly reported in academic papers all fall apart at scale. What actually works is far simpler: treat coordinates as plain numbers — for example, "523, 417." Previous papers invented a specialized vocabulary of positional tokens, hoping the model would learn to speak coordinates like words, but during large-scale training, those new tokens never stuck, and often caused the model to collapse. Another key finding: feed the instruction before the image. Large models read information in one direction — if the model sees "click the blue settings icon" before the pixels, it already knows what to look for. Show it the screenshot first, and it has to blindly scan, producing much worse results.
The team also discovered that reinforcement learning (RL) benefits purely visual tasks. They ran multiple click predictions on the same image, then trained the model by contrasting correct and incorrect outputs — a type of RL known as direct preference optimization (DPO). Even after full fine-tuning, this step significantly improved accuracy. RL is normally reserved for reasoning-heavy language tasks, so seeing it work on a pure perception task — "look and click" — was an unexpected bonus. To handle tiny buttons on 4K screens (a single button can occupy as little as 0.07% of the screen), the team scaled down screenshots and pasted them onto a large white canvas during training, simulating the extreme smallness of UI elements. This trick proved especially effective on complex professional software like Photoshop.