Google has baked computer use directly into its flagship model, Gemini 3.5 Flash, as a built-in tool.
Previously, developers had to call a separate Gemini 2.5 computer use model to perform agent tasks. Now, they and enterprise users can have the main model control devices directly through the Gemini API or Google Cloud’s Gemini Enterprise Agent Platform (formerly Vertex AI). This simplifies the agent development stack.
The built-in computer use tool works by receiving screenshots from browsers, mobile, or desktop environments, then visually perceiving and reasoning about the next steps. It outputs actions like mouse clicks, keyboard inputs, scrolling, and menu navigation for long-running automation tasks such as continuous software testing or cross-page web data collection. For debugging and audit purposes, each generated instruction includes an "intent" field explaining the logic behind the step.
To address the risk of prompt injection agents may encounter in the wild, Google trained the model with adversarial examples and offers two optional safeguards: human-in-the-loop approval for irreversible actions like payments or file deletion, and automatic task halting if injected commands are detected in screenshots.
Currently, Browserbase offers an online hosted demo at gemini.browserbase.com, and Google has open-sourced a reference implementation called computer-use-preview on GitHub.