
Google has launched the Gemini 2.5 Computer Use model, a new system that allows AI agents to interact directly with computer interfaces, much like a human would. It is built on the company’s Gemini 2.5 Pro foundation. The specialized model combines visual understanding and reasoning skills to perform on-screen tasks like clicking buttons, filling forms, and navigating apps or websites.
The new model is available to developers via the Gemini API in Google AI Studio and Vertex AI. Google says it outperforms leading alternatives on multiple web and mobile control benchmarks, all while offering lower latency.