
Google has rolled out a new feature called “conversational image segmentation” to interact with images using natural language inside its Gemini 2.5 model. Unlike earlier methods that required specific labels or bounding boxes, this feature allows you to describe parts of an image in your own words, and Gemini will understand what to highlight.
Instead of just saying “a car,” you can now prompt the model with phrases like “the car that is farthest away” or “the person holding the umbrella.” Gemini can now interpret these instructions and return a segmented result, understanding relationships, ordering, comparative attributes, conditions, and even abstract ideas like “damage” or “a mess.”