
Google has launched a new experimental tool called Stax. It is designed to make testing large language model (LLM)-powered applications more reliable. The feature is built by Google DeepMind and Google Labs to help developers move past the trial-and-error “vibe testing” approach that has long been a pain point in prompt engineering.
Unlike traditional software, AI models are non-deterministic; they don’t always return the same answer for the same input. That makes evaluation tricky, often requiring developers to manually compare outputs or build their own testing pipelines. Stax is designed to solve that problem by offering both human evaluation and LLM-as-a-judge methods, with support for custom testing criteria.