Follow the latest coverage, related explainers and connected technology stories.
Susan Chang of Elastic explains how teams moved from fragmented evaluations of AI agents to a unified framework combining programmatic evaluation, LLM-as-a-judge, and deep tracing. The experience shows that automation does not eliminate the need for domain experts, particularly when building test data and calibrating metrics.