AI agent development teams need more than checking the final answer to determine whether a system is working as it should. In applications that combine retrieval, generation, and tool calling, an incorrect result may be caused by selecting the wrong database or retrieving irrelevant documents, even if the problem appears to be only in the final text. During a presentation delivered through InfoQ, Susan Chang, a principal data scientist at Elastic, explained how the company developed a shared framework for evaluating its agents used in cybersecurity and enterprise chatbots.
From Isolated Evaluations to Shared Tools
Elastic teams began by creating separate datasets, evaluators, and tracing processes for each agent. In the case of an attack-detection agent, the tests included attack and benign scenarios designed by analysts and security researchers, with metrics such as precision and recall, factual correctness, similarity, and MITRE tactic classification. Enterprise-data-based chatbots, meanwhile, tested questions and document retrieval, focusing on answer relevance and completeness, the correctness of product identifiers, and the formulation of ES|QL queries.
The differences between these cases led to substantial duplicated work. Elastic therefore created a shared framework capable of importing different types of datasets into a unified schema, running evaluations based on execution traces, and using shared evaluators for RAG applications, alongside components customized for each product. The process can be run locally, with data loaded, the agent run, results collected, and scores displayed to developers.
Tracing Is Essential to Understanding the Cause of Failure
The experience confirms that agent tracing should not be limited to its final output. The tools it called, vector and keyword searches, retrieved data, token consumption, response time, and sequence of decisions should be recorded. This level of detail makes it possible to evaluate a specific tool call or discover that the agent used the wrong source before the problem appears in the final answer.
Cases reported by users as failures, such as a negative evaluation, can also be turned into new examples in the test set. In this way, production feedback becomes part of regression testing in later releases, rather than remaining manual notes separate from the development cycle.
Why Is LLM-as-a-Judge Not Enough?
Elastic used language models to evaluate the outputs of other models in open-ended or ambiguous tasks, such as style, consistency, and the relevance of an answer to its context. This approach makes it possible to scale evaluation when it is difficult to write a precise rule for judging a long text. However, it may produce inconsistent results from one run to another, and it may fail to detect specific errors such as a nonexistent product identifier or an invalid query formulation.
Elastic therefore combined LLM-as-a-judge with deterministic evaluations based on rules and programming. These evaluations inspect JSON or YAML structure, syntax validity, the presence of required entities, and whether code or a query can be executed, in addition to metrics such as precision, recall, and factual correctness. This combination reduces the cost and speeds up evaluation in cases with a clear answer, while leaving semantic judgment to tasks that are difficult to reduce to rules.
What Cannot Be Generalized?
Elastic did not consider the creation of specialized test data to be fully automatable. In cybersecurity, the team needs analysts and researchers to determine what constitutes a real attack and what behavior is acceptable for the end user. The definition of regression also differs between agents; a security agent may become more inclined to declare that attacks are present even when benign data is entered after a particular update.
Calibrating evaluators also remains the responsibility of each team. If the language evaluator assigns different scores to the same case or does not agree with the intended human judgment, the shared framework will produce misleading numbers regardless of the quality of its software architecture. Chang also pointed to the risk of evaluator bias when the same model family is used for generation and judgment, as it may overestimate the performance of those models.
What Changes Practically for Teams?
Elastic's experience recommends starting with a small set of test examples, potentially ranging from 20 to 50 cases, rather than waiting for a perfect set. In the early stages, fragmented work may be acceptable to accelerate learning, particularly when use cases and metrics are still being discovered. But as multiple agents enter production, basic tracing and evaluation become essential for answering performance questions and diagnosing user problems.
Elastic also moved some evaluation tools from Python to TypeScript to match the TypeScript-written production code and use Playwright and a custom internal tool called Scout. This step does not mean that all data-science evaluations must be migrated; rather, it reflects a specific need to reduce the gap between what the team tests and what the system actually executes. The practical conclusion is that a shared framework can standardize the schema, execution, and general metrics, but it cannot replace domain expertise or judgment about what should be considered success and failure.