Artificial intelligence

How Can You Make AI Agents Predictable in Sensitive Systems?

The Stack Overflow Blog proposes confining the language model to a single nondeterministic node and leaving the rest of the system to ordinary, testable code. The methodology includes fixed flow diagrams, structured outputs, independent validation, human approvals, and an immutable log for every decision.

2026-10-07
5 min read
1 views
certi.news Editorial Team
How Can You Make AI Agents Predictable in Sensitive Systems?

In systems that handle money transfers, healthcare, or infrastructure, it is not enough for an AI agent to succeed most of the time. The central idea in the first level of a six-level maturity model for operating language-model systems in production is to contain the model within a single node, so that everything around it remains ordinary code that can be tested, reviewed, and audited.

The article presents this approach as the “determinism layer”: the system retains the model’s ability to exercise judgment on complex inputs, but prevents it from having direct authority over business state or from executing sensitive actions.

The Agent Proposes but Does Not Execute

The agent operates as a pure function that transforms context into a proposed decision. It produces the decision ID, the required capability, the proposed change, the confidence score, the routing path, the reasons and evidence, but it does not change business state. The result is passed to a separate component, which the article calls the substrate, and which applies it only after obtaining the required approval.

This separation gives the system three practical properties: the ability to test the agent without simulating the outside world, reducing the impact of incorrect outputs to a proposal that can be rejected, and preventing chains of hidden side effects between agents. The auditable question thus becomes: What did the model propose, who approved it, and what was actually applied?

A Fixed Diagram Instead of a Free-Running Loop

Rather than allowing a ReAct model to decide the next step each time, the article proposes a fixed diagram for each capability. According to the example, the request passes through nodes for decision input, input checks, and context loading, followed by a single node for language inference, then output guardrails and validation, optional adjudication, confidence aggregation, routing, proposal preparation, and memory and decision-log writing.

This structure makes the execution path known in advance, limits execution time and cost, and allows each node to be tested separately. The nondeterministic node is llm_decision, which receives a structured input and produces a structured output, while specialized code validates allowed values, schema compliance, and business rules.

Structured Outputs Do Not Mean a Correct Decision

The article warns against relying on free-form text and then attempting to extract the decision from it. The alternative is to use a JSON schema, tool calling, or grammar-constrained generation, then validate the result and retry when it does not conform, with a limited number of attempts and a fail-closed behavior rather than passing a guess to subsequent stages.

However, this procedure controls the output’s form, not the correctness of the judgment. A JSON object may be syntactically valid while still containing an incorrect decision. Therefore, business-rule validation, evaluation, and independent confidence signals must remain in subsequent layers.

Confidence, Escalation, and the Immutable Log

Confidence is assembled from the model signal, validation results, and a review by a second model when sampling sensitive decisions. The system then routes the decision to automated execution, recommends human review, requires it, or rejects the decision. The article emphasizes that the confidence score should not be a field the agent sets for itself, but rather a result calculated from independent signals, with thresholds initially set conservatively and lowered only when the data demonstrates that doing so is safe.

Every decision should also be recorded in a separate immutable log containing the model and prompt version, a summary of the decision inputs, the decision and confidence, and the routing path. Corrections do not modify previous records; instead, they are added as new records that point to the decision they replace. The article recommends storing hashes of sensitive inputs instead of raw data, while never dropping the write even under pressure, because the log is the primary reference, not merely monitoring data.

When Is an Iterative Loop Allowed?

The methodology does not reject ReAct loops entirely, but limits them to cases in which the number and order of steps depend on what the model discovers during research. Even when they are used, an explicit iteration limit, a list of tools permitted for each capability, and logging of every step must be imposed, and the output must be sent back through the same guardrails, validation, confidence, and routing. If the limit is reached before a result is obtained, the path should be human review rather than an endless loop.

Editorial reading: The real change here is not choosing a better model, but shifting the center of trust from “agent autonomy” to the engineering of the boundaries around it. This design does not automatically solve the problem of judgment correctness or confidence calibration, but it makes failure isolatable, reproducible, and reviewable. It therefore works as a foundational principle for high-consequence systems, not as a substitute for continuous evaluation or specialized validation for each capability.

News source
Stack Overflow Blog
Open original source ↗
c
Author

certi.news Editorial Team

In the same category

You may also like

View all news