Running a system based on a large language model does not end when it is made available or its initial accuracy is improved. The system may remain fully available while silently approving incorrect decisions, consuming the budget rapidly, or switching during outages to a model whose quality has not been tested. Therefore, the fifth installment of the Running LLM systems in production series focuses on day-to-day operability as the layer that makes the system observable, controllable, stoppable, and recoverable.
Monitor the Decision, Not Just the Service
Traditional service metrics, such as request and error rates, response time, and CPU consumption, do not reveal whether the agent is making good decisions. The source proposes adding a decision-specific monitoring layer that includes decision volume, the automated execution rate, the distribution of composite confidence, the latency of each node, safety-control block events, token consumption and cost, the rate of human intervention, and the results of hidden evaluations on samples of production traffic.
The source identifies four signal families: decision, safety, cost and performance, and quality. If it is not possible to measure everything, priority should be given to the automated execution rate, safety-control block events, total model cost, and the rate at which humans override system decisions.
It also recommends dividing records into three layers: high-volume operational data, decision metadata within confidence boundaries, and raw or structured data subject to encryption and strict access control. Every decision should carry a stable identifier, decision_id, linking metrics, logs, traces, and the audit log, so that a single decision can be investigated from beginning to end.
Make Cost Measurable and Controllable
Model calls should not be distributed across different parts of the codebase, because this makes cost attribution and control difficult. Instead, the source proposes a single gateway through which all calls pass, responsible for calculating tokens and cost by tenant, capability, and model, in addition to enforcing rate and budget limits.
Cost-reduction measures come in the following order: do not call the model when a deterministic rule is sufficient; batch items into a single call; cache deterministic results; choose a smaller model for simple tasks; and then reduce unnecessary context and instructions. Estimated tokens should also be calculated before the call, rather than counting every request as a single unit, with a total platform cap and limits on retries and unbounded tool loops.
Routing, Fallback, and the Kill Switch
In practice, there is no single model for every task. Low-risk, high-volume tasks can be routed to a smaller model, ambiguous or sensitive cases can be assigned to a more capable model, and a different model family can be used as the judge model so that the models do not share the same weaknesses. Every model present in the routing or fallback path must be evaluated, because switching to an alternative model may change quality.
During a provider outage, a defined fallback chain should be used with a single overall deadline, a circuit breaker that prevents calls to a provider known to be down, and idempotency keys to prevent a decision from being processed twice if the call times out after actually succeeding. When the alternative is unacceptable, the decision should be handed over to a human rather than concealing the degradation.
As for the kill switch, the source proposes that it operate as a shared operational state that can be global or specific to a tenant or capability. The states include normal operation, HUMAN_ONLY to stop automated execution while keeping suggestions available to humans, and HALTED to stop decision-making entirely. If the switch cannot be reached, the safe behavior is to move to human-review mode, not to continue normal operation.
The Architecture That Prevents Chaos
The source connects operability to a hexagonal architecture that separates domain logic from model-provider packages through an interface and adapter layers. This makes it possible to switch providers without rewriting business logic and to test the domain using a mock model without a network or cost.
The core architecture includes stateless agent services, a model gateway, an append-only audit store, shared storage for limits and the kill switch, a secrets store, and object storage for inputs and sensitive data. The source also emphasizes an independent identity for each service, least privilege, and verification that the tenant identity matches the request at every hop, particularly when using short-lived delegation grants for deferred tasks.
Why Does This News Matter?
The practical value here is not in adding a new component, but in bringing critical control points together in one gateway: the gateway that makes model switching possible is also the one that measures cost, enforces limits, manages routing and fallback, and activates the kill switch. The success of this approach remains conditional on actually testing failure scenarios, such as a model-provider outage, activation of the kill switch, a sudden load spike, and restarting a node during an in-progress decision. Without these tests, recovery mechanisms remain unverified assumptions.