AI & Machine Learning

Step-Level Agent Monitoring: Halting Bad Runs at 1ms

Post-hoc log review catches agent failures after the damage is done, and a guardian agent on every step doubles your cost. A streaming trajectory monitor offers a third design, and the access regime you have determines how much it can see.

Share this article
Comments
Share:
Post-hoc log review catches agent failures after the damage is done, and a guardian agent on every step doubles your cost. A streaming trajectory monitor offers a third design, and the access regime you have determines how much it can see.
Table of Contents

If you run agents in production, you already have two ways to find out a run went wrong, and both are bad at the moment that matters. The first is post-hoc review: traces land in your observability store, an eval job or an on-call human reads them, and you learn the agent looped for forty steps or called a destructive tool with the wrong arguments after the run finished. The second is the guardian pattern: put a second model in the loop that judges every step of the first. That catches problems in flight, but it adds a model call, and its latency and cost, to every single step of every single run, including the 90-something percent that were going to be fine.

The design question for the next year of agent platforms is whether there is a third option: a monitor cheap enough to run on every step and informed enough to say “this run is off the rails” before the irreversible action, not after. A paper submitted to arXiv on October 8, OnTrack (arXiv 2610.12375, Barazandeh, Swanson, Kulkarni, Mungel), makes a concrete argument that there is, and, more usefully for architects, it spells out what you must have in place for it to work. This post treats the paper as evidence for a design pattern and walks through what it means for your serving and observability stack.

Why Both Existing Options Fail Differently

Post-hoc review fails on timing. For a read-only research agent, finding the problem late costs you tokens. For an agent that writes to a ticketing system, moves money, or edits a production config, finding it late costs you a rollback, and sometimes there is no rollback. The paper frames it plainly: agents are being used for things like trip planning, stock trading and IT incident triage with little rule-based safeguarding, and existing mitigations either review logs after the run or add a safeguard agent whose cost and latency are paid at every step.

The guardian-agent pattern fails on economics and on correlation. If your agent step costs one model call and your guardian costs another, you have roughly doubled per-step inference spend and added a serial hop to your critical path, which compounds in long loops. Worse, a guardian built on the same model family tends to share the blind spots of the agent it supervises, so the failures you most need caught are the ones it is most likely to wave through.

What both approaches share is that they treat monitoring as a semantic judgment problem. OnTrack’s key move is to treat much of it as a structural problem: does the shape of this run, the sequence of steps and the dependencies between them, look like the shape of runs that succeeded? Loops, stalls, repeated identical tool calls and plan violations are visible in structure without anyone reading the content of the messages.

What a Streaming Trajectory Monitor Actually Does

The mechanism, per the paper, is structure-aware optimal transport applied in a streaming fashion. In plain terms, the monitor keeps a running comparison between the steps the agent has taken so far, along with their dependencies, and the recorded steps of successful reference runs. Each new step updates a score. If the score crosses a threshold, the monitor can alert a human or halt the agent. The reported overhead is about 1 ms per step, which is the number that changes the economics: a check that costs a millisecond and no tokens can run on everything.

The reported results are modest and honest, which makes them more believable. On SWE-bench trajectories, using only the first 8 steps of a run, the monitor’s ranking of failing runs below successful ones improved AUROC by 0.057 over content-similarity baselines. An abort policy built on top of it saved about 18 percent of compute on runs that were going to fail, and of the runs it interrupted, 83 percent (five of six aborts) were genuinely heading to failure. Those are single-benchmark results on a coding-agent workload with a small number of aborts in the sample, so treat them as proof that the approach is viable, not as a number to put in a business case.

The Access Regime Decides What You Can Detect

The part of the paper that matters most for enterprise design is the three-tier access model. The authors study monitoring under three levels of available information: full reference access (past runs plus tool schemas), tool schemas only, and nothing beyond the step logs. Detection capability shrinks as access shrinks. That is exactly the situation most platform teams are in without realizing it.

With full reference access, you can ask whether this run diverges from how successful runs for this task type have gone, which catches plan violations and subtle drift. With only tool schemas, you can check that calls are well-formed and that dependencies between steps make sense, but you cannot say whether the plan itself is reasonable. With only step logs, you are limited to generic pathologies: loops, stalls and repeated tool calls. Those are still worth catching, since they are where a lot of wasted spend and runaway behavior lives, but it is a much weaker monitor.

The architectural consequence is that the quality of your monitor is bounded by the quality of the data your platform retains. Teams that keep traces only for debugging, with short retention, inconsistent step schemas and no outcome labels joined back to runs, are in the logs-only regime by default. Teams that keep structured, labeled trajectories per task type are in the full-reference regime and can build something much better on the same underlying idea.

Common Failure Modes

The first failure mode is reference rot. A monitor calibrated against last quarter’s successful runs will flag a legitimately new behavior after you change the prompt, the toolset or the model. If the threshold is tuned tightly, you get a flood of false aborts after every release. Reference sets need versioning tied to agent versions, and a release should begin with a shadow period where the monitor scores but does not act.

The second is cold start on rare tasks. Structural comparison needs successful examples. For long-tail task types with few reference runs, the score is noisy, and an aggressive abort policy will kill good runs. The usual answer is to fall back to the weaker regime for those task types: loop and stall detection only, with human alerting instead of automatic halting.

The third is the cost of a false abort, which is not symmetric across tasks. Halting a code-generation run that was fine wastes some compute. Halting a multi-step remediation halfway through can leave a system in a partial state that is worse than either finishing or never starting. Abort has to be wired to the same reversibility classification you use elsewhere, and for non-reversible workflows the right action on a bad score is often to pause and escalate, not kill.

The fourth is treating the monitor as a security control. A structural similarity score tells you a run looks unusual; it does not tell you a run is malicious, and an adversarial agent or injected instruction can follow a perfectly ordinary-looking sequence of steps. This is a reliability and cost layer. It complements policy enforcement on actions and does not replace it.

Architecture Impact

What changes in system design? Monitoring moves from an asynchronous consumer of traces to a synchronous participant in the agent loop. The orchestrator emits a structured step event, including tool name, argument shape and dependency on prior steps, to an in-process or sidecar scorer before executing the next action, and the scorer returns continue, alert, or halt. The trace store stops being only a debugging archive and becomes a versioned reference corpus keyed by task type and agent version, with outcome labels attached.

What new failure mode appears? Monitor-induced outages. A scorer with a stale reference set can abort healthy runs en masse after a prompt or model change, and because the abort looks like an agent failure in your dashboards, the team debugs the agent instead of the monitor. A second pattern is silent monitor blindness: if step events lose fields in a schema change, the scorer degrades to the logs-only regime without anyone noticing.

What enterprise teams should evaluate:

  • Platform and orchestration engineers: whether your orchestrator can emit a normalized step event synchronously and honor a halt signal, including cancelling in-flight tool calls cleanly.
  • ML and evaluation teams: whether successful runs are labeled and retained per task type and agent version, so a reference corpus can be built without a data-engineering project.
  • Risk and operations owners: which workflows can tolerate an automatic halt, which require pause-and-escalate, and who is on call for a monitor-triggered pause.

Cost / latency / governance / reliability implications: The reported overhead of about 1 ms per step is negligible next to a model call that typically takes hundreds of milliseconds to seconds, and it adds no tokens, which is the main difference from a guardian agent. The reported 18 percent compute savings on failing runs is only meaningful if failing runs are a large share of your spend; if 10 percent of your runs fail, that is under 2 percent of total compute, so the stronger argument is risk reduction and earlier human attention, not cost. From a governance standpoint, every halt and alert is an auditable decision with a score and a reference set version, which is considerably easier to explain to a model risk or audit function than a free-text judgment from another model.

Implementation Guide

Start with the weakest, cheapest monitor and ship it everywhere: loop, stall and repeated-tool-call detection on the raw step log. It needs no reference corpus, it catches the most embarrassing runaway behaviors, and it forces you to define the step event schema that everything else depends on. Make that schema a contract. Tool name, normalized argument shape, parent step references, timestamps, agent version and task type should be required fields, validated at emit time, with an alert if the field-completeness rate ever drops.

Next, build the reference corpus. For each task type, retain trajectories from runs your evals or downstream outcomes labeled successful, tagged with the agent version. You do not need thousands per task type to get started; you need enough for the high-volume task types, and a clear policy that rare task types stay in the weaker regime. Resist the temptation to sweep in every run as a reference. A corpus that includes quietly failed runs teaches the monitor that failure looks normal, and the fail-plausible pattern, where a run ends in a confident but wrong answer, is exactly how that contamination happens.

Run the scorer in shadow mode for at least a full release cycle before it is allowed to halt anything. Log the score at every step, join it to the eventual outcome, and compute how many runs the monitor would have aborted and how many of those were actually bad. This is the only way to set a threshold you can defend, and it is what turns the paper’s 5-of-6 precision figure into a number measured on your own workloads. Also watch the first 8 steps specifically: the paper’s early-step result suggests that the earliest part of a trajectory carries much of the signal, which is where the savings and the risk reduction live.

When you do turn on enforcement, tier it by consequence. Read-only and easily reversible workflows can auto-halt. Workflows that write to systems of record should pause and page a human with the trajectory so far and the divergence point, not kill the run. Wire the pause to your existing reversibility classification so the resume or compensate step is defined before you need it. And keep action-level policy enforcement in place underneath; the monitor is a probabilistic early warning, and the policy layer is the hard limit.

You will know it is working when three things are true: the false-abort rate on shadowed runs stays stable across releases, the median step at which bad runs are halted is early enough that the actions taken before the halt are all reversible, and on-call engineers can open a halted run and see why the monitor flagged it in under a minute. Teams that get this right over six to twelve months typically end up with per-task-type reference corpora refreshed automatically on each release, monitor thresholds treated as versioned config with their own change review, and a second signal layered on the structural one, such as cheap probes on activations or tool-risk scores, feeding the same continue, alert or halt interface. At that point the guardian model is reserved for the small slice of high-stakes steps where a semantic judgment is actually worth its cost.

Sources

Enterprise AI Architecture

Want more enterprise AI architecture breakdowns?

Subscribe to SuperML.

Comments

Sign in to leave a comment

Back to Blog

Related Posts

View All Posts »

Agent Eval Costs: Stop Runs You Can Already Predict

Full-pass agent benchmarks now cost hundreds to thousands of dollars per run. Early outcome prediction halts runs whose result is already evident, and the architecture around it decides whether your rankings survive.

Agent Inventory Is a Pipeline, Not a Spreadsheet

Most enterprises believe their AI agent inventory is complete. The data says otherwise. Here's how to build agent inventory as a continuously reconciled discovery pipeline, with a canonical agent record, risk tiering, and evidence that is ready before an examiner asks.

Your AI Agents Need Undo, Not Just a Kill Switch

Kill switches stop an agent from doing more damage; they do nothing about the damage already done. Here's how to design agent actions around reversibility, with compensating transactions, state checkpoints, and autonomy tiers set by blast radius.