You Cannot Review a Prompt Diff

How we turned human approval history into a CI gate for agent prompt changes.

September 17, 2026 · 7 min read · Rohan Prashanth

We run agent loops that can propose real business changes. Budget moves, launch decisions, and campaign actions all flow through those loops. A small prompt edit can nudge all of that behavior, even when the diff looks harmless. We learned this in production, and it changed how we ship prompts.

What broke first

A few weeks ago we edited one section of a system prompt to make an agent more decisive. The change was short, and it read cleanly in code review. The next day our operators started seeing a cluster of proposals they had already been turning down. Nothing crashed. Tests passed. Latency was normal. The only signal was that humans trusted fewer suggestions.

That was the bug. Behavior regressed with no error. We had treated prompts like copy changes when they were actually control-surface changes.

Why diff review fails for prompts

We can review typed code because the semantics are explicit and deterministic. Prompts do not work like that. Their behavior lives inside the model. A one-line wording change can shift which actions get proposed under pressure, while leaving every schema contract untouched.

Traditional unit tests still matter, but they mostly tell us whether output is well-formed. They do not tell us whether judgment drifted. Offline benchmark scores do not solve this either, because they are rarely aligned with the exact decisions we need for our own brands.

The artifact we can diff is text. The artifact we deploy is behavior.
Prompt v1prioritize caution under uncertaintyescalate risky changes for reviewPrompt v2prioritize decisive actions under uncertaintyescalate risky changes for reviewone phraseOutput distribution from v1safe_reordermonitor_onlyholdOutput distribution from v2launch_campaignsafe_reordermonitor_only
Figure 1. A tiny wording change can keep the diff almost identical while shifting which actions the model proposes.

The labeled dataset was already in the product

Every write action from our agents already goes through an operator queue. Each item is approved or rejected by someone accountable for the result. Over time, that queue becomes a grounded preference dataset for our actual operating environment.

Once we accepted that, the path was straightforward. We did not need to invent a new evaluation protocol from scratch. We needed to use our own approval history as a regression signal whenever prompts changed.

How the replay gate works

We persist the full input context for each cycle. On a prompt change, CI replays a recent window of those cycles through the candidate prompt in read-only mode. The replay emits proposals but cannot execute them. We then compare those proposals against historical approval patterns by action type.

If too much of the candidate output concentrates in action categories that operators have been rejecting, the gate fails. If mandatory output contracts are missing, the gate fails. Prompt edits, model-routing changes, and playbook changes all run through the same guard.

replay --all-brands --recent-cycles 5 --max-warning-count 0 --max-flag-rate 0.20

# candidate prompt output (illustrative)
{
  "action_type": "launch_campaign",
  "expected_outcome": "increase qualified checkouts over the next 7 days",
  "status": "proposed"
}
Stored cycle inputsstate + context + prompt textCandidate promptpull request versionRead-only replay runemits proposals onlyno execution pathOperator ledgerapprove/reject by typeScore joinwarnings + flag-ratePASSFAIL
Figure 2. Candidate prompts are graded against historical operator decisions before merge.

The parts that were not obvious at first

Replaying from reconstructed prompts sounds fine until you measure it. Reconstruction silently drops context. We now persist verbatim cycle inputs and can require stored inputs only for strict replay runs.

Scoring by textual similarity is also a dead end. We score by action type and operator outcome because that maps to business risk. We also enforce a minimum label count per action class before trusting rejection rates, so one noisy day does not become policy.

Another lesson was model routing discipline. If a model split lingers during replay, we cannot separate prompt impact from model impact. Stable routing during replay made the signal much cleaner.

What this catches, and what it does not

This gate catches exactly the class of regression that bit us first. An agent starts reaching for levers operators regularly refuse, despite clean syntax and valid JSON. It also catches missing intent fields that break later grading and analysis.

It does not prove truth. Approval is a proxy. Humans can be inconsistent, and a genuinely novel good strategy can look suspicious if history is conservative. We are working toward outcome-linked grading so we can compare proposals to observed business results, not only to immediate human preference.

Why we like this direction

The replay gate makes prompt work feel like engineering work again. It gives us fast iteration with a concrete failure mode. Most importantly, it ties quality control to how the system is actually run day to day.

We still debate thresholds, windows, and action bucketing. That is part of the job. But the core principle has held up: if your product already creates accountable decisions, you already have the raw material for regression testing model judgment.

If this kind of problem sounds exciting to you, email rohan@spaceai.so.

Further reading

Dials is the first brand this operator runs. See how India travels with Dials, or go back to the Space homepage.