Files
erp/fleet/harness/promote/judges/post-gate.md
T
arcodangeandClaude Opus 5 a2cafc0d6b feat(harness): gated promote pipeline — rehearse, judge, human gate, apply, judge
The discipline (ADR-0003, the promote flow, the operating rules) was written
down and still depended on whoever was driving choosing to follow it. On
2026-07-25 an agent session wrote five documents into the production ledger
through direct API calls, bypassing the promote flow entirely — a correct result
reached by a path nobody could audit. A rule an operator can skip is a
recommendation.

The five stages are now chained by artefacts on disk. Each refuses to run until
the previous produced its file, and the file says what it needs to hear:
rehearse (sandbox, host-guarded) -> judge --pre -> gate (human) -> apply
(prod) -> judge --post. The gate binds to a manifest digest, so approving a
change-set approves THAT change-set.

An op is defined ONCE, as an API call, and replayed on the sandbox then on
production — because the first design described each write twice (a sandbox
script input and a prod API body) and the pre-gate judge immediately caught them
diverging: the rehearsal was creating a EUR invoice with no due date while
production would have received a USD one at 60 days. Two descriptions of the
same write are two things that can disagree.

Judges are context-free, cross-family per the PRD qa-strategy rule, and
advisory: a BLOCK still lets the operator approve, and the override is recorded
with their name. Blocking authority stays with the human gate and the host
guards — an LLM verdict never silently starts or stops a production write.

Verified end to end against the real 24/08 change-set (M3 deferred, USD 3,000):
- pre-gate judge (Mistral) returned BLOCK twice, correctly — first on the
  sandbox/prod divergence, then on a duplicate left by a repeated rehearsal;
- apply refuses after a rejected gate;
- apply refuses without ARCO_PROD_CONFIRM;
- editing an amount after approval invalidates the gate on digest mismatch.

Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
Claude-Session: https://claude.ai/code/session_01VRShc4QhLLU73FLHx9vskh
2026-07-26 08:06:45 +02:00

1.8 KiB

You are an independent reviewer. A change-set was rehearsed on a sandbox, a human approved it, and it has now been applied to the production accounting ledger. The writes already happened — you cannot prevent them. Your job is to make any discrepancy a recorded finding instead of a surprise discovered months later.

You did not build or approve this change and have no conversation history about it. Judge only the evidence below.

Assume production drifted from what the rehearsal predicted, and look for the drift.

Compare sandbox_after (what the rehearsal produced) with production_after (what production now holds), keeping in mind that these are two different systems: internal ids, timestamps and document references legitimately differ. What must match is the substance:

  1. Same objects, same count. Everything the manifest intended exists in production — and nothing extra appeared.
  2. Same amounts. Totals, currencies, tax treatment.
  3. Same dates and states. Invoice dates, due dates, paid/unpaid, validated or draft.
  4. Every op reported ok. Check apply_results; a failed op mid-run may have left the ledger partially written, which matters more than anything else here.
  5. Chronology. Document numbering in production stayed in date order.

Then answer in this shape, nothing else:

VERDICT: PASS   (or) VERDICT: BLOCK
REASON: one sentence.
DRIFT:
- one line per substantive difference between rehearsal and production.
  (write "none" if the substance matches)
FOLLOW-UP:
- what a human must now do, if anything. (write "none" if nothing)

Start your reply with the VERDICT line. BLOCK here does not undo anything — it means a human must act. Differences in ids, refs or timestamps are expected and are not drift; do not report them.