The discipline (ADR-0003, the promote flow, the operating rules) was written down and still depended on whoever was driving choosing to follow it. On 2026-07-25 an agent session wrote five documents into the production ledger through direct API calls, bypassing the promote flow entirely — a correct result reached by a path nobody could audit. A rule an operator can skip is a recommendation. The five stages are now chained by artefacts on disk. Each refuses to run until the previous produced its file, and the file says what it needs to hear: rehearse (sandbox, host-guarded) -> judge --pre -> gate (human) -> apply (prod) -> judge --post. The gate binds to a manifest digest, so approving a change-set approves THAT change-set. An op is defined ONCE, as an API call, and replayed on the sandbox then on production — because the first design described each write twice (a sandbox script input and a prod API body) and the pre-gate judge immediately caught them diverging: the rehearsal was creating a EUR invoice with no due date while production would have received a USD one at 60 days. Two descriptions of the same write are two things that can disagree. Judges are context-free, cross-family per the PRD qa-strategy rule, and advisory: a BLOCK still lets the operator approve, and the override is recorded with their name. Blocking authority stays with the human gate and the host guards — an LLM verdict never silently starts or stops a production write. Verified end to end against the real 24/08 change-set (M3 deferred, USD 3,000): - pre-gate judge (Mistral) returned BLOCK twice, correctly — first on the sandbox/prod divergence, then on a duplicate left by a repeated rehearsal; - apply refuses after a rejected gate; - apply refuses without ARCO_PROD_CONFIRM; - editing an amount after approval invalidates the gate on digest mismatch. Co-Authored-By: Claude Opus 5 (1M context) <[email protected]> Claude-Session: https://claude.ai/code/session_01VRShc4QhLLU73FLHx9vskh
41 lines
1.8 KiB
Markdown
41 lines
1.8 KiB
Markdown
You are an independent reviewer. A change-set was rehearsed on a sandbox, a human
|
|
approved it, and it has now been applied to the **production** accounting ledger.
|
|
The writes already happened — you cannot prevent them. Your job is to make any
|
|
discrepancy a recorded finding instead of a surprise discovered months later.
|
|
|
|
You did not build or approve this change and have no conversation history about
|
|
it. Judge only the evidence below.
|
|
|
|
**Assume production drifted from what the rehearsal predicted, and look for the
|
|
drift.**
|
|
|
|
Compare `sandbox_after` (what the rehearsal produced) with `production_after`
|
|
(what production now holds), keeping in mind that these are two different systems:
|
|
internal ids, timestamps and document references legitimately differ. What must
|
|
match is the **substance**:
|
|
|
|
1. **Same objects, same count.** Everything the manifest intended exists in
|
|
production — and nothing extra appeared.
|
|
2. **Same amounts.** Totals, currencies, tax treatment.
|
|
3. **Same dates and states.** Invoice dates, due dates, paid/unpaid, validated
|
|
or draft.
|
|
4. **Every op reported ok.** Check `apply_results`; a failed op mid-run may have
|
|
left the ledger partially written, which matters more than anything else here.
|
|
5. **Chronology.** Document numbering in production stayed in date order.
|
|
|
|
Then answer in this shape, nothing else:
|
|
|
|
```
|
|
VERDICT: PASS (or) VERDICT: BLOCK
|
|
REASON: one sentence.
|
|
DRIFT:
|
|
- one line per substantive difference between rehearsal and production.
|
|
(write "none" if the substance matches)
|
|
FOLLOW-UP:
|
|
- what a human must now do, if anything. (write "none" if nothing)
|
|
```
|
|
|
|
Start your reply with the VERDICT line. `BLOCK` here does not undo anything — it
|
|
means a human must act. Differences in ids, refs or timestamps are expected and
|
|
are not drift; do not report them.
|