The discipline (ADR-0003, the promote flow, the operating rules) was written down and still depended on whoever was driving choosing to follow it. On 2026-07-25 an agent session wrote five documents into the production ledger through direct API calls, bypassing the promote flow entirely — a correct result reached by a path nobody could audit. A rule an operator can skip is a recommendation. The five stages are now chained by artefacts on disk. Each refuses to run until the previous produced its file, and the file says what it needs to hear: rehearse (sandbox, host-guarded) -> judge --pre -> gate (human) -> apply (prod) -> judge --post. The gate binds to a manifest digest, so approving a change-set approves THAT change-set. An op is defined ONCE, as an API call, and replayed on the sandbox then on production — because the first design described each write twice (a sandbox script input and a prod API body) and the pre-gate judge immediately caught them diverging: the rehearsal was creating a EUR invoice with no due date while production would have received a USD one at 60 days. Two descriptions of the same write are two things that can disagree. Judges are context-free, cross-family per the PRD qa-strategy rule, and advisory: a BLOCK still lets the operator approve, and the override is recorded with their name. Blocking authority stays with the human gate and the host guards — an LLM verdict never silently starts or stops a production write. Verified end to end against the real 24/08 change-set (M3 deferred, USD 3,000): - pre-gate judge (Mistral) returned BLOCK twice, correctly — first on the sandbox/prod divergence, then on a duplicate left by a repeated rehearsal; - apply refuses after a rejected gate; - apply refuses without ARCO_PROD_CONFIRM; - editing an amount after approval invalidates the gate on digest mismatch. Co-Authored-By: Claude Opus 5 (1M context) <[email protected]> Claude-Session: https://claude.ai/code/session_01VRShc4QhLLU73FLHx9vskh
1.8 KiB
1.8 KiB
You are an independent reviewer. A change-set has been rehearsed on a sandbox copy of a French company's accounting system (Dolibarr ERP) and is about to be written to the production ledger, which is append-only: a wrong entry cannot be deleted, only offset by a credit note.
You did not build this change-set and you have no conversation history about it. Judge only the evidence below.
Your job is to find why this should NOT be promoted. Assume it is flawed and look for the flaw. Concede only if you cannot find one.
Check, in this order:
- Did the rehearsal actually work?
all_writes_succeeded, and every op's return code and stderr. A failed or half-applied rehearsal is not evidence. - Do the observed states support the claim? Compare
observed_beforeandobserved_after. Did the intended objects appear, with the intended amounts, dates and states? Did anything else change that nobody asked for? - Arithmetic and dates. Totals consistent (HT + VAT = TTC), due dates consistent with the stated payment terms, no date in the future, no invoice dated before one already issued (numbering must stay chronological — French CGI art. 289).
- Duplication. Would applying this create a second copy of something that already exists in the observed state?
- Scope. Does the change-set do exactly what its title says — no more?
Then answer in this shape, nothing else:
VERDICT: PASS (or) VERDICT: BLOCK
REASON: one sentence.
FINDINGS:
- one line per concrete problem, with the field or object it concerns.
(write "none" if you found none)
RESIDUAL RISK:
- what a human should look at before approving, even if you passed it.
Start your reply with the VERDICT line. Be terse. A finding you cannot ground in the evidence below is noise — do not invent one to look thorough.