The discipline was written down — ADR-0003, the promote flow, the operating rules in AGENTS.md — and still depended on whoever was driving choosing to follow it. On 2026-07-25 an agent session wrote five documents into the production ledger through direct API calls, bypassing the promote flow entirely. The result was correct; the path was unauditable. A rule an operator can skip is a recommendation.
What lands
Five stages chained by artefacts on disk, not by intentions. Each refuses to run until the previous produced its file, and that file says what this stage needs to hear:
Stage
Produces
Refuses unless
rehearse
01-rehearsal.json
the target is the sandbox (host-guarded)
judge --pre
02-pre-verdict.json
a rehearsal exists
gate
03-gate.json
a pre-verdict exists; a human types the decision
apply
04-applied.json
the gate says approved, by a named human, for this exact change-set
judge --post
05-post-verdict.json
production was applied
The gate binds to a manifest digest: edit one amount after approval and apply refuses, because the approval no longer describes what is about to be written.
One definition per op
The first design described each write twice — a sandbox script input and a production API body. The pre-gate judge caught them diverging on the very first run: the rehearsal was creating a EUR invoice with no due date while production would have received a USD one at 60 days. We would have validated one thing and applied another. An op is now defined once, as an API call, replayed on the sandbox then on production.
The judges
Context-free, cross-family per the PRD qa-strategy rule, and advisory: a BLOCK still lets the operator approve, and the override is recorded with their name. Blocking authority stays with the human gate and the host guards — an LLM verdict never silently starts or stops a production write.
Verified end to end
Against the real 24/08 change-set (M3 deferred, USD 3,000):
pre-gate judge (Mistral) returned BLOCK twice, correctly — first on the sandbox/production divergence above, then on a duplicate left by a repeated rehearsal;
apply refuses after a rejected gate;
apply refuses without ARCO_PROD_CONFIRM;
editing an amount after approval invalidates the gate on digest mismatch.
The discipline was written down — [ADR-0003](https://gitea.arcodange.lab/arcodange-org/factory/src/branch/main/vibe/ADR/0003-sandbox-state-lifecycle.md), the promote flow, the operating rules in `AGENTS.md` — and still depended on whoever was driving choosing to follow it. **On 2026-07-25 an agent session wrote five documents into the production ledger through direct API calls, bypassing the promote flow entirely.** The result was correct; the path was unauditable. A rule an operator can skip is a recommendation.
## What lands
Five stages **chained by artefacts on disk**, not by intentions. Each refuses to run until the previous produced its file, and that file says what this stage needs to hear:
| Stage | Produces | Refuses unless |
| --- | --- | --- |
| `rehearse` | `01-rehearsal.json` | the target is the sandbox (host-guarded) |
| `judge --pre` | `02-pre-verdict.json` | a rehearsal exists |
| `gate` | `03-gate.json` | a pre-verdict exists; **a human types the decision** |
| `apply` | `04-applied.json` | the gate says `approved`, by a named human, **for this exact change-set** |
| `judge --post` | `05-post-verdict.json` | production was applied |
The gate binds to a **manifest digest**: edit one amount after approval and `apply` refuses, because the approval no longer describes what is about to be written.
## One definition per op
The first design described each write **twice** — a sandbox script input and a production API body. The pre-gate judge caught them diverging on the very first run: the rehearsal was creating a **EUR invoice with no due date** while production would have received a **USD one at 60 days**. We would have validated one thing and applied another. An op is now defined once, as an API call, replayed on the sandbox then on production.
## The judges
Context-free, cross-family per the PRD [qa-strategy rule](https://gitea.arcodange.lab/arcodange-org/factory/src/branch/main/vibe/PRD/ai-back-office/qa-strategy.md#independent-verification--no-self-grading), and **advisory**: a `BLOCK` still lets the operator approve, and the override is recorded with their name. Blocking authority stays with the human gate and the host guards — an LLM verdict never silently starts or stops a production write.
## Verified end to end
Against the real 24/08 change-set (M3 deferred, USD 3,000):
- pre-gate judge (Mistral) returned **BLOCK twice, correctly** — first on the sandbox/production divergence above, then on a duplicate left by a repeated rehearsal;
- `apply` refuses after a rejected gate;
- `apply` refuses without `ARCO_PROD_CONFIRM`;
- editing an amount after approval invalidates the gate on digest mismatch.
🤖 Generated with [Claude Code](https://claude.com/claude-code)
https://claude.ai/code/session_01VRShc4QhLLU73FLHx9vskh
The discipline (ADR-0003, the promote flow, the operating rules) was written
down and still depended on whoever was driving choosing to follow it. On
2026-07-25 an agent session wrote five documents into the production ledger
through direct API calls, bypassing the promote flow entirely — a correct result
reached by a path nobody could audit. A rule an operator can skip is a
recommendation.
The five stages are now chained by artefacts on disk. Each refuses to run until
the previous produced its file, and the file says what it needs to hear:
rehearse (sandbox, host-guarded) -> judge --pre -> gate (human) -> apply
(prod) -> judge --post. The gate binds to a manifest digest, so approving a
change-set approves THAT change-set.
An op is defined ONCE, as an API call, and replayed on the sandbox then on
production — because the first design described each write twice (a sandbox
script input and a prod API body) and the pre-gate judge immediately caught them
diverging: the rehearsal was creating a EUR invoice with no due date while
production would have received a USD one at 60 days. Two descriptions of the
same write are two things that can disagree.
Judges are context-free, cross-family per the PRD qa-strategy rule, and
advisory: a BLOCK still lets the operator approve, and the override is recorded
with their name. Blocking authority stays with the human gate and the host
guards — an LLM verdict never silently starts or stops a production write.
Verified end to end against the real 24/08 change-set (M3 deferred, USD 3,000):
- pre-gate judge (Mistral) returned BLOCK twice, correctly — first on the
sandbox/prod divergence, then on a duplicate left by a repeated rehearsal;
- apply refuses after a rejected gate;
- apply refuses without ARCO_PROD_CONFIRM;
- editing an amount after approval invalidates the gate on digest mismatch.
Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
Claude-Session: https://claude.ai/code/session_01VRShc4QhLLU73FLHx9vskh
Blocking a user prevents them from interacting with repositories, such as opening or commenting on pull requests or issues. Learn more about blocking a user.
The discipline was written down — ADR-0003, the promote flow, the operating rules in
AGENTS.md— and still depended on whoever was driving choosing to follow it. On 2026-07-25 an agent session wrote five documents into the production ledger through direct API calls, bypassing the promote flow entirely. The result was correct; the path was unauditable. A rule an operator can skip is a recommendation.What lands
Five stages chained by artefacts on disk, not by intentions. Each refuses to run until the previous produced its file, and that file says what this stage needs to hear:
rehearse01-rehearsal.jsonjudge --pre02-pre-verdict.jsongate03-gate.jsonapply04-applied.jsonapproved, by a named human, for this exact change-setjudge --post05-post-verdict.jsonThe gate binds to a manifest digest: edit one amount after approval and
applyrefuses, because the approval no longer describes what is about to be written.One definition per op
The first design described each write twice — a sandbox script input and a production API body. The pre-gate judge caught them diverging on the very first run: the rehearsal was creating a EUR invoice with no due date while production would have received a USD one at 60 days. We would have validated one thing and applied another. An op is now defined once, as an API call, replayed on the sandbox then on production.
The judges
Context-free, cross-family per the PRD qa-strategy rule, and advisory: a
BLOCKstill lets the operator approve, and the override is recorded with their name. Blocking authority stays with the human gate and the host guards — an LLM verdict never silently starts or stops a production write.Verified end to end
Against the real 24/08 change-set (M3 deferred, USD 3,000):
applyrefuses after a rejected gate;applyrefuses withoutARCO_PROD_CONFIRM;🤖 Generated with Claude Code
https://claude.ai/code/session_01VRShc4QhLLU73FLHx9vskh