The discipline (ADR-0003, the promote flow, the operating rules) was written down and still depended on whoever was driving choosing to follow it. On 2026-07-25 an agent session wrote five documents into the production ledger through direct API calls, bypassing the promote flow entirely — a correct result reached by a path nobody could audit. A rule an operator can skip is a recommendation. The five stages are now chained by artefacts on disk. Each refuses to run until the previous produced its file, and the file says what it needs to hear: rehearse (sandbox, host-guarded) -> judge --pre -> gate (human) -> apply (prod) -> judge --post. The gate binds to a manifest digest, so approving a change-set approves THAT change-set. An op is defined ONCE, as an API call, and replayed on the sandbox then on production — because the first design described each write twice (a sandbox script input and a prod API body) and the pre-gate judge immediately caught them diverging: the rehearsal was creating a EUR invoice with no due date while production would have received a USD one at 60 days. Two descriptions of the same write are two things that can disagree. Judges are context-free, cross-family per the PRD qa-strategy rule, and advisory: a BLOCK still lets the operator approve, and the override is recorded with their name. Blocking authority stays with the human gate and the host guards — an LLM verdict never silently starts or stops a production write. Verified end to end against the real 24/08 change-set (M3 deferred, USD 3,000): - pre-gate judge (Mistral) returned BLOCK twice, correctly — first on the sandbox/prod divergence, then on a duplicate left by a repeated rehearsal; - apply refuses after a rejected gate; - apply refuses without ARCO_PROD_CONFIRM; - editing an amount after approval invalidates the gate on digest mismatch. Co-Authored-By: Claude Opus 5 (1M context) <[email protected]> Claude-Session: https://claude.ai/code/session_01VRShc4QhLLU73FLHx9vskh
fleet/harness/ — the multi-runtime harness layer
The harness is the orchestration layer around the atoms: builder sessions that
execute backlog issues, cold verifiers that check them (locate-tests, backlog
audits, refutation passes), and the evidence flow into Gitea. Per the PRD
model-fleet › harness portability
(operator direction 2026-07-15), this layer must not have Anthropic as a hard
dependency: the same loop runs on Mistral (vibe -p, mistral-medium-3.5)
or on hermes-served local models (Ornith / MLX, 127.0.0.1:18080). Claude is
an escalation tier, not a prerequisite. Admission of a runtime to a role is
evidence-gated (erp#63):
verifier roles first, scoped builders benched second, and no acceptance gate is
ever relaxed for a cheaper runtime.
Layout
| Path | Role |
|---|---|
verifier/locate-test.md |
canonical locate-test: prompt, inputs, ground truth, pass rule |
verifier/backlog-audit.md |
canonical cold-reader backlog audit: prompt, inputs, rubric |
bin/run-verifier.sh |
run a verifier test against a runtime; emits a JSON transcript |
bin/vibe-builder.sh |
run a scoped builder bench (vibe -p) inside a worktree, with caps + journal |
runs/<date>/ |
committed evidence transcripts, when they back an issue comment |
Runtimes
| Runtime | How the harness reaches it | Typical role |
|---|---|---|
claude |
a context-free subagent in a Claude Code session, given the exact assembled prompt (run-verifier.sh <test> --print-prompt) and nothing else |
baseline verifier; multi-file builder (default per the PRD complexity ceiling) |
ornith |
hermes MLX server, OpenAI-style POST /v1/chat/completions on 127.0.0.1:18080, model leonsarmiento/Ornith-1.0-35B-5bit-mlx |
verifier (candidate) |
mlx --model <id> |
same endpoint, any model the server lists under /v1/models |
verifier (candidate) |
mistral |
vibe -p programmatic mode, tools disabled, model = the vibe active_model (today mistral-medium-3.5) |
verifier (candidate); scoped builder via vibe-builder.sh |
Verifier protocol — no self-grading
- Assemble the prompt from the canonical test file + the pinned input documents
(
run-verifier.shembeds file contents verbatim and records their sha256). - Run every candidate runtime on the same assembled prompt, temperature 0.
- An independent, context-free judge (never the session that built the thing, per the PRD qa-strategy) scores each transcript against the test's ground truth and emits the parity table. A runtime is admitted to verifier duty when it reaches verdict parity with the Claude baseline on both tests.
- Once a non-Claude verifier is admitted, prefer cross-family verification: the verifier SHOULD be a different model family than the builder — a foreign family refuting the builder is stronger evidence than the builder's own family agreeing with itself.
Builder bench protocol
vibe-builder.sh runs one tightly-footered backlog issue end-to-end under a
non-Claude runtime, against the unchanged Execution footer and acceptance
gates. It measures completion, intervention count and wall-clock; a failed bench
is a valid result — it sets the complexity ceiling honestly. Safety bounds:
- refuses to run anywhere that is not a linked worktree (never the trunk —
same structural-guard pattern as
dol-write.sh); - hard caps:
--max-turnsand--max-priceare always set; --auto-approveis acceptable only because the blast radius is bounded: a disposable worktree, read-only API credentials, and the caps above;- the full
vibeJSON journal is kept per run.
Recurring tasks on the Mistral tier
A recurring task (T11 reminders, T13 drift checks, T14 backup freshness) is a
scoped builder with a standing prompt: cron (hermes cron or the operator's
scheduler) calls vibe-builder.sh <worktree> <task-prompt.md> and routes the
journal into the digest. The task prompt lives with the atom
(fleet/atoms/<atom>/prompt.md + its class skeleton); the harness only supplies
the bounded execution shell. No recurring task writes outside its worktree, and
anything ERP-write-shaped still goes through the sandbox + promote gate —
runtime choice never changes the gates.