Files
erp/fleet/harness
arcodangeandClaude Opus 5 a2cafc0d6b feat(harness): gated promote pipeline — rehearse, judge, human gate, apply, judge
The discipline (ADR-0003, the promote flow, the operating rules) was written
down and still depended on whoever was driving choosing to follow it. On
2026-07-25 an agent session wrote five documents into the production ledger
through direct API calls, bypassing the promote flow entirely — a correct result
reached by a path nobody could audit. A rule an operator can skip is a
recommendation.

The five stages are now chained by artefacts on disk. Each refuses to run until
the previous produced its file, and the file says what it needs to hear:
rehearse (sandbox, host-guarded) -> judge --pre -> gate (human) -> apply
(prod) -> judge --post. The gate binds to a manifest digest, so approving a
change-set approves THAT change-set.

An op is defined ONCE, as an API call, and replayed on the sandbox then on
production — because the first design described each write twice (a sandbox
script input and a prod API body) and the pre-gate judge immediately caught them
diverging: the rehearsal was creating a EUR invoice with no due date while
production would have received a USD one at 60 days. Two descriptions of the
same write are two things that can disagree.

Judges are context-free, cross-family per the PRD qa-strategy rule, and
advisory: a BLOCK still lets the operator approve, and the override is recorded
with their name. Blocking authority stays with the human gate and the host
guards — an LLM verdict never silently starts or stops a production write.

Verified end to end against the real 24/08 change-set (M3 deferred, USD 3,000):
- pre-gate judge (Mistral) returned BLOCK twice, correctly — first on the
  sandbox/prod divergence, then on a duplicate left by a repeated rehearsal;
- apply refuses after a rejected gate;
- apply refuses without ARCO_PROD_CONFIRM;
- editing an amount after approval invalidates the gate on digest mismatch.

Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
Claude-Session: https://claude.ai/code/session_01VRShc4QhLLU73FLHx9vskh
2026-07-26 08:06:45 +02:00
..

fleet/harness/ — the multi-runtime harness layer

The harness is the orchestration layer around the atoms: builder sessions that execute backlog issues, cold verifiers that check them (locate-tests, backlog audits, refutation passes), and the evidence flow into Gitea. Per the PRD model-fleet harness portability (operator direction 2026-07-15), this layer must not have Anthropic as a hard dependency: the same loop runs on Mistral (vibe -p, mistral-medium-3.5) or on hermes-served local models (Ornith / MLX, 127.0.0.1:18080). Claude is an escalation tier, not a prerequisite. Admission of a runtime to a role is evidence-gated (erp#63): verifier roles first, scoped builders benched second, and no acceptance gate is ever relaxed for a cheaper runtime.

Layout

Path Role
verifier/locate-test.md canonical locate-test: prompt, inputs, ground truth, pass rule
verifier/backlog-audit.md canonical cold-reader backlog audit: prompt, inputs, rubric
bin/run-verifier.sh run a verifier test against a runtime; emits a JSON transcript
bin/vibe-builder.sh run a scoped builder bench (vibe -p) inside a worktree, with caps + journal
runs/<date>/ committed evidence transcripts, when they back an issue comment

Runtimes

Runtime How the harness reaches it Typical role
claude a context-free subagent in a Claude Code session, given the exact assembled prompt (run-verifier.sh <test> --print-prompt) and nothing else baseline verifier; multi-file builder (default per the PRD complexity ceiling)
ornith hermes MLX server, OpenAI-style POST /v1/chat/completions on 127.0.0.1:18080, model leonsarmiento/Ornith-1.0-35B-5bit-mlx verifier (candidate)
mlx --model <id> same endpoint, any model the server lists under /v1/models verifier (candidate)
mistral vibe -p programmatic mode, tools disabled, model = the vibe active_model (today mistral-medium-3.5) verifier (candidate); scoped builder via vibe-builder.sh

Verifier protocol — no self-grading

  1. Assemble the prompt from the canonical test file + the pinned input documents (run-verifier.sh embeds file contents verbatim and records their sha256).
  2. Run every candidate runtime on the same assembled prompt, temperature 0.
  3. An independent, context-free judge (never the session that built the thing, per the PRD qa-strategy) scores each transcript against the test's ground truth and emits the parity table. A runtime is admitted to verifier duty when it reaches verdict parity with the Claude baseline on both tests.
  4. Once a non-Claude verifier is admitted, prefer cross-family verification: the verifier SHOULD be a different model family than the builder — a foreign family refuting the builder is stronger evidence than the builder's own family agreeing with itself.

Builder bench protocol

vibe-builder.sh runs one tightly-footered backlog issue end-to-end under a non-Claude runtime, against the unchanged Execution footer and acceptance gates. It measures completion, intervention count and wall-clock; a failed bench is a valid result — it sets the complexity ceiling honestly. Safety bounds:

  • refuses to run anywhere that is not a linked worktree (never the trunk — same structural-guard pattern as dol-write.sh);
  • hard caps: --max-turns and --max-price are always set;
  • --auto-approve is acceptable only because the blast radius is bounded: a disposable worktree, read-only API credentials, and the caps above;
  • the full vibe JSON journal is kept per run.

Recurring tasks on the Mistral tier

A recurring task (T11 reminders, T13 drift checks, T14 backup freshness) is a scoped builder with a standing prompt: cron (hermes cron or the operator's scheduler) calls vibe-builder.sh <worktree> <task-prompt.md> and routes the journal into the digest. The task prompt lives with the atom (fleet/atoms/<atom>/prompt.md + its class skeleton); the harness only supplies the bounded execution shell. No recurring task writes outside its worktree, and anything ERP-write-shaped still goes through the sandbox + promote gate — runtime choice never changes the gates.