Files
arcodangeandClaude Fable 5 f2a60817e2 feat(fleet): multi-runtime harness — verifier tests + capped builder shell
The harness layer (builder sessions, cold verifiers, evidence flow) gets a
committable home, per the PRD model-fleet § harness portability and erp#63:

- fleet/harness/verifier/: the two canonical verifier tests (locate-test,
  cold-reader backlog audit) with pinned inputs, verbatim prompts, ground
  truth and pass rules — judged context-free, never self-graded.
- fleet/harness/bin/run-verifier.sh: runs a test against any OpenAI-style
  local endpoint (Ornith/MLX) or vibe -p (Mistral); emits sha256-pinned
  JSON transcripts.
- fleet/harness/bin/vibe-builder.sh: the bounded shell for scoped builders
  and recurring tasks — refuses the trunk (linked-worktree guard), hard
  --max-turns/--max-price caps, full JSON journal per run.
- fleet/README.md layout + AGENTS.md Fleet section updated in the same
  change (same-change freshness rule).

Part of erp#63 (harness portability spike, D2).

Co-Authored-By: Claude Fable 5 <[email protected]>
Claude-Session: https://claude.ai/code/session_01VRShc4QhLLU73FLHx9vskh
2026-07-18 18:52:53 +02:00

73 lines
4.5 KiB
Markdown
Raw Permalink Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# fleet/harness/ — the multi-runtime harness layer
The **harness** is the orchestration layer around the atoms: builder sessions that
execute backlog issues, cold verifiers that check them (locate-tests, backlog
audits, refutation passes), and the evidence flow into Gitea. Per the PRD
[model-fleet harness portability](https://gitea.arcodange.lab/arcodange-org/factory/src/branch/main/vibe/PRD/ai-back-office/model-fleet.md#harness-portability)
(operator direction 2026-07-15), this layer must not have Anthropic as a hard
dependency: the same loop runs on **Mistral** (`vibe -p`, `mistral-medium-3.5`)
or on **hermes-served local models** (Ornith / MLX, `127.0.0.1:18080`). Claude is
an escalation tier, not a prerequisite. Admission of a runtime to a role is
**evidence-gated** ([erp#63](https://gitea.arcodange.lab/arcodange-org/erp/issues/63)):
verifier roles first, scoped builders benched second, and no acceptance gate is
ever relaxed for a cheaper runtime.
## Layout
| Path | Role |
| --- | --- |
| `verifier/locate-test.md` | canonical locate-test: prompt, inputs, ground truth, pass rule |
| `verifier/backlog-audit.md` | canonical cold-reader backlog audit: prompt, inputs, rubric |
| `bin/run-verifier.sh` | run a verifier test against a runtime; emits a JSON transcript |
| `bin/vibe-builder.sh` | run a scoped builder bench (`vibe -p`) inside a worktree, with caps + journal |
| `runs/<date>/` | committed evidence transcripts, when they back an issue comment |
## Runtimes
| Runtime | How the harness reaches it | Typical role |
| --- | --- | --- |
| `claude` | a **context-free subagent** in a Claude Code session, given the exact assembled prompt (`run-verifier.sh <test> --print-prompt`) and nothing else | baseline verifier; multi-file builder (default per the PRD complexity ceiling) |
| `ornith` | hermes MLX server, OpenAI-style `POST /v1/chat/completions` on `127.0.0.1:18080`, model `leonsarmiento/Ornith-1.0-35B-5bit-mlx` | verifier (candidate) |
| `mlx --model <id>` | same endpoint, any model the server lists under `/v1/models` | verifier (candidate) |
| `mistral` | `vibe -p` programmatic mode, tools disabled, model = the vibe `active_model` (today `mistral-medium-3.5`) | verifier (candidate); scoped builder via `vibe-builder.sh` |
## Verifier protocol — no self-grading
1. Assemble the prompt from the canonical test file + the pinned input documents
(`run-verifier.sh` embeds file contents verbatim and records their sha256).
2. Run every candidate runtime on the **same assembled prompt**, temperature 0.
3. **An independent, context-free judge** (never the session that built the thing,
per the PRD [qa-strategy](https://gitea.arcodange.lab/arcodange-org/factory/src/branch/main/vibe/PRD/ai-back-office/qa-strategy.md#independent-verification--no-self-grading))
scores each transcript against the test's ground truth and emits the parity
table. A runtime is **admitted to verifier duty** when it reaches verdict
parity with the Claude baseline on both tests.
4. Once a non-Claude verifier is admitted, **prefer cross-family verification**:
the verifier SHOULD be a different model family than the builder — a foreign
family refuting the builder is stronger evidence than the builder's own family
agreeing with itself.
## Builder bench protocol
`vibe-builder.sh` runs one tightly-footered backlog issue end-to-end under a
non-Claude runtime, against the **unchanged** Execution footer and acceptance
gates. It measures completion, intervention count and wall-clock; a failed bench
is a valid result — it sets the complexity ceiling honestly. Safety bounds:
- refuses to run anywhere that is not a **linked worktree** (never the trunk —
same structural-guard pattern as `dol-write.sh`);
- hard caps: `--max-turns` and `--max-price` are always set;
- `--auto-approve` is acceptable only because the blast radius is bounded: a
disposable worktree, read-only API credentials, and the caps above;
- the full `vibe` JSON journal is kept per run.
## Recurring tasks on the Mistral tier
A recurring task (T11 reminders, T13 drift checks, T14 backup freshness) is a
**scoped builder with a standing prompt**: cron (hermes `cron` or the operator's
scheduler) calls `vibe-builder.sh <worktree> <task-prompt.md>` and routes the
journal into the digest. The task prompt lives with the atom
(`fleet/atoms/<atom>/prompt.md` + its class skeleton); the harness only supplies
the bounded execution shell. No recurring task writes outside its worktree, and
anything ERP-write-shaped still goes through the sandbox + promote gate —
runtime choice never changes the gates.