feat(fleet): multi-runtime harness — verifier tests + capped builder shell
The harness layer (builder sessions, cold verifiers, evidence flow) gets a committable home, per the PRD model-fleet § harness portability and erp#63: - fleet/harness/verifier/: the two canonical verifier tests (locate-test, cold-reader backlog audit) with pinned inputs, verbatim prompts, ground truth and pass rules — judged context-free, never self-graded. - fleet/harness/bin/run-verifier.sh: runs a test against any OpenAI-style local endpoint (Ornith/MLX) or vibe -p (Mistral); emits sha256-pinned JSON transcripts. - fleet/harness/bin/vibe-builder.sh: the bounded shell for scoped builders and recurring tasks — refuses the trunk (linked-worktree guard), hard --max-turns/--max-price caps, full JSON journal per run. - fleet/README.md layout + AGENTS.md Fleet section updated in the same change (same-change freshness rule). Part of erp#63 (harness portability spike, D2). Co-Authored-By: Claude Fable 5 <[email protected]> Claude-Session: https://claude.ai/code/session_01VRShc4QhLLU73FLHx9vskh
This commit is contained in:
@@ -0,0 +1,72 @@
|
||||
# fleet/harness/ — the multi-runtime harness layer
|
||||
|
||||
The **harness** is the orchestration layer around the atoms: builder sessions that
|
||||
execute backlog issues, cold verifiers that check them (locate-tests, backlog
|
||||
audits, refutation passes), and the evidence flow into Gitea. Per the PRD
|
||||
[model-fleet › harness portability](https://gitea.arcodange.lab/arcodange-org/factory/src/branch/main/vibe/PRD/ai-back-office/model-fleet.md#harness-portability)
|
||||
(operator direction 2026-07-15), this layer must not have Anthropic as a hard
|
||||
dependency: the same loop runs on **Mistral** (`vibe -p`, `mistral-medium-3.5`)
|
||||
or on **hermes-served local models** (Ornith / MLX, `127.0.0.1:18080`). Claude is
|
||||
an escalation tier, not a prerequisite. Admission of a runtime to a role is
|
||||
**evidence-gated** ([erp#63](https://gitea.arcodange.lab/arcodange-org/erp/issues/63)):
|
||||
verifier roles first, scoped builders benched second, and no acceptance gate is
|
||||
ever relaxed for a cheaper runtime.
|
||||
|
||||
## Layout
|
||||
|
||||
| Path | Role |
|
||||
| --- | --- |
|
||||
| `verifier/locate-test.md` | canonical locate-test: prompt, inputs, ground truth, pass rule |
|
||||
| `verifier/backlog-audit.md` | canonical cold-reader backlog audit: prompt, inputs, rubric |
|
||||
| `bin/run-verifier.sh` | run a verifier test against a runtime; emits a JSON transcript |
|
||||
| `bin/vibe-builder.sh` | run a scoped builder bench (`vibe -p`) inside a worktree, with caps + journal |
|
||||
| `runs/<date>/` | committed evidence transcripts, when they back an issue comment |
|
||||
|
||||
## Runtimes
|
||||
|
||||
| Runtime | How the harness reaches it | Typical role |
|
||||
| --- | --- | --- |
|
||||
| `claude` | a **context-free subagent** in a Claude Code session, given the exact assembled prompt (`run-verifier.sh <test> --print-prompt`) and nothing else | baseline verifier; multi-file builder (default per the PRD complexity ceiling) |
|
||||
| `ornith` | hermes MLX server, OpenAI-style `POST /v1/chat/completions` on `127.0.0.1:18080`, model `leonsarmiento/Ornith-1.0-35B-5bit-mlx` | verifier (candidate) |
|
||||
| `mlx --model <id>` | same endpoint, any model the server lists under `/v1/models` | verifier (candidate) |
|
||||
| `mistral` | `vibe -p` programmatic mode, tools disabled, model = the vibe `active_model` (today `mistral-medium-3.5`) | verifier (candidate); scoped builder via `vibe-builder.sh` |
|
||||
|
||||
## Verifier protocol — no self-grading
|
||||
|
||||
1. Assemble the prompt from the canonical test file + the pinned input documents
|
||||
(`run-verifier.sh` embeds file contents verbatim and records their sha256).
|
||||
2. Run every candidate runtime on the **same assembled prompt**, temperature 0.
|
||||
3. **An independent, context-free judge** (never the session that built the thing,
|
||||
per the PRD [qa-strategy](https://gitea.arcodange.lab/arcodange-org/factory/src/branch/main/vibe/PRD/ai-back-office/qa-strategy.md#independent-verification--no-self-grading))
|
||||
scores each transcript against the test's ground truth and emits the parity
|
||||
table. A runtime is **admitted to verifier duty** when it reaches verdict
|
||||
parity with the Claude baseline on both tests.
|
||||
4. Once a non-Claude verifier is admitted, **prefer cross-family verification**:
|
||||
the verifier SHOULD be a different model family than the builder — a foreign
|
||||
family refuting the builder is stronger evidence than the builder's own family
|
||||
agreeing with itself.
|
||||
|
||||
## Builder bench protocol
|
||||
|
||||
`vibe-builder.sh` runs one tightly-footered backlog issue end-to-end under a
|
||||
non-Claude runtime, against the **unchanged** Execution footer and acceptance
|
||||
gates. It measures completion, intervention count and wall-clock; a failed bench
|
||||
is a valid result — it sets the complexity ceiling honestly. Safety bounds:
|
||||
|
||||
- refuses to run anywhere that is not a **linked worktree** (never the trunk —
|
||||
same structural-guard pattern as `dol-write.sh`);
|
||||
- hard caps: `--max-turns` and `--max-price` are always set;
|
||||
- `--auto-approve` is acceptable only because the blast radius is bounded: a
|
||||
disposable worktree, read-only API credentials, and the caps above;
|
||||
- the full `vibe` JSON journal is kept per run.
|
||||
|
||||
## Recurring tasks on the Mistral tier
|
||||
|
||||
A recurring task (T11 reminders, T13 drift checks, T14 backup freshness) is a
|
||||
**scoped builder with a standing prompt**: cron (hermes `cron` or the operator's
|
||||
scheduler) calls `vibe-builder.sh <worktree> <task-prompt.md>` and routes the
|
||||
journal into the digest. The task prompt lives with the atom
|
||||
(`fleet/atoms/<atom>/prompt.md` + its class skeleton); the harness only supplies
|
||||
the bounded execution shell. No recurring task writes outside its worktree, and
|
||||
anything ERP-write-shaped still goes through the sandbox + promote gate —
|
||||
runtime choice never changes the gates.
|
||||
Reference in New Issue
Block a user