From f2a60817e28588eeba8a8cca0aae2014c3009ef2 Mon Sep 17 00:00:00 2001 From: Gabriel Radureau Date: Sat, 18 Jul 2026 18:52:53 +0200 Subject: [PATCH] =?UTF-8?q?feat(fleet):=20multi-runtime=20harness=20?= =?UTF-8?q?=E2=80=94=20verifier=20tests=20+=20capped=20builder=20shell?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit The harness layer (builder sessions, cold verifiers, evidence flow) gets a committable home, per the PRD model-fleet § harness portability and erp#63: - fleet/harness/verifier/: the two canonical verifier tests (locate-test, cold-reader backlog audit) with pinned inputs, verbatim prompts, ground truth and pass rules — judged context-free, never self-graded. - fleet/harness/bin/run-verifier.sh: runs a test against any OpenAI-style local endpoint (Ornith/MLX) or vibe -p (Mistral); emits sha256-pinned JSON transcripts. - fleet/harness/bin/vibe-builder.sh: the bounded shell for scoped builders and recurring tasks — refuses the trunk (linked-worktree guard), hard --max-turns/--max-price caps, full JSON journal per run. - fleet/README.md layout + AGENTS.md Fleet section updated in the same change (same-change freshness rule). Part of erp#63 (harness portability spike, D2). Co-Authored-By: Claude Fable 5 Claude-Session: https://claude.ai/code/session_01VRShc4QhLLU73FLHx9vskh --- AGENTS.md | 1 + fleet/README.md | 3 +- fleet/harness/README.md | 72 ++++++++++++++ fleet/harness/bin/run-verifier.sh | 125 ++++++++++++++++++++++++ fleet/harness/bin/vibe-builder.sh | 69 +++++++++++++ fleet/harness/verifier/backlog-audit.md | 44 +++++++++ fleet/harness/verifier/locate-test.md | 36 +++++++ 7 files changed, 349 insertions(+), 1 deletion(-) create mode 100644 fleet/harness/README.md create mode 100755 fleet/harness/bin/run-verifier.sh create mode 100755 fleet/harness/bin/vibe-builder.sh create mode 100644 fleet/harness/verifier/backlog-audit.md create mode 100644 fleet/harness/verifier/locate-test.md diff --git a/AGENTS.md b/AGENTS.md index 9919a05..aaf0924 100644 --- a/AGENTS.md +++ b/AGENTS.md @@ -38,6 +38,7 @@ The [AI back-office PRD](https://gitea.arcodange.lab/arcodange-org/factory/src/b - **Environment rules**: the [operating rules](#operating-rules-for-agents) above + [`.claude/skills/dolibarr-sandbox-write/SKILL.md`](.claude/skills/dolibarr-sandbox-write/SKILL.md) (the host-guarded sandbox write path and its promote gate). - **Autonomy ladder**: levels A0–A3 in the PRD [hub](https://gitea.arcodange.lab/arcodange-org/factory/src/branch/main/vibe/PRD/ai-back-office/README.md#the-autonomy-ladder); promotion/demotion per the PRD [qa-strategy gates](https://gitea.arcodange.lab/arcodange-org/factory/src/branch/main/vibe/PRD/ai-back-office/qa-strategy.md#autonomy-promotion-gates). - **Graduation**: an atom earns autonomy through its golden-set evals and unedited-approval streaks — the earned level + eval evidence live in its `atom.yaml` `autonomy` field, and a promotion is a PR changing that field with the evidence linked. +- **Harness**: [`fleet/harness/`](fleet/harness/README.md) — the multi-runtime execution layer around the atoms: canonical verifier tests (locate-test, backlog audit), `run-verifier.sh` for any OpenAI-style local endpoint or `vibe -p` (Mistral), and `vibe-builder.sh` (the capped, worktree-guarded shell for scoped builders and recurring tasks). Runtimes are admitted per role by evidence ([erp#63](https://gitea.arcodange.lab/arcodange-org/erp/issues/63)); Claude is the escalation tier, not a prerequisite, per the PRD [harness portability](https://gitea.arcodange.lab/arcodange-org/factory/src/branch/main/vibe/PRD/ai-back-office/model-fleet.md#harness-portability). ## Before building anything diff --git a/fleet/README.md b/fleet/README.md index bc88d8f..0346a97 100644 --- a/fleet/README.md +++ b/fleet/README.md @@ -47,7 +47,8 @@ fleet/ │ ├── prompt.md │ └── scripts/ ├── golden/ # per-atom golden sets — land with erp#39 -└── profile/ # fiscal.yaml + calendar.yaml + ADC register — land with erp#54 +├── profile/ # fiscal.yaml + calendar.yaml + ADC register — land with erp#54 +└── harness/ # multi-runtime harness layer: verifier tests + builder bench (harness/README.md) ``` ## `atom.yaml` — the contract, field by field diff --git a/fleet/harness/README.md b/fleet/harness/README.md new file mode 100644 index 0000000..5fdbcb7 --- /dev/null +++ b/fleet/harness/README.md @@ -0,0 +1,72 @@ +# fleet/harness/ — the multi-runtime harness layer + +The **harness** is the orchestration layer around the atoms: builder sessions that +execute backlog issues, cold verifiers that check them (locate-tests, backlog +audits, refutation passes), and the evidence flow into Gitea. Per the PRD +[model-fleet › harness portability](https://gitea.arcodange.lab/arcodange-org/factory/src/branch/main/vibe/PRD/ai-back-office/model-fleet.md#harness-portability) +(operator direction 2026-07-15), this layer must not have Anthropic as a hard +dependency: the same loop runs on **Mistral** (`vibe -p`, `mistral-medium-3.5`) +or on **hermes-served local models** (Ornith / MLX, `127.0.0.1:18080`). Claude is +an escalation tier, not a prerequisite. Admission of a runtime to a role is +**evidence-gated** ([erp#63](https://gitea.arcodange.lab/arcodange-org/erp/issues/63)): +verifier roles first, scoped builders benched second, and no acceptance gate is +ever relaxed for a cheaper runtime. + +## Layout + +| Path | Role | +| --- | --- | +| `verifier/locate-test.md` | canonical locate-test: prompt, inputs, ground truth, pass rule | +| `verifier/backlog-audit.md` | canonical cold-reader backlog audit: prompt, inputs, rubric | +| `bin/run-verifier.sh` | run a verifier test against a runtime; emits a JSON transcript | +| `bin/vibe-builder.sh` | run a scoped builder bench (`vibe -p`) inside a worktree, with caps + journal | +| `runs//` | committed evidence transcripts, when they back an issue comment | + +## Runtimes + +| Runtime | How the harness reaches it | Typical role | +| --- | --- | --- | +| `claude` | a **context-free subagent** in a Claude Code session, given the exact assembled prompt (`run-verifier.sh --print-prompt`) and nothing else | baseline verifier; multi-file builder (default per the PRD complexity ceiling) | +| `ornith` | hermes MLX server, OpenAI-style `POST /v1/chat/completions` on `127.0.0.1:18080`, model `leonsarmiento/Ornith-1.0-35B-5bit-mlx` | verifier (candidate) | +| `mlx --model ` | same endpoint, any model the server lists under `/v1/models` | verifier (candidate) | +| `mistral` | `vibe -p` programmatic mode, tools disabled, model = the vibe `active_model` (today `mistral-medium-3.5`) | verifier (candidate); scoped builder via `vibe-builder.sh` | + +## Verifier protocol — no self-grading + +1. Assemble the prompt from the canonical test file + the pinned input documents + (`run-verifier.sh` embeds file contents verbatim and records their sha256). +2. Run every candidate runtime on the **same assembled prompt**, temperature 0. +3. **An independent, context-free judge** (never the session that built the thing, + per the PRD [qa-strategy](https://gitea.arcodange.lab/arcodange-org/factory/src/branch/main/vibe/PRD/ai-back-office/qa-strategy.md#independent-verification--no-self-grading)) + scores each transcript against the test's ground truth and emits the parity + table. A runtime is **admitted to verifier duty** when it reaches verdict + parity with the Claude baseline on both tests. +4. Once a non-Claude verifier is admitted, **prefer cross-family verification**: + the verifier SHOULD be a different model family than the builder — a foreign + family refuting the builder is stronger evidence than the builder's own family + agreeing with itself. + +## Builder bench protocol + +`vibe-builder.sh` runs one tightly-footered backlog issue end-to-end under a +non-Claude runtime, against the **unchanged** Execution footer and acceptance +gates. It measures completion, intervention count and wall-clock; a failed bench +is a valid result — it sets the complexity ceiling honestly. Safety bounds: + +- refuses to run anywhere that is not a **linked worktree** (never the trunk — + same structural-guard pattern as `dol-write.sh`); +- hard caps: `--max-turns` and `--max-price` are always set; +- `--auto-approve` is acceptable only because the blast radius is bounded: a + disposable worktree, read-only API credentials, and the caps above; +- the full `vibe` JSON journal is kept per run. + +## Recurring tasks on the Mistral tier + +A recurring task (T11 reminders, T13 drift checks, T14 backup freshness) is a +**scoped builder with a standing prompt**: cron (hermes `cron` or the operator's +scheduler) calls `vibe-builder.sh ` and routes the +journal into the digest. The task prompt lives with the atom +(`fleet/atoms//prompt.md` + its class skeleton); the harness only supplies +the bounded execution shell. No recurring task writes outside its worktree, and +anything ERP-write-shaped still goes through the sandbox + promote gate — +runtime choice never changes the gates. diff --git a/fleet/harness/bin/run-verifier.sh b/fleet/harness/bin/run-verifier.sh new file mode 100755 index 0000000..e20972a --- /dev/null +++ b/fleet/harness/bin/run-verifier.sh @@ -0,0 +1,125 @@ +#!/usr/bin/env bash +# Run a canonical verifier test against a runtime; emit a JSON transcript. +# See fleet/harness/README.md (runtimes, no-self-grading protocol). +set -euo pipefail + +usage() { + cat >&2 <<'EOF' +usage: run-verifier.sh [options] + run-verifier.sh --print-prompt [options] + +options: + --agents-file PATH AGENTS.md to pin (default: repo-root AGENTS.md) + --status-file PATH STATUS.md to pin (required for the backlog test) + --model ID model id (required for mlx; optional override for ornith) + --endpoint URL OpenAI-style base (default: http://127.0.0.1:18080/v1) + --out DIR transcript dir (default: $TMPDIR/harness-runs) + --print-prompt print the assembled prompt to stdout and exit +EOF + exit 2 +} + +TEST="${1:-}"; shift || usage +case "$TEST" in locate|backlog) ;; *) usage ;; esac + +RUNTIME="" PRINT_ONLY=0 MODEL="" ENDPOINT="http://127.0.0.1:18080/v1" +AGENTS_FILE="" STATUS_FILE="" OUT_DIR="${TMPDIR:-/tmp}/harness-runs" +while [ $# -gt 0 ]; do + case "$1" in + ornith|mlx|mistral) RUNTIME="$1" ;; + --print-prompt) PRINT_ONLY=1 ;; + --agents-file) AGENTS_FILE="$2"; shift ;; + --status-file) STATUS_FILE="$2"; shift ;; + --model) MODEL="$2"; shift ;; + --endpoint) ENDPOINT="$2"; shift ;; + --out) OUT_DIR="$2"; shift ;; + *) usage ;; + esac + shift +done +[ "$PRINT_ONLY" = 1 ] || [ -n "$RUNTIME" ] || usage + +HERE="$(cd "$(dirname "$0")/.." && pwd)" +REPO_ROOT="$(git -C "$HERE" rev-parse --show-toplevel)" +[ -n "$AGENTS_FILE" ] || AGENTS_FILE="$REPO_ROOT/AGENTS.md" +[ -f "$AGENTS_FILE" ] || { echo "missing AGENTS.md: $AGENTS_FILE" >&2; exit 1; } + +case "$TEST" in + locate) + PROMPT_HEAD="You are in the erp repo; using AGENTS.md alone, state where the atom registry, the class skeletons and the environment rules live." + ;; + backlog) + [ -f "${STATUS_FILE:-}" ] || { echo "backlog test requires --status-file" >&2; exit 1; } + PROMPT_HEAD="You are a cold reader auditing the Arcodange AI back-office backlog. Using ONLY the two documents below — no other knowledge, no tools — answer: +1. What shipped most recently? +2. What should be worked on next, and why that item? +3. What would you verify before trusting these documents, and in what order?" + ;; +esac + +PROMPT="$PROMPT_HEAD + +--- AGENTS.md --- +$(cat "$AGENTS_FILE")" +if [ "$TEST" = backlog ]; then + PROMPT="$PROMPT + +--- STATUS.md --- +$(cat "$STATUS_FILE")" +fi + +if [ "$PRINT_ONLY" = 1 ]; then + printf '%s\n' "$PROMPT" + exit 0 +fi + +mkdir -p "$OUT_DIR" +TS="$(date +%Y%m%dT%H%M%S)" +OUT_FILE="$OUT_DIR/${TEST}-${RUNTIME}${MODEL:+-$(basename "$MODEL")}-$TS.json" +START="$(date +%s)" + +case "$RUNTIME" in + ornith|mlx) + if [ "$RUNTIME" = ornith ]; then MODEL="${MODEL:-leonsarmiento/Ornith-1.0-35B-5bit-mlx}"; fi + [ -n "$MODEL" ] || { echo "mlx runtime requires --model" >&2; exit 1; } + RESPONSE="$(PROMPT="$PROMPT" MODEL="$MODEL" python3 - "$ENDPOINT" <<'PY' +import json, os, sys, urllib.request +body = json.dumps({ + "model": os.environ["MODEL"], + "messages": [{"role": "user", "content": os.environ["PROMPT"]}], + "temperature": 0, + "max_tokens": 3000, +}).encode() +req = urllib.request.Request(sys.argv[1].rstrip("/") + "/chat/completions", + data=body, headers={"Content-Type": "application/json"}) +with urllib.request.urlopen(req, timeout=900) as r: + print(json.load(r)["choices"][0]["message"]["content"]) +PY +)" + ;; + mistral) + MODEL="vibe-active-model" + RESPONSE="$(vibe -p "$PROMPT" --max-turns 1 --enabled-tools __none__ --output text)" + ;; +esac + +LATENCY=$(( $(date +%s) - START )) +RESPONSE="$RESPONSE" PROMPT="$PROMPT" TEST="$TEST" RUNTIME="$RUNTIME" MODEL="$MODEL" \ +ENDPOINT="$ENDPOINT" LATENCY="$LATENCY" TS="$TS" AGENTS_FILE="$AGENTS_FILE" STATUS_FILE="${STATUS_FILE:-}" \ +python3 - > "$OUT_FILE" <<'PY' +import hashlib, json, os +def sha(p): + return hashlib.sha256(open(p, "rb").read()).hexdigest() if p and os.path.exists(p) else None +e = os.environ +inputs = {"AGENTS.md": {"path": e["AGENTS_FILE"], "sha256": sha(e["AGENTS_FILE"])}} +if e["STATUS_FILE"]: + inputs["STATUS.md"] = {"path": e["STATUS_FILE"], "sha256": sha(e["STATUS_FILE"])} +print(json.dumps({ + "test": e["TEST"], "runtime": e["RUNTIME"], "model": e["MODEL"], + "endpoint": e["ENDPOINT"] if e["RUNTIME"] != "mistral" else "vibe -p", + "timestamp": e["TS"], "latency_s": int(e["LATENCY"]), + "prompt_sha256": hashlib.sha256(e["PROMPT"].encode()).hexdigest(), + "inputs": inputs, "response": e["RESPONSE"], +}, indent=2, ensure_ascii=False)) +PY +echo "$OUT_FILE" diff --git a/fleet/harness/bin/vibe-builder.sh b/fleet/harness/bin/vibe-builder.sh new file mode 100755 index 0000000..2f57fda --- /dev/null +++ b/fleet/harness/bin/vibe-builder.sh @@ -0,0 +1,69 @@ +#!/usr/bin/env bash +# Scoped builder bench / recurring-task shell: run `vibe -p` inside a linked +# worktree with hard caps and a JSON journal. Refuses to run in the trunk. +# See fleet/harness/README.md (builder bench protocol, safety bounds). +set -euo pipefail + +usage() { + cat >&2 <<'EOF' +usage: vibe-builder.sh [--max-turns N] [--max-price DOLLARS] [--out DIR] + +Runs: vibe -p "$(cat prompt-file)" --auto-approve --max-turns N --max-price D --output json +inside , which MUST be a linked git worktree (never the trunk). +Journal: /builder--.json (+ .meta.json with wall-clock and exit code). +Defaults: --max-turns 60, --max-price 3, --out $TMPDIR/harness-runs. +EOF + exit 2 +} + +WORKTREE="${1:-}"; PROMPT_FILE="${2:-}" +[ -d "$WORKTREE" ] && [ -f "${PROMPT_FILE:-}" ] || usage +shift 2 +MAX_TURNS=60 MAX_PRICE=3 OUT_DIR="${TMPDIR:-/tmp}/harness-runs" +while [ $# -gt 0 ]; do + case "$1" in + --max-turns) MAX_TURNS="$2"; shift ;; + --max-price) MAX_PRICE="$2"; shift ;; + --out) OUT_DIR="$2"; shift ;; + *) usage ;; + esac + shift +done + +# Structural guard: a linked worktree has .git as a FILE (gitdir pointer); +# the trunk has .git as a directory. Same never-the-trunk guarantee as dol-write.sh. +if [ ! -f "$WORKTREE/.git" ]; then + echo "REFUSED: $WORKTREE is not a linked git worktree (trunk is reserved for the user)" >&2 + exit 1 +fi + +mkdir -p "$OUT_DIR" +TS="$(date +%Y%m%dT%H%M%S)" +NAME="$(basename "$WORKTREE")" +JOURNAL="$OUT_DIR/builder-$NAME-$TS.json" +META="$OUT_DIR/builder-$NAME-$TS.meta.json" + +START="$(date +%s)" +set +e +(cd "$WORKTREE" && vibe -p "$(cat "$PROMPT_FILE")" --auto-approve \ + --max-turns "$MAX_TURNS" --max-price "$MAX_PRICE" --output json) > "$JOURNAL" 2>"$JOURNAL.stderr" +EXIT_CODE=$? +set -e +WALL=$(( $(date +%s) - START )) + +PROMPT_SHA="$(shasum -a 256 "$PROMPT_FILE" | cut -d' ' -f1)" +cat > "$META" < You are a cold reader auditing the Arcodange AI back-office backlog. Using ONLY +> the two documents below — no other knowledge, no tools — answer: +> 1. What shipped most recently? +> 2. What should be worked on next, and why that item? +> 3. What would you verify before trusting these documents, and in what order? + +`bin/run-verifier.sh backlog --status-file ` assembles this +prompt with both documents appended under `--- AGENTS.md ---` and +`--- STATUS.md ---` delimiters. + +## Rubric + +| Question | Substantively correct means | +| --- | --- | +| 1 — shipped | names the most recent ✅ item(s) the STATUS actually records (e.g. the fleet scaffold / latest merged PR in its ledger) — not an older phase, not an open issue | +| 2 — next | applies the resume protocol: top **unblocked** issue of the **earliest open milestone by due date**, and skips issues whose only open step is a `[HUMAN]` gate (surfaced, not stalled on) | +| 3 — trust | states the trust order live system > code/git log > STATUS > PRD leaves > memories (wording may vary; the direction must be right), and proposes checking Last Updated / verifying claims against the live repo | + +## Pass rule + +PASS = substantively correct on all three questions. Minor omissions that do not +invert the protocol (e.g. not listing every milestone) do not fail; inverting +the trust order, proposing a blocked/`[HUMAN]`-gated issue as "next" without +flagging the gate, or citing shipped work the documents do not support, fails. +Judged by an independent context-free judge — never by the session that produced +the run. diff --git a/fleet/harness/verifier/locate-test.md b/fleet/harness/verifier/locate-test.md new file mode 100644 index 0000000..063cb07 --- /dev/null +++ b/fleet/harness/verifier/locate-test.md @@ -0,0 +1,36 @@ +# Verifier test — locate-test + +First run: erp#38 acceptance gate (Claude, 2026-07-15, PASS — evidence on the +issue). Re-run on candidate runtimes for verifier admission (erp#63). + +## Inputs + +| Document | Pinned source | +| --- | --- | +| `AGENTS.md` | the erp repo root, at the commit under test | + +The runtime gets the document content embedded in the prompt and **nothing else** +— no repo access, no tools, no conversation history. + +## Prompt (verbatim — do not paraphrase when re-running) + +> You are in the erp repo; using AGENTS.md alone, state where the atom registry, +> the class skeletons and the environment rules live. + +`bin/run-verifier.sh locate ` assembles this prompt with the document +appended under a `--- AGENTS.md ---` delimiter. + +## Ground truth + +| Item | Correct answer | +| --- | --- | +| atom registry | `fleet/README.md` | +| class skeletons | `fleet/classes/` (7 skeleton files) | +| environment rules | the AGENTS.md **operating rules** section + `.claude/skills/dolibarr-sandbox-write/SKILL.md` | + +## Pass rule + +PASS = all three locations correct (path-level; exact anchor wording not +required; for environment rules both halves must appear). Judged by an +independent context-free judge against this table — never by the session that +produced the run.