Migration applied to production, in the only order that does not break the promote flow: 1. Created `ai_agent_prod_prod_write` (id=5) with the narrow `prod-write` scope — invoices, payments, and document submission (needed by builddoc to regenerate a modified invoice's PDF). Verified functionally: reads pass, DELETE on an invoice returns 403. 2. Repointed the promote pipeline at that user's key. It no longer borrows the read skills' credential; if the key is absent it dies with the provisioning command rather than silently falling back. 3. Revoked 12 write/delete rights from `ai_agent` (id=3), the credential every read skill holds: create/modify on customer AND supplier invoices, thirdparties, contacts, thirdparty payment details, proposals, exports, accounting links — and delete on proposals, events, and GED documents. Verified after: invoices, thirdparties, contacts, products, proposals, supplier invoices and bank accounts all still read; creating an invoice returns `403 Forbidden: Insuffisant rights`. The documented posture and the real one finally agree. scopes.ts corrected against the live instance: 262 is NOT "créer/modifier les produits" as the first catalogue guessed but the `voir_tous` ACL extension — a READ right the skills depend on (without it, list endpoints return empty arrays instead of 403). Revoking it would have silently blinded every read skill. This is why the audit reads labels off /user/perms.php rather than trusting ids in code. The READ_ONLY baseline is now the audited read surface (35 rights), not a guess. Co-Authored-By: Claude Opus 5 (1M context) <[email protected]> Claude-Session: https://claude.ai/code/session_01VRShc4QhLLU73FLHx9vskh
fleet/harness/ — the multi-runtime harness layer
The harness is the orchestration layer around the atoms: builder sessions that
execute backlog issues, cold verifiers that check them (locate-tests, backlog
audits, refutation passes), and the evidence flow into Gitea. Per the PRD
model-fleet › harness portability
(operator direction 2026-07-15), this layer must not have Anthropic as a hard
dependency: the same loop runs on Mistral (vibe -p, mistral-medium-3.5)
or on hermes-served local models (Ornith / MLX, 127.0.0.1:18080). Claude is
an escalation tier, not a prerequisite. Admission of a runtime to a role is
evidence-gated (erp#63):
verifier roles first, scoped builders benched second, and no acceptance gate is
ever relaxed for a cheaper runtime.
Layout
| Path | Role |
|---|---|
verifier/locate-test.md |
canonical locate-test: prompt, inputs, ground truth, pass rule |
verifier/backlog-audit.md |
canonical cold-reader backlog audit: prompt, inputs, rubric |
bin/run-verifier.sh |
run a verifier test against a runtime; emits a JSON transcript |
bin/vibe-builder.sh |
run a scoped builder bench (vibe -p) inside a worktree, with caps + journal |
runs/<date>/ |
committed evidence transcripts, when they back an issue comment |
Runtimes
| Runtime | How the harness reaches it | Typical role |
|---|---|---|
claude |
a context-free subagent in a Claude Code session, given the exact assembled prompt (run-verifier.sh <test> --print-prompt) and nothing else |
baseline verifier; multi-file builder (default per the PRD complexity ceiling) |
ornith |
hermes MLX server, OpenAI-style POST /v1/chat/completions on 127.0.0.1:18080, model leonsarmiento/Ornith-1.0-35B-5bit-mlx |
verifier (candidate) |
mlx --model <id> |
same endpoint, any model the server lists under /v1/models |
verifier (candidate) |
mistral |
vibe -p programmatic mode, tools disabled, model = the vibe active_model (today mistral-medium-3.5) |
verifier (candidate); scoped builder via vibe-builder.sh |
Verifier protocol — no self-grading
- Assemble the prompt from the canonical test file + the pinned input documents
(
run-verifier.shembeds file contents verbatim and records their sha256). - Run every candidate runtime on the same assembled prompt, temperature 0.
- An independent, context-free judge (never the session that built the thing, per the PRD qa-strategy) scores each transcript against the test's ground truth and emits the parity table. A runtime is admitted to verifier duty when it reaches verdict parity with the Claude baseline on both tests.
- Once a non-Claude verifier is admitted, prefer cross-family verification: the verifier SHOULD be a different model family than the builder — a foreign family refuting the builder is stronger evidence than the builder's own family agreeing with itself.
Builder bench protocol
vibe-builder.sh runs one tightly-footered backlog issue end-to-end under a
non-Claude runtime, against the unchanged Execution footer and acceptance
gates. It measures completion, intervention count and wall-clock; a failed bench
is a valid result — it sets the complexity ceiling honestly. Safety bounds:
- refuses to run anywhere that is not a linked worktree (never the trunk —
same structural-guard pattern as
dol-write.sh); - hard caps:
--max-turnsand--max-priceare always set; --auto-approveis acceptable only because the blast radius is bounded: a disposable worktree, read-only API credentials, and the caps above;- the full
vibeJSON journal is kept per run.
Recurring tasks on the Mistral tier
A recurring task (T11 reminders, T13 drift checks, T14 backup freshness) is a
scoped builder with a standing prompt: cron (hermes cron or the operator's
scheduler) calls vibe-builder.sh <worktree> <task-prompt.md> and routes the
journal into the digest. The task prompt lives with the atom
(fleet/atoms/<atom>/prompt.md + its class skeleton); the harness only supplies
the bounded execution shell. No recurring task writes outside its worktree, and
anything ERP-write-shaped still goes through the sandbox + promote gate —
runtime choice never changes the gates.