Files
arcodangeandClaude Opus 5 1cba032512 fix(security): production read agent is now actually read-only
Migration applied to production, in the only order that does not break the
promote flow:

1. Created `ai_agent_prod_prod_write` (id=5) with the narrow `prod-write` scope —
   invoices, payments, and document submission (needed by builddoc to regenerate
   a modified invoice's PDF). Verified functionally: reads pass, DELETE on an
   invoice returns 403.
2. Repointed the promote pipeline at that user's key. It no longer borrows the
   read skills' credential; if the key is absent it dies with the provisioning
   command rather than silently falling back.
3. Revoked 12 write/delete rights from `ai_agent` (id=3), the credential every
   read skill holds: create/modify on customer AND supplier invoices,
   thirdparties, contacts, thirdparty payment details, proposals, exports,
   accounting links — and delete on proposals, events, and GED documents.

Verified after: invoices, thirdparties, contacts, products, proposals, supplier
invoices and bank accounts all still read; creating an invoice returns
`403 Forbidden: Insuffisant rights`. The documented posture and the real one
finally agree.

scopes.ts corrected against the live instance: 262 is NOT "créer/modifier les
produits" as the first catalogue guessed but the `voir_tous` ACL extension — a
READ right the skills depend on (without it, list endpoints return empty arrays
instead of 403). Revoking it would have silently blinded every read skill. This
is why the audit reads labels off /user/perms.php rather than trusting ids in
code. The READ_ONLY baseline is now the audited read surface (35 rights), not a
guess.

Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
Claude-Session: https://claude.ai/code/session_01VRShc4QhLLU73FLHx9vskh
2026-08-09 19:33:00 +02:00
..

fleet/harness/ — the multi-runtime harness layer

The harness is the orchestration layer around the atoms: builder sessions that execute backlog issues, cold verifiers that check them (locate-tests, backlog audits, refutation passes), and the evidence flow into Gitea. Per the PRD model-fleet harness portability (operator direction 2026-07-15), this layer must not have Anthropic as a hard dependency: the same loop runs on Mistral (vibe -p, mistral-medium-3.5) or on hermes-served local models (Ornith / MLX, 127.0.0.1:18080). Claude is an escalation tier, not a prerequisite. Admission of a runtime to a role is evidence-gated (erp#63): verifier roles first, scoped builders benched second, and no acceptance gate is ever relaxed for a cheaper runtime.

Layout

Path Role
verifier/locate-test.md canonical locate-test: prompt, inputs, ground truth, pass rule
verifier/backlog-audit.md canonical cold-reader backlog audit: prompt, inputs, rubric
bin/run-verifier.sh run a verifier test against a runtime; emits a JSON transcript
bin/vibe-builder.sh run a scoped builder bench (vibe -p) inside a worktree, with caps + journal
runs/<date>/ committed evidence transcripts, when they back an issue comment

Runtimes

Runtime How the harness reaches it Typical role
claude a context-free subagent in a Claude Code session, given the exact assembled prompt (run-verifier.sh <test> --print-prompt) and nothing else baseline verifier; multi-file builder (default per the PRD complexity ceiling)
ornith hermes MLX server, OpenAI-style POST /v1/chat/completions on 127.0.0.1:18080, model leonsarmiento/Ornith-1.0-35B-5bit-mlx verifier (candidate)
mlx --model <id> same endpoint, any model the server lists under /v1/models verifier (candidate)
mistral vibe -p programmatic mode, tools disabled, model = the vibe active_model (today mistral-medium-3.5) verifier (candidate); scoped builder via vibe-builder.sh

Verifier protocol — no self-grading

  1. Assemble the prompt from the canonical test file + the pinned input documents (run-verifier.sh embeds file contents verbatim and records their sha256).
  2. Run every candidate runtime on the same assembled prompt, temperature 0.
  3. An independent, context-free judge (never the session that built the thing, per the PRD qa-strategy) scores each transcript against the test's ground truth and emits the parity table. A runtime is admitted to verifier duty when it reaches verdict parity with the Claude baseline on both tests.
  4. Once a non-Claude verifier is admitted, prefer cross-family verification: the verifier SHOULD be a different model family than the builder — a foreign family refuting the builder is stronger evidence than the builder's own family agreeing with itself.

Builder bench protocol

vibe-builder.sh runs one tightly-footered backlog issue end-to-end under a non-Claude runtime, against the unchanged Execution footer and acceptance gates. It measures completion, intervention count and wall-clock; a failed bench is a valid result — it sets the complexity ceiling honestly. Safety bounds:

  • refuses to run anywhere that is not a linked worktree (never the trunk — same structural-guard pattern as dol-write.sh);
  • hard caps: --max-turns and --max-price are always set;
  • --auto-approve is acceptable only because the blast radius is bounded: a disposable worktree, read-only API credentials, and the caps above;
  • the full vibe JSON journal is kept per run.

Recurring tasks on the Mistral tier

A recurring task (T11 reminders, T13 drift checks, T14 backup freshness) is a scoped builder with a standing prompt: cron (hermes cron or the operator's scheduler) calls vibe-builder.sh <worktree> <task-prompt.md> and routes the journal into the digest. The task prompt lives with the atom (fleet/atoms/<atom>/prompt.md + its class skeleton); the harness only supplies the bounded execution shell. No recurring task writes outside its worktree, and anything ERP-write-shaped still goes through the sandbox + promote gate — runtime choice never changes the gates.