Files
erp/fleet/harness
arcodangeandClaude Opus 5 67a74924e9 feat(erp): indemnité d'occupation janv→juil 2026 — paiements divers, sans tiers
Appliqué en production : 7 écritures, 1 483,23 EUR, compte courant d'associé
porté de -429,75 à -1 912,98. Grand livre auxiliaire intact — 12 tiers, 15
factures fournisseur, inchangés.

La première approche créait un tiers fournisseur « Radureau Gabriel » et lui
adressait 7 factures. C'était faux : le gérant n'est pas un fournisseur de sa
société, et lui ouvrir une fiche l'aurait fait apparaître au grand livre
auxiliaire, dans les balances âgées et les états de dettes fournisseurs —
l'objection exacte déjà opposée à l'URSSAF dans RUNBOOK_charges_sociales.md,
que j'ai reproduite en la contredisant. L'existant le disait pourtant : les 8
dettes déjà portées au compte courant sont toutes des factures de fournisseurs
RÉELS payées personnellement, le tiers n'étant jamais le gérant.

Le modèle correct est direct. Le compte bancaire CCA1 (id 3) porte le numéro
comptable 45511 et son propre journal ; un paiement divers en sens débit, code
613000 Locations, produit

    débit  613000  Locations                     (la charge)
    crédit 45511   G. RADUREAU, compte courant   (la dette envers l'associé)

Correction au passage : le plan comptable EST chargé (358 comptes, dont un
455110 dédié au compte courant du gérant). adc-009 affirme « module comptabilité
pas déployé » en confondant trois choses — le plan chargé, l'API REST absente,
et le dictionnaire des types de charges sans code comptable.

Deux bugs de ma main, trouvés en répétition :

- l'idempotence comparait les 24 PREMIERS CARACTÈRES du libellé, or
  « Indemnité d'occupation — » en fait exactement 24 : mars reconnaissait
  février et se déclarait déjà enregistré. Six mois silencieusement sautés.
  On compare désormais le libellé entier, avec tolérance à la troncature
  « … » de Dolibarr. recordSocialCharge.ts porte le même défaut, latent :
  ses libellés diffèrent avant le 24e caractère, aujourd'hui seulement.
- une boucle shell utilisait `set -- $m`, qui ne découpe pas les mots en zsh :
  la date devenait « 2026-- ». Le script a correctement refusé d'écrire.

/variouspayments répond « API not found » : le pipeline gated, qui parle REST,
ne peut pas porter cette opération — comme pour les charges sociales. Le script
en garde la discipline (répétition sandbox, relecture par la liste, opt-in
production explicite) sans le juge ni l'artefact de gate.

Le run gated abandonné est conservé sous fleet/harness/runs/ avec ABANDONNE.md :
il documente ce que le harness a vu, et surtout ce qu'il n'a pas vu.

Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
2026-08-13 21:04:55 +02:00
..

fleet/harness/ — the multi-runtime harness layer

The harness is the orchestration layer around the atoms: builder sessions that execute backlog issues, cold verifiers that check them (locate-tests, backlog audits, refutation passes), and the evidence flow into Gitea. Per the PRD model-fleet harness portability (operator direction 2026-07-15), this layer must not have Anthropic as a hard dependency: the same loop runs on Mistral (vibe -p, mistral-medium-3.5) or on hermes-served local models (Ornith / MLX, 127.0.0.1:18080). Claude is an escalation tier, not a prerequisite. Admission of a runtime to a role is evidence-gated (erp#63): verifier roles first, scoped builders benched second, and no acceptance gate is ever relaxed for a cheaper runtime.

Layout

Path Role
verifier/locate-test.md canonical locate-test: prompt, inputs, ground truth, pass rule
verifier/backlog-audit.md canonical cold-reader backlog audit: prompt, inputs, rubric
bin/run-verifier.sh run a verifier test against a runtime; emits a JSON transcript
bin/vibe-builder.sh run a scoped builder bench (vibe -p) inside a worktree, with caps + journal
runs/<date>/ committed evidence transcripts, when they back an issue comment

Runtimes

Runtime How the harness reaches it Typical role
claude a context-free subagent in a Claude Code session, given the exact assembled prompt (run-verifier.sh <test> --print-prompt) and nothing else baseline verifier; multi-file builder (default per the PRD complexity ceiling)
ornith hermes MLX server, OpenAI-style POST /v1/chat/completions on 127.0.0.1:18080, model leonsarmiento/Ornith-1.0-35B-5bit-mlx verifier (candidate)
mlx --model <id> same endpoint, any model the server lists under /v1/models verifier (candidate)
mistral vibe -p programmatic mode, tools disabled, model = the vibe active_model (today mistral-medium-3.5) verifier (candidate); scoped builder via vibe-builder.sh

Verifier protocol — no self-grading

  1. Assemble the prompt from the canonical test file + the pinned input documents (run-verifier.sh embeds file contents verbatim and records their sha256).
  2. Run every candidate runtime on the same assembled prompt, temperature 0.
  3. An independent, context-free judge (never the session that built the thing, per the PRD qa-strategy) scores each transcript against the test's ground truth and emits the parity table. A runtime is admitted to verifier duty when it reaches verdict parity with the Claude baseline on both tests.
  4. Once a non-Claude verifier is admitted, prefer cross-family verification: the verifier SHOULD be a different model family than the builder — a foreign family refuting the builder is stronger evidence than the builder's own family agreeing with itself.

Builder bench protocol

vibe-builder.sh runs one tightly-footered backlog issue end-to-end under a non-Claude runtime, against the unchanged Execution footer and acceptance gates. It measures completion, intervention count and wall-clock; a failed bench is a valid result — it sets the complexity ceiling honestly. Safety bounds:

  • refuses to run anywhere that is not a linked worktree (never the trunk — same structural-guard pattern as dol-write.sh);
  • hard caps: --max-turns and --max-price are always set;
  • --auto-approve is acceptable only because the blast radius is bounded: a disposable worktree, read-only API credentials, and the caps above;
  • the full vibe JSON journal is kept per run.

Recurring tasks on the Mistral tier

A recurring task (T11 reminders, T13 drift checks, T14 backup freshness) is a scoped builder with a standing prompt: cron (hermes cron or the operator's scheduler) calls vibe-builder.sh <worktree> <task-prompt.md> and routes the journal into the digest. The task prompt lives with the atom (fleet/atoms/<atom>/prompt.md + its class skeleton); the harness only supplies the bounded execution shell. No recurring task writes outside its worktree, and anything ERP-write-shaped still goes through the sandbox + promote gate — runtime choice never changes the gates.