3 000 USD manquants, découverts en assemblant le dossier client complet. La cartographie des cycles ne laisse pas de doute : M1, M2 et M4 ont chacun leur part fixe ET leur part différée ; M3 n'avait que sa part fixe (FAC008, réglée le 20/07). La facture pro forma PR2607-0001, remise au client en juillet, annonçait pourtant l'émission de cette part différée pour le 24/08/2026 — soit le jour même. C'est la date retenue : la numérotation reste chronologique au sens de l'article 289 du CGI, et l'échéance reste au 23/10/2026, trois mois après la fin du cycle M3. Le retard est écrit dans la note de la facture plutôt que masqué par une date d'émission rétroactive, qui elle romprait la chronologie. LE JUGE A POSÉ LA BONNE QUESTION. Il demandait si le différé de trois mois court depuis l'émission ou depuis la fin du cycle — les deux lectures ne divergent que pour M3, dont la facture est tardive : 23/10 contre 24/11. La pro forma tranche en toutes lettres et dans ses deux versions : « émise le 24/08/2026, à échéance du 23/10/2026 », et « paiement différé à 3 mois DE LA PRESTATION ». Retenir une autre date reviendrait à s'écarter d'une annonce déjà faite au client. Preuve en 02b-preuve-echeance.txt. CE QUI A ÉTÉ RETIRÉ DU CHANGE-SET EN COURS DE ROUTE. Une première version ajoutait la clause de pénalités L.441-10 à FAC001, FAC002 et FAC003, qui en sont dépourvues. L'opérateur a précisé le circuit réel : pendant la phase d'audit (janvier-février 2026), c'est WISE qui a émis les factures, en reprenant les identifiants internes de Dolibarr. Ces trois-là ONT donc été transmises au client. Modifier leur note aujourd'hui ferait diverger le registre interne du document que le client détient — l'inverse du but recherché. Le défaut de mention est réel mais historique, sur des factures réglées : consigné, pas réécrit. Depuis avril 2026 en revanche, KissMetrics vire le montant sur le compte Wise sans recevoir aucune facture. C'est ce qui a permis de reprendre FAC004, FAC006 et FAC009 plus tôt dans la journée, et c'est le dossier du 24/08 qui porte ces factures au client pour la première fois. Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
fleet/harness/ — the multi-runtime harness layer
The harness is the orchestration layer around the atoms: builder sessions that
execute backlog issues, cold verifiers that check them (locate-tests, backlog
audits, refutation passes), and the evidence flow into Gitea. Per the PRD
model-fleet › harness portability
(operator direction 2026-07-15), this layer must not have Anthropic as a hard
dependency: the same loop runs on Mistral (vibe -p, mistral-medium-3.5)
or on hermes-served local models (Ornith / MLX, 127.0.0.1:18080). Claude is
an escalation tier, not a prerequisite. Admission of a runtime to a role is
evidence-gated (erp#63):
verifier roles first, scoped builders benched second, and no acceptance gate is
ever relaxed for a cheaper runtime.
Layout
| Path | Role |
|---|---|
verifier/locate-test.md |
canonical locate-test: prompt, inputs, ground truth, pass rule |
verifier/backlog-audit.md |
canonical cold-reader backlog audit: prompt, inputs, rubric |
bin/run-verifier.sh |
run a verifier test against a runtime; emits a JSON transcript |
bin/vibe-builder.sh |
run a scoped builder bench (vibe -p) inside a worktree, with caps + journal |
runs/<date>/ |
committed evidence transcripts, when they back an issue comment |
Runtimes
| Runtime | How the harness reaches it | Typical role |
|---|---|---|
claude |
a context-free subagent in a Claude Code session, given the exact assembled prompt (run-verifier.sh <test> --print-prompt) and nothing else |
baseline verifier; multi-file builder (default per the PRD complexity ceiling) |
ornith |
hermes MLX server, OpenAI-style POST /v1/chat/completions on 127.0.0.1:18080, model leonsarmiento/Ornith-1.0-35B-5bit-mlx |
verifier (candidate) |
mlx --model <id> |
same endpoint, any model the server lists under /v1/models |
verifier (candidate) |
mistral |
vibe -p programmatic mode, tools disabled, model = the vibe active_model (today mistral-medium-3.5) |
verifier (candidate); scoped builder via vibe-builder.sh |
Verifier protocol — no self-grading
- Assemble the prompt from the canonical test file + the pinned input documents
(
run-verifier.shembeds file contents verbatim and records their sha256). - Run every candidate runtime on the same assembled prompt, temperature 0.
- An independent, context-free judge (never the session that built the thing, per the PRD qa-strategy) scores each transcript against the test's ground truth and emits the parity table. A runtime is admitted to verifier duty when it reaches verdict parity with the Claude baseline on both tests.
- Once a non-Claude verifier is admitted, prefer cross-family verification: the verifier SHOULD be a different model family than the builder — a foreign family refuting the builder is stronger evidence than the builder's own family agreeing with itself.
Builder bench protocol
vibe-builder.sh runs one tightly-footered backlog issue end-to-end under a
non-Claude runtime, against the unchanged Execution footer and acceptance
gates. It measures completion, intervention count and wall-clock; a failed bench
is a valid result — it sets the complexity ceiling honestly. Safety bounds:
- refuses to run anywhere that is not a linked worktree (never the trunk —
same structural-guard pattern as
dol-write.sh); - hard caps:
--max-turnsand--max-priceare always set; --auto-approveis acceptable only because the blast radius is bounded: a disposable worktree, read-only API credentials, and the caps above;- the full
vibeJSON journal is kept per run.
Recurring tasks on the Mistral tier
A recurring task (T11 reminders, T13 drift checks, T14 backup freshness) is a
scoped builder with a standing prompt: cron (hermes cron or the operator's
scheduler) calls vibe-builder.sh <worktree> <task-prompt.md> and routes the
journal into the digest. The task prompt lives with the atom
(fleet/atoms/<atom>/prompt.md + its class skeleton); the harness only supplies
the bounded execution shell. No recurring task writes outside its worktree, and
anything ERP-write-shaped still goes through the sandbox + promote gate —
runtime choice never changes the gates.