Commit Graph
6 Commits
Author SHA1 Message Date
arcodangeandClaude Opus 5 5607e2af86 feat(erp): cycle M4 KissMetrics, comblement des écarts bancaires, clause de pénalités rétablie
Deux passages gated en production, et l'outil qui manquait pour clore le dossier
client.

CYCLE M4 (2026-08-24-m4-et-ecarts). Cinq opérations : le tiers Anthropic PBC —
entité américaine, distincte d'Anthropic Ireland Limited — et sa facture de
90 EUR réglée le 19/07 ; le règlement Darnis de 262,20 EUR ; les deux factures
du cycle M4 (part fixe 2 500 USD → 2 140,45 EUR échéance 22/09, part différée
3 000 USD → 2 568,54 EUR échéance 23/11). Les trois écarts bancaires relevés
sont comblés ; Qonto tombe à 3 806,08 EUR.

Le juge pré-gate a bloqué : le règlement Darnis du 14/08 précède la facture
datée du 31/08. C'est un fait bancaire, pas une erreur de saisie — Darnis
facture en fin de mois — et adc-008 avait EXPRESSÉMENT prévu ce blocage et
autorisé le passage outre tracé. L'art. 289 du CGI qu'invoque le juge régit la
numérotation des factures ÉMISES, non l'ordre entre un paiement et sa facture ;
c'est la troisième fois de la session qu'il l'étend hors de son champ.

CLAUSE DE PÉNALITÉS (2026-08-24-clause-bilingue). Les factures M4 sont sorties
avec une clause AMPUTÉE : sans la traduction anglaise, sans la mention « Ces
stipulations sont des minima légaux d'ordre public auxquels il ne peut être
renoncé ». Détecté par le contrôle des mentions obligatoires sur les PDF —
avant tout envoi. La note publique est rétablie dans sa rédaction de référence,
à droit constant : aucun montant, aucune date, aucune ligne touchés.

Deux raisons de corriger plutôt que de laisser courir. La permanence des
méthodes (PCG art. 121-5, adc-011) : une clause légale identique doit être
rédigée identiquement, sans quoi la variation se lit comme une intention. Et le
fond : le destinataire est américain, c'est la version anglaise qui lui rend la
clause opposable en fait — et c'est sur elle que s'appuie la relance.

Le juge a bloqué là aussi, en supposant que la référence portait « ce montant ».
Il ne l'avait pas vérifié, et son propre residual risk demandait de le faire.
Vérification faite et consignée (02b-preuve-reference.txt) : FAC005 à FAC008
portent toutes « ce forfait », et la clause proposée leur est identique au
caractère près — 1283 car. Le finding était inversé.

test/buildInvoicePdf.ts. Valider une facture par l'API ne produit AUCUN PDF : le
fichier n'existe que si quelqu'un l'a demandé. `PUT /documents/builddoc` le
ferait mais répond 403 — ce droit n'est accordé à aucun scope agent. Le script
emprunte donc l'interface, puis relit le fichier par l'API : on ne croit pas la
page de retour, on vérifie que le PDF est déposé et lisible par un tiers.

Les deux juges post-gate passent sans dérive.

Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
2026-08-24 13:12:34 +02:00
arcodange 480888937f feat(erp): indemnité d'occupation janv→juil 2026 (1 483,23 €) — paiements divers, sans tiers (#89)
Co-authored-by: Gabriel Radureau <[email protected]>
2026-08-13 21:09:10 +02:00
arcodangeandClaude Opus 5 1cba032512 fix(security): production read agent is now actually read-only
Migration applied to production, in the only order that does not break the
promote flow:

1. Created `ai_agent_prod_prod_write` (id=5) with the narrow `prod-write` scope —
   invoices, payments, and document submission (needed by builddoc to regenerate
   a modified invoice's PDF). Verified functionally: reads pass, DELETE on an
   invoice returns 403.
2. Repointed the promote pipeline at that user's key. It no longer borrows the
   read skills' credential; if the key is absent it dies with the provisioning
   command rather than silently falling back.
3. Revoked 12 write/delete rights from `ai_agent` (id=3), the credential every
   read skill holds: create/modify on customer AND supplier invoices,
   thirdparties, contacts, thirdparty payment details, proposals, exports,
   accounting links — and delete on proposals, events, and GED documents.

Verified after: invoices, thirdparties, contacts, products, proposals, supplier
invoices and bank accounts all still read; creating an invoice returns
`403 Forbidden: Insuffisant rights`. The documented posture and the real one
finally agree.

scopes.ts corrected against the live instance: 262 is NOT "créer/modifier les
produits" as the first catalogue guessed but the `voir_tous` ACL extension — a
READ right the skills depend on (without it, list endpoints return empty arrays
instead of 403). Revoking it would have silently blinded every read skill. This
is why the audit reads labels off /user/perms.php rather than trusting ids in
code. The READ_ONLY baseline is now the audited read surface (35 rights), not a
guess.

Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
Claude-Session: https://claude.ai/code/session_01VRShc4QhLLU73FLHx9vskh
2026-08-09 19:33:00 +02:00
arcodangeandClaude Opus 5 a2cafc0d6b feat(harness): gated promote pipeline — rehearse, judge, human gate, apply, judge
The discipline (ADR-0003, the promote flow, the operating rules) was written
down and still depended on whoever was driving choosing to follow it. On
2026-07-25 an agent session wrote five documents into the production ledger
through direct API calls, bypassing the promote flow entirely — a correct result
reached by a path nobody could audit. A rule an operator can skip is a
recommendation.

The five stages are now chained by artefacts on disk. Each refuses to run until
the previous produced its file, and the file says what it needs to hear:
rehearse (sandbox, host-guarded) -> judge --pre -> gate (human) -> apply
(prod) -> judge --post. The gate binds to a manifest digest, so approving a
change-set approves THAT change-set.

An op is defined ONCE, as an API call, and replayed on the sandbox then on
production — because the first design described each write twice (a sandbox
script input and a prod API body) and the pre-gate judge immediately caught them
diverging: the rehearsal was creating a EUR invoice with no due date while
production would have received a USD one at 60 days. Two descriptions of the
same write are two things that can disagree.

Judges are context-free, cross-family per the PRD qa-strategy rule, and
advisory: a BLOCK still lets the operator approve, and the override is recorded
with their name. Blocking authority stays with the human gate and the host
guards — an LLM verdict never silently starts or stops a production write.

Verified end to end against the real 24/08 change-set (M3 deferred, USD 3,000):
- pre-gate judge (Mistral) returned BLOCK twice, correctly — first on the
  sandbox/prod divergence, then on a duplicate left by a repeated rehearsal;
- apply refuses after a rejected gate;
- apply refuses without ARCO_PROD_CONFIRM;
- editing an amount after approval invalidates the gate on digest mismatch.

Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
Claude-Session: https://claude.ai/code/session_01VRShc4QhLLU73FLHx9vskh
2026-07-26 08:06:45 +02:00
arcodangeandClaude Fable 5 ceb4321224 chore(fleet): erp#63 evidence — verifier parity + builder bench transcripts
- runs/2026-07-18/: the 8 sha256-pinned verifier transcripts (4 runtimes ×
  2 tests), blind-judging verdicts (2 independent judges per cell, unanimous),
  the erp#56 builder-bench journal + prompt + caps, and the evidence README
  with the parity table.
- run-verifier.sh: mistral runtime drops the tool-filter flag (--enabled-tools
  with a no-match pattern hangs vibe 2.21.0); plain -p with --max-turns 1.

Verdicts: Mistral (vibe -p, mistral-medium-3.5) and Ornith 35B (hermes MLX)
reach verdict parity with the Claude baseline on both tests → admitted to
verifier duty. Qwen2.5-7B-4bit fails both → the honest small-model floor.
Builder bench: erp#56 completed by the Mistral runtime, 0 code corrections,
261 s, acceptance run clean (0 bank-UNKNOWN) → merged as PR #68.

Closes #63 (with the paired factory qa-strategy PR).

Co-Authored-By: Claude Fable 5 <[email protected]>
Claude-Session: https://claude.ai/code/session_01VRShc4QhLLU73FLHx9vskh
2026-07-18 19:56:55 +02:00
arcodangeandClaude Fable 5 f2a60817e2 feat(fleet): multi-runtime harness — verifier tests + capped builder shell
The harness layer (builder sessions, cold verifiers, evidence flow) gets a
committable home, per the PRD model-fleet § harness portability and erp#63:

- fleet/harness/verifier/: the two canonical verifier tests (locate-test,
  cold-reader backlog audit) with pinned inputs, verbatim prompts, ground
  truth and pass rules — judged context-free, never self-graded.
- fleet/harness/bin/run-verifier.sh: runs a test against any OpenAI-style
  local endpoint (Ornith/MLX) or vibe -p (Mistral); emits sha256-pinned
  JSON transcripts.
- fleet/harness/bin/vibe-builder.sh: the bounded shell for scoped builders
  and recurring tasks — refuses the trunk (linked-worktree guard), hard
  --max-turns/--max-price caps, full JSON journal per run.
- fleet/README.md layout + AGENTS.md Fleet section updated in the same
  change (same-change freshness rule).

Part of erp#63 (harness portability spike, D2).

Co-Authored-By: Claude Fable 5 <[email protected]>
Claude-Session: https://claude.ai/code/session_01VRShc4QhLLU73FLHx9vskh
2026-07-18 18:52:53 +02:00