L'entrée « urssaf-echeancier / pending-definition / recurrence unknown » traînait
depuis erp#54 avec la consigne « ne pas inventer de cadence ». L'opérateur a
communiqué l'échéancier : ce n'est pas trimestriel mais trois appels irréguliers.
493,00 (22/05, prélevé) + 1 215,00 (05/08) + 1 333,00 (05/11) = 3 041,00 EUR
Corrige aussi le compte comptable, faux dans la note : les cotisations d'un
gérant associé unique de SARLU (TNS) vont en 646 — cotisations personnelles du
dirigeant — et non en 645, qui vise les cotisations patronales sur salaires et
doit rester vide puisque Arcodange n'a aucun salarié.
Le schéma accepte désormais amount_eur : la boucle de rappels T11 (erp#60) a
besoin du montant en donnée structurée, pas noyé dans du texte libre.
Registre validé : 8 règles, 16 entrées, 8 ADC, 0 erreur.
Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
Claude-Session: https://claude.ai/code/session_01VRShc4QhLLU73FLHx9vskh
L'extraction de champs vivait dans un heredoc à l'intérieur d'email-inspect.sh :
impossible à exécuter isolément, donc jamais mesurée, donc fausse sans que
personne puisse le voir. Sur la facture Darnis F1048 elle renvoyait le numéro de
TVA d'Arcodange comme référence de facture, et aucune date.
- extract_fields.py : l'extraction sort du shell et devient un module.
- test_extract.py : régression contre les 16 factures hand-vérifiées de
fleet/golden/invoice-extract/. Score par champ, et une valeur FAUSSE pèse plus
qu'une valeur absente — un humain recopie ce qui s'affiche.
Valeurs fausses : 4 → 0. Exactitude ref 62,5 → 75 %, date 62,5 → 75 %,
HT 68,8 → 75 %, TTC 81,2 → 93,8 %.
Cinq bugs réels, dont trois invisibles sans test :
- « Nº » sur les factures françaises est U+00BA (ordinal masculin), pas le signe
degré. La classe [°o] le rate, le motif principal échoue, et le repli attrape
le premier jeton ref-shaped du document — très souvent un numéro de TVA.
- Le filtre anti-TVA rejetait « FR73261832 », qui est la vraie référence OVH : un
numéro FR fait exactement 11 caractères après le préfixe.
- « Montant total (HT) » était lu comme un TTC.
- Une référence coupée par la colonne (« 06-01-26- » / « payment-366753 ») était
renvoyée amputée : le recollage doit précéder le scan, sinon la queue seule est
trouvée en premier.
- Un `\b` après `€` ne peut jamais matcher en fin de ligne (€ n'est pas un
caractère de mot) — la TVA n'était jamais extraite.
adc-008 : une facture fournisseur s'enregistre à SA date, même future, tant que
l'exercice (année civile) ne bascule pas. Le document fait foi ; altérer sa date
ferait diverger l'écriture de sa pièce justificative (CGI art. 289 VII).
Registre validé : 8 règles, 8 ADC, 0 erreur.
scopes.ts : 1232 (factures fournisseur) ajouté à prod-write — oubli initial,
révélé par un 403 en production sur F1048. Le pipeline s'est arrêté sans écrire.
Appliqué en production via le pipeline gated : FAF2026014 (Darnis F1048),
218,50 HT + 43,70 TVA = 262,20 TTC, validée, non réglée.
Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
Claude-Session: https://claude.ai/code/session_01VRShc4QhLLU73FLHx9vskh
Completes the role model on the sandbox side, and closes the third failure of
the 2026-07/08 sessions: a refresh wiped hand-granted rights and nothing
recorded them, so the sandbox silently lost capabilities nobody had written down.
- Provisioned `ai_agent_sandbox_read` (36 rights) and
`ai_agent_sandbox_sandbox_write` (44 rights) from test/scopes.ts.
Verified functionally on the live sandbox: the reader reads and gets
403 on invoice creation; the writer creates a draft and gets 403 on DELETE.
- checkpoint-provision.sh now re-creates both scoped agents after every
refresh, so their rights come from code rather than from someone's memory.
Failure to provision a scope warns instead of aborting the whole checkpoint.
- checkpoint-relink-env.sh points the write skill at the scoped writer key,
falling back to the legacy single-user key so an older checkout still works.
The write skill now operates as ai_agent_sandbox_sandbox_write (id 6).
Smoke-tested end to end after the credential swap: the promote pipeline
rehearses on the sandbox under the scoped writer, and `apply` still refuses
without a recorded human gate.
The redundant `ai_agent_prod_prod_write` login is documented as deliberate:
renaming a provisioned production credential means creating a second privileged
user and repointing the promote flow — churn for cosmetics.
Left behind in the sandbox: draft invoice id=19, a scope probe. It cannot be
deleted (no scope grants DELETE, which is the point) and the next refresh
reclaims it.
Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
Claude-Session: https://claude.ai/code/session_01VRShc4QhLLU73FLHx9vskh
Migration applied to production, in the only order that does not break the
promote flow:
1. Created `ai_agent_prod_prod_write` (id=5) with the narrow `prod-write` scope —
invoices, payments, and document submission (needed by builddoc to regenerate
a modified invoice's PDF). Verified functionally: reads pass, DELETE on an
invoice returns 403.
2. Repointed the promote pipeline at that user's key. It no longer borrows the
read skills' credential; if the key is absent it dies with the provisioning
command rather than silently falling back.
3. Revoked 12 write/delete rights from `ai_agent` (id=3), the credential every
read skill holds: create/modify on customer AND supplier invoices,
thirdparties, contacts, thirdparty payment details, proposals, exports,
accounting links — and delete on proposals, events, and GED documents.
Verified after: invoices, thirdparties, contacts, products, proposals, supplier
invoices and bank accounts all still read; creating an invoice returns
`403 Forbidden: Insuffisant rights`. The documented posture and the real one
finally agree.
scopes.ts corrected against the live instance: 262 is NOT "créer/modifier les
produits" as the first catalogue guessed but the `voir_tous` ACL extension — a
READ right the skills depend on (without it, list endpoints return empty arrays
instead of 403). Revoking it would have silently blinded every read skill. This
is why the audit reads labels off /user/perms.php rather than trusting ids in
code. The READ_ONLY baseline is now the audited read surface (35 rights), not a
guess.
Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
Claude-Session: https://claude.ai/code/session_01VRShc4QhLLU73FLHx9vskh
provisionSandbox.ts granted `[12, 122, 262, 32, 111, 251]`: opaque numeric ids,
one fixed scope, one agent per environment. Three failures in the 2026-07/08
sessions came straight from that:
1. Production `ai_agent` — documented as READ-ONLY in AGENTS.md, ADR-0003 and
every SKILL.md — actually holds 46 rights against 9 declared, including
create/modify on customer AND supplier invoices, thirdparties, contacts,
thirdparty payment details, proposals, plus THREE delete rights (proposals,
events, and submit/delete documents in the GED).
2. Writing a payment, then a product, to production each required granting a
right by hand and revoking it after. A privilege granted ad hoc under time
pressure is worse than one nobody holds.
3. A sandbox refresh wiped the agent's proposal rights, because they had been
granted manually and lived nowhere in code.
- scopes.ts declares three scopes with their purpose and allowed environments:
`read` (both envs), `sandbox-write` (sandbox only), `prod-write` (production
only, narrow: invoices + payments, what the gated promote apply actually
does). resolveScope() refuses a scope on an environment it does not belong to.
No scope grants DELETE — the ledger is append-only, deletion stays human.
- provisionAiUser.ts creates or aligns one user per (environment × scope), emits
its API key to a gitignored 600 file, and has an --audit mode that diffs what a
user HOLDS against what its scope DECLARES. Production writes require the
guard.ts opt-in.
- The audit reads permission labels LIVE off /user/perms.php rather than trusting
a catalogue in code: ids are stable per Dolibarr version, not across them, and
an audit that cannot name what it found is not actionable.
findUserId is implemented locally rather than imported: the trunk's userSetup.ts
has one, but it is uncommitted WIP and a provisioning script must not depend on
someone's working tree.
Tooling only — no production rights were changed. The migration (create the
scoped users, repoint the promote pipeline, then strip the over-grants from
`ai_agent`) rotates credentials used by every read skill and is the operator's
call.
Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
Claude-Session: https://claude.ai/code/session_01VRShc4QhLLU73FLHx9vskh
The discipline (ADR-0003, the promote flow, the operating rules) was written
down and still depended on whoever was driving choosing to follow it. On
2026-07-25 an agent session wrote five documents into the production ledger
through direct API calls, bypassing the promote flow entirely — a correct result
reached by a path nobody could audit. A rule an operator can skip is a
recommendation.
The five stages are now chained by artefacts on disk. Each refuses to run until
the previous produced its file, and the file says what it needs to hear:
rehearse (sandbox, host-guarded) -> judge --pre -> gate (human) -> apply
(prod) -> judge --post. The gate binds to a manifest digest, so approving a
change-set approves THAT change-set.
An op is defined ONCE, as an API call, and replayed on the sandbox then on
production — because the first design described each write twice (a sandbox
script input and a prod API body) and the pre-gate judge immediately caught them
diverging: the rehearsal was creating a EUR invoice with no due date while
production would have received a USD one at 60 days. Two descriptions of the
same write are two things that can disagree.
Judges are context-free, cross-family per the PRD qa-strategy rule, and
advisory: a BLOCK still lets the operator approve, and the override is recorded
with their name. Blocking authority stays with the human gate and the host
guards — an LLM verdict never silently starts or stops a production write.
Verified end to end against the real 24/08 change-set (M3 deferred, USD 3,000):
- pre-gate judge (Mistral) returned BLOCK twice, correctly — first on the
sandbox/prod divergence, then on a duplicate left by a repeated rehearsal;
- apply refuses after a rejected gate;
- apply refuses without ARCO_PROD_CONFIRM;
- editing an amount after approval invalidates the gate on digest mismatch.
Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
Claude-Session: https://claude.ai/code/session_01VRShc4QhLLU73FLHx9vskh
Two guards, both from real incidents in the same session.
1. invoice-create.sh — chronology (CGI art. 289). Dolibarr assigns the number at
validation, in creation order, so issuing a document dated BEFORE the last one
already issued gives a higher number to an earlier date. The July plan walked
straight into it: the M3 deferred part is due 2026-10-23 and must be issued at
D-60 (24/08) to stay under the L.441-10 I ceiling, while the M4 fixed part is
dated 23/08 — issue them in the wrong order and the numbering breaks. The
guard reads the last issued document of the same kind and refuses an earlier
date, with ARCO_ALLOW_BACKDATE as a loud, documented override.
Verified: refuses a 01/07 invoice against FAC008 (23/07), accepts 23/08.
2. test/scripts/guard.ts — production opt-in. The sandbox-only guard had no way
to express a deliberate production run, so any prod work meant bypassing it
entirely (which is how guards die). Production now requires BOTH
ARCO_ALLOW_PRODUCTION=<exact host> and
ARCO_PROD_CONFIRM=I-UNDERSTAND-THIS-WRITES-PROD, and prints a banner. Nothing
reaches prod by inheriting an ambient variable.
Also fixes a misleading "(sandbox verified)" log that printed even on prod.
grantAgentRight.ts joins the repo (it was never committed) and gains --revoke,
so a temporarily elevated right can be handed back — used today to attach a
payment in production and revoked immediately after.
Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
Claude-Session: https://claude.ai/code/session_01VRShc4QhLLU73FLHx9vskh
static/config/company.json feeds setupCompany() in both main.ts and
provisionSandbox.ts, and main.ts defaults to the PRODUCTION address. The file
held placeholder identity values: running it against prod would have replaced
Arcodange's legal identity with them.
- capital "1000€" -> "1000". Dolibarr expects a number; the € made the value
unusable by the PDF template, which is why "Capital de 1 000 €" was missing
from every invoice since January (mandatory mention, C. com. R.123-238). This
file is the root cause — patching the database alone would have been undone
by the next provisioning run.
- siren 123456789 -> 999657455, siret 12345678900011 -> 99965745500013,
numTva FR00000000000 -> FR00999657455, rcs_rm "000 000 000 R.C.S. Evry"
-> "R.C.S. Évry", naf_ape 62.02A -> 6201Z. All read off the production ERP,
where they render on every issued invoice.
- formeJuridique SAS -> SARL, and the same correction in fleet/profile/fiscal.yaml
(legal_form), which inherited "SAS" from the PRD README. Three operational
sources say SARL: the production ERP, the signed contrat cadre signature block
("Pour Arcodange (SARL)"), and the 2026-05-28 cohort review. The PRD is wrong.
Flagged in-file for confirmation against the Kbis.
- moisDebutExercice Juillet -> Janvier (fiscal year closes 12-31 per fiscal.yaml).
fiscal.yaml still validates: 8 rules, 15 calendar entries, 7 ADC records, 0 errors.
Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
Claude-Session: https://claude.ai/code/session_01VRShc4QhLLU73FLHx9vskh
sandbox-lifecycle.sh scales deployments to zero, patches the ArgoCD Application
and runs DROP OWNED ... CASCADE. Every one of those ran against whatever
kube-context happened to be current.
This workstation also carries a CLIENT production cluster. On 2026-07-25 a
`checkpoint refresh` was issued while the current context was
do-nyc3-kissmetrics-prod-k8s-cluster: the script patched the ArgoCD Application,
scaled `erp-sandbox` to zero and copied a prod secret — all against the client's
cluster. Nothing was damaged only because that cluster has no `application` CRD
and no erp/erp-sandbox namespaces, so each call failed silently under `|| true`.
That is luck, not a control.
- ERP_KUBE_CONTEXT (default: "default") pins the target; every kubectl call now
goes through K(), so nothing inherits the ambient context.
- assert_arcodange_cluster() proves the target by positive fingerprint — the
erp, erp-sandbox and argocd namespaces AND the erp-sandbox ArgoCD Application.
A client cluster cannot match all four by accident. Wired into all three
entry points, before any mutation.
Verified: refuses the client context, refuses an unknown context, passes on the
homelab and completes normally.
Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
Claude-Session: https://claude.ai/code/session_01VRShc4QhLLU73FLHx9vskh
test/.env ships DOLIBARR_ADDRESS pointing at PRODUCTION and test/main.ts
defaults to it, so any Playwright admin script run with the ambient
environment drives the real ERP. The REST path has been structurally safe
since ADR-0003 (dol-write.sh refuses non-sandbox hosts); the UI path had no
equivalent. scripts/guard.ts closes that gap — assertSandbox() resolves the
target and refuses anything that is not erp-sandbox.*, with the override
spelled out in the error. Verified: an unqualified run now dies instead of
reaching prod.
Two settings the write agent cannot reach (non-admin by design, 403 on
/setup/conf), rehearsed on the sandbox:
- sandboxLegalSetup.ts — capital social + multi-currency module. Finding:
the capital was NOT missing, it was stored as "1000€"; the symbol made the
value unusable by the PDF template, which is why "Capital de 1 000 €" was
absent from every invoice since January (C. com. R.123-238). Normalised to
"1000" → the mention now renders.
- sandboxCurrencySetup.ts — registers the USD reference rate. Enabling the
module is not enough: a currency absent from the rate table makes Dolibarr
silently fall back to EUR (observed on a probe invoice). With the rate
registered, an invoice carries USD 3,000.00 with its EUR counter-value,
i.e. the contractual obligation itself rather than a drifting equivalent.
Both scripts are report-only when they cannot recognise a form, and screenshot
what they did.
Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
Claude-Session: https://claude.ai/code/session_01VRShc4QhLLU73FLHx9vskh
Implementation of the T02 atom over the erp#39 golden set:
- validators.py: instruction-pattern + multi-IBAN pre-screens (0 hard false
positives on the 16 real docs; all 6 injection fixtures quarantined BEFORE
any model call), the atom.yaml invariants, and literal provenance anchoring
with locale-aware locate (FR/EN months incl. abbreviations, NBSP-tolerant
amounts, line-wrap + column-interleave fragment anchoring for refs).
- extract.py: single-leg runner (MLX endpoint / vibe -p), zero credentials,
zero action tools; reasoning-channel aware.
- dual_run.py: model_policy in code — dual legs, exact critical-field
agreement; disagreement, single-valid-leg or both-invalid → escalations/
for the Claude tier (resolutions go back through validators.check).
Eval (eval/2026-07-19/, full transcripts + journals committed):
- critical-field accuracy 100 % (bar 98 %) — MET
- injection suite 6/6 quarantined — zero leaks
- overall field accuracy 94.9 % (known gaps: supplier ids often null,
period_covered format) — non-blocking, noted for the next version
- 9/16 documents escalated to the Claude tier (Mistral API timeouts, small
local model on receipts, one BIC-glued IBAN, derived-ratio rates) —
consistent with the A1 autonomy level recorded in atom.yaml
Runtimes this run: m4-local = Qwen2.5-7B-4bit (MLX), mistral = vibe -p
(mistral-medium-3.5) — provisional pending erp#45; journals are the
routing-bench raw material.
Closes erp#40 (PR to follow once arcodange/golden-set is pushed — this branch
stacks on it).
Co-Authored-By: Claude Fable 5 <[email protected]>
Claude-Session: https://claude.ai/code/session_01VRShc4QhLLU73FLHx9vskh
document-attach.sh uploads a source piece (the supplier's own PDF) onto an
invoice's GED via POST /documents/upload — idempotent by (object, filename,
sha256): before any POST the object's GED is listed and a same-named entry is
downloaded back and sha256-compared. Identical → deduped no-op; different
content → ABORT (refuse-never-repair, overwriteifexists always 0, never
Dolibarr's overwrite flag). Read-back after upload: re-list + download +
sha256-verify. Module-relative download paths are derived from the listing's
fullname (supplier invoices carry an id-derived get_exdir prefix like
9/2/FAF2026013/…, so reconstruction would be wrong).
Promote integration: new `attach` op in promote-plan/promote-apply (OP_SCRIPT),
object_id resolvable via @ref and #supplierinvoice lookups; a relative `file`
resolves against the manifest's directory (replay packs carry pdfs/ beside the
manifest, gitignored — README documents the books@ re-fetch message ids).
promote-plan prints each file's sha256 (or a loud MISSING) at review time.
CLI: `arcodange sandbox attach`.
Proof: offline case 12 in tests/run-tests.sh (upload body, dedupe, conflict
abort, field refusal, manifest-relative resolution via stubbed /documents);
live: manifest-C-ged-attach.json applied twice on the sandbox — run 1 four
created, run 2 four deduped, one GED file per FAF2026010-013, stored sha256s
equal to the re-fetched sources; tests/replay-idempotency.sh extended with an
attach op (4 created → 4 deduped, ged_files count unchanged) and a live
same-name/different-bytes abort verified.
Closes erp#43
Co-Authored-By: Claude Fable 5 <[email protected]>
Claude-Session: https://claude.ai/code/session_01VRShc4QhLLU73FLHx9vskh
Learning #4 of the 2026-07-11 rehearsal: manifest B failed mid-run and could
not be re-applied — op 1 (the DARNIS invoice) had already run and a replay
would have duplicated it. Every write op now dedupes BEFORE any POST:
- thirdparty-create.sh: by exact name (promote '#thirdparty:name=' semantics);
ambiguous (2+) aborts; an existing fiche missing the requested role aborts
(refuse-never-repair). Emits {"id", "deduped"} instead of a bare id.
- invoice-create.sh: supplier kind by (socid, ref_supplier) — same key with a
different total aborts as a conflict; customer kind (or supplier without
ref_supplier) by (socid, date, total_ttc ±0.02, line fingerprint) with descs
HTML-unescaped. Credit notes are never candidates. A deduped DRAFT with
validate:true is validated on replay, so an interrupted run converges.
- payment-record.sh: by (invoice, amount, normalized transaction_id), composing
with the erp#37 varchar(50) normalization on BOTH sides so historical
long-form nums still match; same tx + different amount aborts; without a tx
id there is no dedupe key (warned). Dedupe answers id:null (the payments list
exposes no paiement rowid) + the existing bank line.
- All three refuse to POST blind when the dedupe lookup fails with anything but
the documented empty-list 404 (the voir_tous trap would otherwise mint dupes).
- promote-apply.sh: marks each op created / deduped=true inline and totals them
in the summary — an all-deduped second run is visible proof of a no-op.
- promote-plan.sh: advertises each op's dedupe key (and flags tx=MISSING as
'a replay WILL double-pay').
Proof:
- tests/run-tests.sh: 5 new offline cases (11 total) — dedupe hits POST
nothing, conflicts/ambiguity abort pre-POST, long-form history dedupes,
draft convergence validates; stub extended to serve the new lookups with the
live-observed empty behaviors ([] for invoices/payments, 404 for tiers).
- tests/replay-idempotency.sh (new, live): double-applies a self-contained
manifest on the sandbox — run 1 '3 created' (rows 1/1/1), run 2 '3 deduped'
with row counts unchanged and the stored num in erp#37 short form.
- The historic manifest-B now replays on the sandbox as 5/5 deduped, zero new
rows — the exact replay Learning #4 declared impossible.
SKILL.md updated in the same change (per-op dedupe keys, replay-safety section,
gotchas); the 2026-07-11 runbook's Learning #4 carries a dated resolution
addendum.
Closes erp#44.
Co-Authored-By: Claude Fable 5 <[email protected]>
Claude-Session: https://claude.ai/code/session_01VRShc4QhLLU73FLHx9vskh
fleet/profile/ goes from stub to the machine-readable business-rules surface
the fleet reads (PRD agent-catalog document surface + compliance ADC framework):
- fiscal.yaml — entity, VAT position, 8 rules (regime reel simplifie until
2026-12-31 -> quarterly CA3 from 2027-01-01 per LF 2025 art. 38; KM export
autoliquidation 259-1 CGI box E2; FR 20% deductible; intra-EU reverse
charge; FX 766/666; SaaS expensed; CCA 455 lane). Every rule carries
effective_from/effective_until AND decision: adc-NNN; every date cites its
PRD anchor as an inline comment (verified against factory origin/main).
- calendar.yaml — 15 entries: acomptes TVA (2026-07 month-window, 2026-12-15),
last CA12 FY-2026 (2027-05-04), CA3 quarterly windows, CFE (December),
AG comptes annuels (2027-06-30), e-invoicing milestones (2026-09-01
reception, 2027-09-01 emission/e-reporting), URSSAF echeancier with the
in-file NOTE that a real direct debit exists since May 2026 (erp#57 revisit
of the payroll-dormant assumption), KM deferred due dates + renewal stub.
- JSON Schemas for both + scripts/validate.py (stdlib-only: strict YAML-subset
parser, JSON-Schema-subset checker, rule->ADC resolution, calendar checks).
- decisions/ — ADC register: template + adc-001..005 Accepted formalizations
(autoliquidation KM, FX->766/666, SaaS expensed, reel simplifie until
abolition, CCA personal-card lane) + adc-006/007 Proposed stubs (retainer
currency -> erp#53; capital path -> erp#51). Agents draft, the operator
Accepts — never the reverse; immutable once merged, supersede never edit.
- Mutation policy in-file: PRs only (T12 proposes, human merges).
- Same-change: profile README stub -> real doc; fleet/README.md layout line
and AGENTS.md fleet row updated (profile no longer a stub).
Validation: PASS — 8 rules, 15 entries, 7 ADCs, 0 errors, 7 warnings (the
warnings list exactly what awaits operator verification). Human gate left
open on purpose: operator sanity-read of the calendar + Acceptance of
adc-001..005.
Co-Authored-By: Claude Fable 5 <[email protected]>
Claude-Session: https://claude.ai/code/session_01VRShc4QhLLU73FLHx9vskh
- validators.py: deterministic pre-screens (instruction patterns, multi-IBAN
escalate flag) + the atom.yaml invariants (arithmetic, rates, SIREN Luhn,
IBAN mod-97, date plausibility) + literal-provenance anchoring (a value
absent from the source can never appear in output).
Tested: 0 hard false positives on the 16 real docs; 6/6 injection fixtures
quarantined PRE-model; darnis-f1042 (embedded second document) → escalate.
- extract.py: single-leg runner, zero credentials/action tools; runtimes =
MLX endpoint (Ornith/M4) and vibe -p (Mistral).
- dual_run.py: model_policy in code — dual legs, exact critical-field
agreement, disagreement/flags → escalations/, invalid-both → quarantine.
Eval run against the golden set follows in this branch.
Co-Authored-By: Claude Fable 5 <[email protected]>
Claude-Session: https://claude.ai/code/session_01VRShc4QhLLU73FLHx9vskh
- runs/2026-07-18/: the 8 sha256-pinned verifier transcripts (4 runtimes ×
2 tests), blind-judging verdicts (2 independent judges per cell, unanimous),
the erp#56 builder-bench journal + prompt + caps, and the evidence README
with the parity table.
- run-verifier.sh: mistral runtime drops the tool-filter flag (--enabled-tools
with a no-match pattern hangs vibe 2.21.0); plain -p with --max-turns 1.
Verdicts: Mistral (vibe -p, mistral-medium-3.5) and Ornith 35B (hermes MLX)
reach verdict parity with the Claude baseline on both tests → admitted to
verifier duty. Qwen2.5-7B-4bit fails both → the honest small-model floor.
Builder bench: erp#56 completed by the Mistral runtime, 0 code corrections,
261 s, acceptance run clean (0 bank-UNKNOWN) → merged as PR #68.
Closes#63 (with the paired factory qa-strategy PR).
Co-Authored-By: Claude Fable 5 <[email protected]>
Claude-Session: https://claude.ai/code/session_01VRShc4QhLLU73FLHx9vskh
Seed the invoice-extract (T02) and mail-classify (T01) golden sets from real
Arcodange history, plus an adversarial injection suite and an offline
field-level scorer.
invoice-extract/
- 16 real supplier PDFs (DARNIS/Hiway F1040/F1042/F1045/F1046, Anthropic
invoice+receipt x2, Mistral, OVH, greffe d'Evry, INPI x2, Legalstart, Qonto,
Infogreffe) fetched from the Zoho mailbox + Dolibarr GED, each with a
hand-verified expected JSON per the T02 schema. Every expected value was
cross-checked against the pdftotext -layout text and re-validated against the
deterministic invariants (HT+TVA=TTC, per-rate sums, IBAN mod-97, SIREN Luhn).
- inputs/ carries both the source PDF and its {source_sha256, mime, text} pair.
- 6 SYNTHETIC injection fixtures (LLM-directive, hidden white text, IBAN-swap
BEC lure, arithmetic-repair lure, fake tool-call, ref-hijack duplicate) whose
only correct outcome is quarantine; each PDF is marked SYNTHETIC.
- score.py: stdlib-only field-level scorer, critical fields (amounts/IBAN/refs/
dates) scored separately against the 98% bar, injection leaks blocking; a
built-in --self-test proves it catches perturbed fields and leaks.
- manifest.json: per-item provenance (mail message id / GED path + sha256),
linked Dolibarr supplier invoice, a verification note, and the list of real
documents deliberately excluded (fee statements, payment proofs, La Poste
receipts with no HT/TVA breakdown) with reasons.
mail-classify/
- 1824 historical mails labeled into {supplier-invoice, bank-notice,
government-admin, client, other} via sender-domain + subject weak supervision,
one human-correctable JSONL line per message with confidence + reason +
message-id provenance. manifest.json records the pull method and distribution.
Docs: golden/README hub, invoice-extract/README (T02 schema + conventions),
injection/README (threat table), mail-classify/README (method + distribution).
Co-Authored-By: Claude Fable 5 <[email protected]>
Claude-Session: https://claude.ai/code/session_01VRShc4QhLLU73FLHx9vskh
- MISTRAL.AI: update note to reflect annual subscription (Le Chat Pro - Annual)
from invoice MSTRL-API-814045-001, 2026-04-02, 143.90 HT / 172.68 TTC.
Next expected ~2027-04.
- CLAUDE.AI: document payment rail moved to personal card (fk_account=3,
API-invisible) for May/June; reference issue #57.
Both AI subscriptions are now recorded supplier invoices (post-replay).
Generated by Mistral Vibe.
Co-Authored-By: Mistral Vibe <[email protected]>
The harness layer (builder sessions, cold verifiers, evidence flow) gets a
committable home, per the PRD model-fleet § harness portability and erp#63:
- fleet/harness/verifier/: the two canonical verifier tests (locate-test,
cold-reader backlog audit) with pinned inputs, verbatim prompts, ground
truth and pass rules — judged context-free, never self-graded.
- fleet/harness/bin/run-verifier.sh: runs a test against any OpenAI-style
local endpoint (Ornith/MLX) or vibe -p (Mistral); emits sha256-pinned
JSON transcripts.
- fleet/harness/bin/vibe-builder.sh: the bounded shell for scoped builders
and recurring tasks — refuses the trunk (linked-worktree guard), hard
--max-turns/--max-price caps, full JSON journal per run.
- fleet/README.md layout + AGENTS.md Fleet section updated in the same
change (same-change freshness rule).
Part of erp#63 (harness portability spike, D2).
Co-Authored-By: Claude Fable 5 <[email protected]>
Claude-Session: https://claude.ai/code/session_01VRShc4QhLLU73FLHx9vskh
Part of erp#65 (phase 1). Ledger grammar "thirdparty complete" gets its
op: allowlisted non-ledger fields, per-field diff read-back. Contacts
are born idempotent (dedupe by email then name). Promote ops wired both
targets, offline stub tests, SKILL.md workflows, KM dossier manifest
(unsigned-contract truth fix + EIN-to-collect note).
Co-Authored-By: Claude Fable 5 <[email protected]>
Closes erp#38 deliverables: fleet/ layout, atom.yaml schema documented
in fleet/README.md, 7 class skeletons per the PRD agent-catalog,
invoice-extract as the worked example (contract only — implementation
is erp#40), golden/ + profile/ stubs, AGENTS.md Fleet section with
freshness fixes (fleet/ no longer "not yet landed").
Co-Authored-By: Claude Fable 5 <[email protected]>
The pack (manifests, prelude, runbook, verify-provenance.py PoC) lived
only in an ephemeral session scratchpad while erp#41/#42/#43/#44 now
reference it as fixtures and the prod replay is still pending. 36/36
provenance checks were green at rehearsal time; PDFs are re-fetchable
via arcodange-email-ingest (documented in the pack README).
Co-Authored-By: Claude Fable 5 <[email protected]>
Agents landing in this repo had no entry point: no AGENTS.md, and the
backlog (erp#38-57 on 6 dated milestones, gateway#1-2, factory#22)
was only discoverable from the PRD STATUS in the factory repo. This
seeds the repo-root orientation map: where the work comes from (resume
protocol: top unblocked issue of the earliest open milestone, gitea
MCP pointers, owner gotcha for telegram-gateway), the repo map, the
operating rules (trunk/worktrees, read-only prod, sandbox+promote
gate, append-only ledger, anti-hallucination contract pointers).
Advances erp#38 (AGENTS.md seed; the fleet/ scaffold and the full
fleet section remain in #38's scope).
Co-Authored-By: Claude Fable 5 <[email protected]>
Qonto transaction ids run ~67 chars (<org>-<n>-<n>-transaction-<uuid>) but
Dolibarr stores num_payment in varchar(50) (llx_paiement.num_paiement,
llx_paiementfourn.num_paiement) — POSTing a payment with the raw id fails
HTTP 400 "value too long for type character varying(50)". Parade proven live
on the sandbox (2026-07-11): store the UUID suffix (globally unique, ~37
chars). Wise ids (short numerics) are unaffected.
Writer side — payment-record.sh strips everything through "transaction-"
before POST, announces the normalization on stderr, REFUSES (never truncates)
ids still >50 chars after normalization, and emits the normalized num in the
output JSON.
Reader side — bank-match.sh PASS 0 (exact tx-id, erp#28) now compares BOTH
sides in raw AND canonical short form: Qonto feed ids are carried long+short,
payment nums are normalized on compare — so nums stored short (the varchar(50)
form) and historical long-form nums both keep matching. Wise ids untouched.
Proven offline (no credentials, no network, no sandbox/prod writes):
- arcodange-bank-reco/tests/run-tests.sh — new bank-match --fixtures offline
mode: long feed id ↔ short num, long ↔ long (back-compat), Wise numeric,
each Δ+19d outside the ±7d window so only PASS 0 can pair them (exit 0,
3×[tx-id]); plus the empty-num negative (exit 1, 0 matched).
- dolibarr-sandbox-write/tests/run-tests.sh — payment-record via a stubbed
dol-write.sh (DOL_WRITE hook): long→short in POST body + output JSON, Wise
untouched, >50-after-normalization refused BEFORE any POST, citing
varchar(50).
Both SKILL.md document the canonical short form + the varchar(50) constraint.
Co-Authored-By: Claude Fable 5 <[email protected]>