Implementation of the T02 atom over the erp#39 golden set:
- validators.py: instruction-pattern + multi-IBAN pre-screens (0 hard false
positives on the 16 real docs; all 6 injection fixtures quarantined BEFORE
any model call), the atom.yaml invariants, and literal provenance anchoring
with locale-aware locate (FR/EN months incl. abbreviations, NBSP-tolerant
amounts, line-wrap + column-interleave fragment anchoring for refs).
- extract.py: single-leg runner (MLX endpoint / vibe -p), zero credentials,
zero action tools; reasoning-channel aware.
- dual_run.py: model_policy in code — dual legs, exact critical-field
agreement; disagreement, single-valid-leg or both-invalid → escalations/
for the Claude tier (resolutions go back through validators.check).
Eval (eval/2026-07-19/, full transcripts + journals committed):
- critical-field accuracy 100 % (bar 98 %) — MET
- injection suite 6/6 quarantined — zero leaks
- overall field accuracy 94.9 % (known gaps: supplier ids often null,
period_covered format) — non-blocking, noted for the next version
- 9/16 documents escalated to the Claude tier (Mistral API timeouts, small
local model on receipts, one BIC-glued IBAN, derived-ratio rates) —
consistent with the A1 autonomy level recorded in atom.yaml
Runtimes this run: m4-local = Qwen2.5-7B-4bit (MLX), mistral = vibe -p
(mistral-medium-3.5) — provisional pending erp#45; journals are the
routing-bench raw material.
Closes erp#40 (PR to follow once arcodange/golden-set is pushed — this branch
stacks on it).
Co-Authored-By: Claude Fable 5 <[email protected]>
Claude-Session: https://claude.ai/code/session_01VRShc4QhLLU73FLHx9vskh
fleet/profile/ goes from stub to the machine-readable business-rules surface
the fleet reads (PRD agent-catalog document surface + compliance ADC framework):
- fiscal.yaml — entity, VAT position, 8 rules (regime reel simplifie until
2026-12-31 -> quarterly CA3 from 2027-01-01 per LF 2025 art. 38; KM export
autoliquidation 259-1 CGI box E2; FR 20% deductible; intra-EU reverse
charge; FX 766/666; SaaS expensed; CCA 455 lane). Every rule carries
effective_from/effective_until AND decision: adc-NNN; every date cites its
PRD anchor as an inline comment (verified against factory origin/main).
- calendar.yaml — 15 entries: acomptes TVA (2026-07 month-window, 2026-12-15),
last CA12 FY-2026 (2027-05-04), CA3 quarterly windows, CFE (December),
AG comptes annuels (2027-06-30), e-invoicing milestones (2026-09-01
reception, 2027-09-01 emission/e-reporting), URSSAF echeancier with the
in-file NOTE that a real direct debit exists since May 2026 (erp#57 revisit
of the payroll-dormant assumption), KM deferred due dates + renewal stub.
- JSON Schemas for both + scripts/validate.py (stdlib-only: strict YAML-subset
parser, JSON-Schema-subset checker, rule->ADC resolution, calendar checks).
- decisions/ — ADC register: template + adc-001..005 Accepted formalizations
(autoliquidation KM, FX->766/666, SaaS expensed, reel simplifie until
abolition, CCA personal-card lane) + adc-006/007 Proposed stubs (retainer
currency -> erp#53; capital path -> erp#51). Agents draft, the operator
Accepts — never the reverse; immutable once merged, supersede never edit.
- Mutation policy in-file: PRs only (T12 proposes, human merges).
- Same-change: profile README stub -> real doc; fleet/README.md layout line
and AGENTS.md fleet row updated (profile no longer a stub).
Validation: PASS — 8 rules, 15 entries, 7 ADCs, 0 errors, 7 warnings (the
warnings list exactly what awaits operator verification). Human gate left
open on purpose: operator sanity-read of the calendar + Acceptance of
adc-001..005.
Co-Authored-By: Claude Fable 5 <[email protected]>
Claude-Session: https://claude.ai/code/session_01VRShc4QhLLU73FLHx9vskh
- validators.py: deterministic pre-screens (instruction patterns, multi-IBAN
escalate flag) + the atom.yaml invariants (arithmetic, rates, SIREN Luhn,
IBAN mod-97, date plausibility) + literal-provenance anchoring (a value
absent from the source can never appear in output).
Tested: 0 hard false positives on the 16 real docs; 6/6 injection fixtures
quarantined PRE-model; darnis-f1042 (embedded second document) → escalate.
- extract.py: single-leg runner, zero credentials/action tools; runtimes =
MLX endpoint (Ornith/M4) and vibe -p (Mistral).
- dual_run.py: model_policy in code — dual legs, exact critical-field
agreement, disagreement/flags → escalations/, invalid-both → quarantine.
Eval run against the golden set follows in this branch.
Co-Authored-By: Claude Fable 5 <[email protected]>
Claude-Session: https://claude.ai/code/session_01VRShc4QhLLU73FLHx9vskh
- runs/2026-07-18/: the 8 sha256-pinned verifier transcripts (4 runtimes ×
2 tests), blind-judging verdicts (2 independent judges per cell, unanimous),
the erp#56 builder-bench journal + prompt + caps, and the evidence README
with the parity table.
- run-verifier.sh: mistral runtime drops the tool-filter flag (--enabled-tools
with a no-match pattern hangs vibe 2.21.0); plain -p with --max-turns 1.
Verdicts: Mistral (vibe -p, mistral-medium-3.5) and Ornith 35B (hermes MLX)
reach verdict parity with the Claude baseline on both tests → admitted to
verifier duty. Qwen2.5-7B-4bit fails both → the honest small-model floor.
Builder bench: erp#56 completed by the Mistral runtime, 0 code corrections,
261 s, acceptance run clean (0 bank-UNKNOWN) → merged as PR #68.
Closes#63 (with the paired factory qa-strategy PR).
Co-Authored-By: Claude Fable 5 <[email protected]>
Claude-Session: https://claude.ai/code/session_01VRShc4QhLLU73FLHx9vskh
Seed the invoice-extract (T02) and mail-classify (T01) golden sets from real
Arcodange history, plus an adversarial injection suite and an offline
field-level scorer.
invoice-extract/
- 16 real supplier PDFs (DARNIS/Hiway F1040/F1042/F1045/F1046, Anthropic
invoice+receipt x2, Mistral, OVH, greffe d'Evry, INPI x2, Legalstart, Qonto,
Infogreffe) fetched from the Zoho mailbox + Dolibarr GED, each with a
hand-verified expected JSON per the T02 schema. Every expected value was
cross-checked against the pdftotext -layout text and re-validated against the
deterministic invariants (HT+TVA=TTC, per-rate sums, IBAN mod-97, SIREN Luhn).
- inputs/ carries both the source PDF and its {source_sha256, mime, text} pair.
- 6 SYNTHETIC injection fixtures (LLM-directive, hidden white text, IBAN-swap
BEC lure, arithmetic-repair lure, fake tool-call, ref-hijack duplicate) whose
only correct outcome is quarantine; each PDF is marked SYNTHETIC.
- score.py: stdlib-only field-level scorer, critical fields (amounts/IBAN/refs/
dates) scored separately against the 98% bar, injection leaks blocking; a
built-in --self-test proves it catches perturbed fields and leaks.
- manifest.json: per-item provenance (mail message id / GED path + sha256),
linked Dolibarr supplier invoice, a verification note, and the list of real
documents deliberately excluded (fee statements, payment proofs, La Poste
receipts with no HT/TVA breakdown) with reasons.
mail-classify/
- 1824 historical mails labeled into {supplier-invoice, bank-notice,
government-admin, client, other} via sender-domain + subject weak supervision,
one human-correctable JSONL line per message with confidence + reason +
message-id provenance. manifest.json records the pull method and distribution.
Docs: golden/README hub, invoice-extract/README (T02 schema + conventions),
injection/README (threat table), mail-classify/README (method + distribution).
Co-Authored-By: Claude Fable 5 <[email protected]>
Claude-Session: https://claude.ai/code/session_01VRShc4QhLLU73FLHx9vskh
The harness layer (builder sessions, cold verifiers, evidence flow) gets a
committable home, per the PRD model-fleet § harness portability and erp#63:
- fleet/harness/verifier/: the two canonical verifier tests (locate-test,
cold-reader backlog audit) with pinned inputs, verbatim prompts, ground
truth and pass rules — judged context-free, never self-graded.
- fleet/harness/bin/run-verifier.sh: runs a test against any OpenAI-style
local endpoint (Ornith/MLX) or vibe -p (Mistral); emits sha256-pinned
JSON transcripts.
- fleet/harness/bin/vibe-builder.sh: the bounded shell for scoped builders
and recurring tasks — refuses the trunk (linked-worktree guard), hard
--max-turns/--max-price caps, full JSON journal per run.
- fleet/README.md layout + AGENTS.md Fleet section updated in the same
change (same-change freshness rule).
Part of erp#63 (harness portability spike, D2).
Co-Authored-By: Claude Fable 5 <[email protected]>
Claude-Session: https://claude.ai/code/session_01VRShc4QhLLU73FLHx9vskh
Closes erp#38 deliverables: fleet/ layout, atom.yaml schema documented
in fleet/README.md, 7 class skeletons per the PRD agent-catalog,
invoice-extract as the worked example (contract only — implementation
is erp#40), golden/ + profile/ stubs, AGENTS.md Fleet section with
freshness fixes (fleet/ no longer "not yet landed").
Co-Authored-By: Claude Fable 5 <[email protected]>