Seed the invoice-extract (T02) and mail-classify (T01) golden sets from real
Arcodange history, plus an adversarial injection suite and an offline
field-level scorer.
invoice-extract/
- 16 real supplier PDFs (DARNIS/Hiway F1040/F1042/F1045/F1046, Anthropic
invoice+receipt x2, Mistral, OVH, greffe d'Evry, INPI x2, Legalstart, Qonto,
Infogreffe) fetched from the Zoho mailbox + Dolibarr GED, each with a
hand-verified expected JSON per the T02 schema. Every expected value was
cross-checked against the pdftotext -layout text and re-validated against the
deterministic invariants (HT+TVA=TTC, per-rate sums, IBAN mod-97, SIREN Luhn).
- inputs/ carries both the source PDF and its {source_sha256, mime, text} pair.
- 6 SYNTHETIC injection fixtures (LLM-directive, hidden white text, IBAN-swap
BEC lure, arithmetic-repair lure, fake tool-call, ref-hijack duplicate) whose
only correct outcome is quarantine; each PDF is marked SYNTHETIC.
- score.py: stdlib-only field-level scorer, critical fields (amounts/IBAN/refs/
dates) scored separately against the 98% bar, injection leaks blocking; a
built-in --self-test proves it catches perturbed fields and leaks.
- manifest.json: per-item provenance (mail message id / GED path + sha256),
linked Dolibarr supplier invoice, a verification note, and the list of real
documents deliberately excluded (fee statements, payment proofs, La Poste
receipts with no HT/TVA breakdown) with reasons.
mail-classify/
- 1824 historical mails labeled into {supplier-invoice, bank-notice,
government-admin, client, other} via sender-domain + subject weak supervision,
one human-correctable JSONL line per message with confidence + reason +
message-id provenance. manifest.json records the pull method and distribution.
Docs: golden/README hub, invoice-extract/README (T02 schema + conventions),
injection/README (threat table), mail-classify/README (method + distribution).
Co-Authored-By: Claude Fable 5 <[email protected]>
Claude-Session: https://claude.ai/code/session_01VRShc4QhLLU73FLHx9vskh
2.0 KiB
2.0 KiB
fleet/golden/ — per-atom golden sets
fleet > golden
Per-atom golden sets — <atom>/{inputs,expected}/ + a field-level scoring script,
adversarial injection fixtures where the atom reads untrusted content. Seeded from
real Arcodange history (the 2026 mailbox, every recorded supplier invoice, the
GED) per the PRD golden datasets.
Landed with erp#39.
Sets
| Set | Atom / task | Items | What |
|---|---|---|---|
invoice-extract/ |
invoice-extract (T02) |
16 real + 6 injection | supplier PDFs → hand-verified T02 JSON; adversarial quarantine suite; score.py |
mail-classify/ |
mailbox triage (T01) | 1824 labeled | historical mail labeled into the 5 T01 classes, human-correctable JSONL |
Principles (shared)
- Every real item is a test case. Volumes are small, so the set is the history, not a sample of it.
- Field-level scoring, not document-level: a 9/10-field extraction is a failed document but 90 % field accuracy — both are tracked. Critical fields (amounts, IBAN, refs, dates) are scored separately and hold the 98 % bar.
- Hand-verified ground truth. Expected values are checked against the source
text; a value the document does not state is
null, never a guess. - Provenance per item. Each set's
manifest.json(or the JSONL's per-linemessage_id) records the source id (mail message id / GED path) + sha256 of the source file, so any label is traceable to its origin. - Adversarial fixtures are clearly synthetic and their only correct outcome is quarantine; a single injection leak is a blocking failure regardless of accuracy.
- The set grows as a by-product of operation — every human correction, rejection reason and reclassification is captured back into it.