feat(fleet): invoice-extract atom — dual extraction + validators + provenance (erp#40)

Implementation of the T02 atom over the erp#39 golden set:

- validators.py: instruction-pattern + multi-IBAN pre-screens (0 hard false
  positives on the 16 real docs; all 6 injection fixtures quarantined BEFORE
  any model call), the atom.yaml invariants, and literal provenance anchoring
  with locale-aware locate (FR/EN months incl. abbreviations, NBSP-tolerant
  amounts, line-wrap + column-interleave fragment anchoring for refs).
- extract.py: single-leg runner (MLX endpoint / vibe -p), zero credentials,
  zero action tools; reasoning-channel aware.
- dual_run.py: model_policy in code — dual legs, exact critical-field
  agreement; disagreement, single-valid-leg or both-invalid → escalations/
  for the Claude tier (resolutions go back through validators.check).

Eval (eval/2026-07-19/, full transcripts + journals committed):
- critical-field accuracy 100 % (bar 98 %) — MET
- injection suite 6/6 quarantined — zero leaks
- overall field accuracy 94.9 % (known gaps: supplier ids often null,
  period_covered format) — non-blocking, noted for the next version
- 9/16 documents escalated to the Claude tier (Mistral API timeouts, small
  local model on receipts, one BIC-glued IBAN, derived-ratio rates) —
  consistent with the A1 autonomy level recorded in atom.yaml

Runtimes this run: m4-local = Qwen2.5-7B-4bit (MLX), mistral = vibe -p
(mistral-medium-3.5) — provisional pending erp#45; journals are the
routing-bench raw material.

Closes erp#40 (PR to follow once arcodange/golden-set is pushed — this branch
stacks on it).

Co-Authored-By: Claude Fable 5 <[email protected]>
Claude-Session: https://claude.ai/code/session_01VRShc4QhLLU73FLHx9vskh
This commit is contained in:
2026-07-19 00:31:21 +02:00
co-authored by Claude Fable 5
parent bdd3d63b61
commit e9d4a2bcb2
38 changed files with 1678 additions and 44 deletions
+41 -7
View File
@@ -1,8 +1,42 @@
# invoice-extract/scripts — intentionally empty
# invoice-extractscripts (erp#40)
The implementation — dual-run extraction drivers (M4 local ∥ Mistral API),
deterministic validators (arithmetic, rates, SIREN/IBAN checksums, dedupe,
provenance re-verification), the stubbed OCR fallback and the scoring hooks —
lands with [erp#40](https://gitea.arcodange.lab/arcodange-org/erp/issues/40).
This scaffold ships the contract only ([`../atom.yaml`](../atom.yaml)); do not
fake extraction code here.
The deterministic implementation around the [atom contract](../atom.yaml). The
LLM proposes, this code disposes; a failed check refuses, never repairs.
| File | Role |
| --- | --- |
| `validators.py` | pre-screens (instruction patterns → quarantine; multi-IBAN → escalate flag) + the `atom.yaml` invariants (arithmetic, rate whitelist, SIREN Luhn, IBAN mod-97, date plausibility) + **literal provenance anchoring**: every critical value must be locatable verbatim in the source text (locale-aware) or the leg fails — a value absent from its source can never appear in output |
| `extract.py` | single-leg runner, zero credentials, zero action tools. Runtimes: `mlx` (OpenAI-style local endpoint, default `127.0.0.1:18080` — hermes MLX; handles reasoning-channel models) and `vibe` (Mistral via `vibe -p`, the harness's admitted runtime) |
| `dual_run.py` | the `model_policy` in code: pre-screen → two independent legs → validators per leg → **exact critical-field agreement** required. Disagreement, single-valid-leg, escalate flag, or both-legs-invalid → `escalations/` for the **Claude tier**, whose resolution goes back through `validators.check` (same bar) and may itself be a quarantine. Hostile documents never reach a model |
## Running the eval
```bash
python3 scripts/dual_run.py \
--inputs ../../golden/invoice-extract/inputs \
--injection ../../golden/invoice-extract/injection/inputs \
--out /tmp/run/predicted --journal /tmp/run/journal.jsonl \
--mlx-model mlx-community/Qwen2.5-7B-Instruct-4bit
# escalations resolved (Claude tier, through validators.check), then:
python3 ../../golden/invoice-extract/score.py --predicted /tmp/run/predicted
```
`score.py` (the golden set's scorer) owns the verdict: critical-field bar 98 %,
any injection leak is blocking. The dual-run journal records every leg (runtime,
model, latency, invariant failures) — it is the routing-bench raw material for
erp#45.
## Model notes (provisional until erp#45 closes D5/model_policy)
- Local leg: `Qwen2.5-7B-Instruct-4bit` (resident on the M4) — fast (~3-5 s/doc),
weaker grounding on receipts; its misses surface as escalations, never as
silent output (the validators see to that).
- `Ornith-1.0-35B` emits a `reasoning` channel that consumes the token budget
before `content`; usable with `max_tokens ≥ 8000` at minutes-per-doc latency —
benched properly in erp#45.
- Mistral leg: `vibe -p` (`mistral-medium-3.5`, thinking on) — ~15-100 s/doc,
strong grounding; JSON shape is prompt-enforced + parsed defensively
(constrained decoding is not exposed through the CLI; the validators guarantee
truth conditions regardless).
- OCR fallback for scanned inputs: stubbed — provider choice is D5 (erp#45).