Files
arcodangeandClaude Fable 5 e9d4a2bcb2 feat(fleet): invoice-extract atom — dual extraction + validators + provenance (erp#40)
Implementation of the T02 atom over the erp#39 golden set:

- validators.py: instruction-pattern + multi-IBAN pre-screens (0 hard false
  positives on the 16 real docs; all 6 injection fixtures quarantined BEFORE
  any model call), the atom.yaml invariants, and literal provenance anchoring
  with locale-aware locate (FR/EN months incl. abbreviations, NBSP-tolerant
  amounts, line-wrap + column-interleave fragment anchoring for refs).
- extract.py: single-leg runner (MLX endpoint / vibe -p), zero credentials,
  zero action tools; reasoning-channel aware.
- dual_run.py: model_policy in code — dual legs, exact critical-field
  agreement; disagreement, single-valid-leg or both-invalid → escalations/
  for the Claude tier (resolutions go back through validators.check).

Eval (eval/2026-07-19/, full transcripts + journals committed):
- critical-field accuracy 100 % (bar 98 %) — MET
- injection suite 6/6 quarantined — zero leaks
- overall field accuracy 94.9 % (known gaps: supplier ids often null,
  period_covered format) — non-blocking, noted for the next version
- 9/16 documents escalated to the Claude tier (Mistral API timeouts, small
  local model on receipts, one BIC-glued IBAN, derived-ratio rates) —
  consistent with the A1 autonomy level recorded in atom.yaml

Runtimes this run: m4-local = Qwen2.5-7B-4bit (MLX), mistral = vibe -p
(mistral-medium-3.5) — provisional pending erp#45; journals are the
routing-bench raw material.

Closes erp#40 (PR to follow once arcodange/golden-set is pushed — this branch
stacks on it).

Co-Authored-By: Claude Fable 5 <[email protected]>
Claude-Session: https://claude.ai/code/session_01VRShc4QhLLU73FLHx9vskh
2026-07-19 00:31:21 +02:00
..

invoice-extract — scripts (erp#40)

The deterministic implementation around the atom contract. The LLM proposes, this code disposes; a failed check refuses, never repairs.

File Role
validators.py pre-screens (instruction patterns → quarantine; multi-IBAN → escalate flag) + the atom.yaml invariants (arithmetic, rate whitelist, SIREN Luhn, IBAN mod-97, date plausibility) + literal provenance anchoring: every critical value must be locatable verbatim in the source text (locale-aware) or the leg fails — a value absent from its source can never appear in output
extract.py single-leg runner, zero credentials, zero action tools. Runtimes: mlx (OpenAI-style local endpoint, default 127.0.0.1:18080 — hermes MLX; handles reasoning-channel models) and vibe (Mistral via vibe -p, the harness's admitted runtime)
dual_run.py the model_policy in code: pre-screen → two independent legs → validators per leg → exact critical-field agreement required. Disagreement, single-valid-leg, escalate flag, or both-legs-invalid → escalations/ for the Claude tier, whose resolution goes back through validators.check (same bar) and may itself be a quarantine. Hostile documents never reach a model

Running the eval

python3 scripts/dual_run.py \
  --inputs  ../../golden/invoice-extract/inputs \
  --injection ../../golden/invoice-extract/injection/inputs \
  --out /tmp/run/predicted --journal /tmp/run/journal.jsonl \
  --mlx-model mlx-community/Qwen2.5-7B-Instruct-4bit

# escalations resolved (Claude tier, through validators.check), then:
python3 ../../golden/invoice-extract/score.py --predicted /tmp/run/predicted

score.py (the golden set's scorer) owns the verdict: critical-field bar 98 %, any injection leak is blocking. The dual-run journal records every leg (runtime, model, latency, invariant failures) — it is the routing-bench raw material for erp#45.

Model notes (provisional until erp#45 closes D5/model_policy)

  • Local leg: Qwen2.5-7B-Instruct-4bit (resident on the M4) — fast (~3-5 s/doc), weaker grounding on receipts; its misses surface as escalations, never as silent output (the validators see to that).
  • Ornith-1.0-35B emits a reasoning channel that consumes the token budget before content; usable with max_tokens ≥ 8000 at minutes-per-doc latency — benched properly in erp#45.
  • Mistral leg: vibe -p (mistral-medium-3.5, thinking on) — ~15-100 s/doc, strong grounding; JSON shape is prompt-enforced + parsed defensively (constrained decoding is not exposed through the CLI; the validators guarantee truth conditions regardless).
  • OCR fallback for scanned inputs: stubbed — provider choice is D5 (erp#45).