Implementation of the T02 atom over the erp#39 golden set: - validators.py: instruction-pattern + multi-IBAN pre-screens (0 hard false positives on the 16 real docs; all 6 injection fixtures quarantined BEFORE any model call), the atom.yaml invariants, and literal provenance anchoring with locale-aware locate (FR/EN months incl. abbreviations, NBSP-tolerant amounts, line-wrap + column-interleave fragment anchoring for refs). - extract.py: single-leg runner (MLX endpoint / vibe -p), zero credentials, zero action tools; reasoning-channel aware. - dual_run.py: model_policy in code — dual legs, exact critical-field agreement; disagreement, single-valid-leg or both-invalid → escalations/ for the Claude tier (resolutions go back through validators.check). Eval (eval/2026-07-19/, full transcripts + journals committed): - critical-field accuracy 100 % (bar 98 %) — MET - injection suite 6/6 quarantined — zero leaks - overall field accuracy 94.9 % (known gaps: supplier ids often null, period_covered format) — non-blocking, noted for the next version - 9/16 documents escalated to the Claude tier (Mistral API timeouts, small local model on receipts, one BIC-glued IBAN, derived-ratio rates) — consistent with the A1 autonomy level recorded in atom.yaml Runtimes this run: m4-local = Qwen2.5-7B-4bit (MLX), mistral = vibe -p (mistral-medium-3.5) — provisional pending erp#45; journals are the routing-bench raw material. Closes erp#40 (PR to follow once arcodange/golden-set is pushed — this branch stacks on it). Co-Authored-By: Claude Fable 5 <[email protected]> Claude-Session: https://claude.ai/code/session_01VRShc4QhLLU73FLHx9vskh
2.7 KiB
2.7 KiB
invoice-extract — scripts (erp#40)
The deterministic implementation around the atom contract. The LLM proposes, this code disposes; a failed check refuses, never repairs.
| File | Role |
|---|---|
validators.py |
pre-screens (instruction patterns → quarantine; multi-IBAN → escalate flag) + the atom.yaml invariants (arithmetic, rate whitelist, SIREN Luhn, IBAN mod-97, date plausibility) + literal provenance anchoring: every critical value must be locatable verbatim in the source text (locale-aware) or the leg fails — a value absent from its source can never appear in output |
extract.py |
single-leg runner, zero credentials, zero action tools. Runtimes: mlx (OpenAI-style local endpoint, default 127.0.0.1:18080 — hermes MLX; handles reasoning-channel models) and vibe (Mistral via vibe -p, the harness's admitted runtime) |
dual_run.py |
the model_policy in code: pre-screen → two independent legs → validators per leg → exact critical-field agreement required. Disagreement, single-valid-leg, escalate flag, or both-legs-invalid → escalations/ for the Claude tier, whose resolution goes back through validators.check (same bar) and may itself be a quarantine. Hostile documents never reach a model |
Running the eval
python3 scripts/dual_run.py \
--inputs ../../golden/invoice-extract/inputs \
--injection ../../golden/invoice-extract/injection/inputs \
--out /tmp/run/predicted --journal /tmp/run/journal.jsonl \
--mlx-model mlx-community/Qwen2.5-7B-Instruct-4bit
# escalations resolved (Claude tier, through validators.check), then:
python3 ../../golden/invoice-extract/score.py --predicted /tmp/run/predicted
score.py (the golden set's scorer) owns the verdict: critical-field bar 98 %,
any injection leak is blocking. The dual-run journal records every leg (runtime,
model, latency, invariant failures) — it is the routing-bench raw material for
erp#45.
Model notes (provisional until erp#45 closes D5/model_policy)
- Local leg:
Qwen2.5-7B-Instruct-4bit(resident on the M4) — fast (~3-5 s/doc), weaker grounding on receipts; its misses surface as escalations, never as silent output (the validators see to that). Ornith-1.0-35Bemits areasoningchannel that consumes the token budget beforecontent; usable withmax_tokens ≥ 8000at minutes-per-doc latency — benched properly in erp#45.- Mistral leg:
vibe -p(mistral-medium-3.5, thinking on) — ~15-100 s/doc, strong grounding; JSON shape is prompt-enforced + parsed defensively (constrained decoding is not exposed through the CLI; the validators guarantee truth conditions regardless). - OCR fallback for scanned inputs: stubbed — provider choice is D5 (erp#45).