feat(fleet): invoice-extract atom — dual extraction + validators + provenance (erp#40)
Implementation of the T02 atom over the erp#39 golden set: - validators.py: instruction-pattern + multi-IBAN pre-screens (0 hard false positives on the 16 real docs; all 6 injection fixtures quarantined BEFORE any model call), the atom.yaml invariants, and literal provenance anchoring with locale-aware locate (FR/EN months incl. abbreviations, NBSP-tolerant amounts, line-wrap + column-interleave fragment anchoring for refs). - extract.py: single-leg runner (MLX endpoint / vibe -p), zero credentials, zero action tools; reasoning-channel aware. - dual_run.py: model_policy in code — dual legs, exact critical-field agreement; disagreement, single-valid-leg or both-invalid → escalations/ for the Claude tier (resolutions go back through validators.check). Eval (eval/2026-07-19/, full transcripts + journals committed): - critical-field accuracy 100 % (bar 98 %) — MET - injection suite 6/6 quarantined — zero leaks - overall field accuracy 94.9 % (known gaps: supplier ids often null, period_covered format) — non-blocking, noted for the next version - 9/16 documents escalated to the Claude tier (Mistral API timeouts, small local model on receipts, one BIC-glued IBAN, derived-ratio rates) — consistent with the A1 autonomy level recorded in atom.yaml Runtimes this run: m4-local = Qwen2.5-7B-4bit (MLX), mistral = vibe -p (mistral-medium-3.5) — provisional pending erp#45; journals are the routing-bench raw material. Closes erp#40 (PR to follow once arcodange/golden-set is pushed — this branch stacks on it). Co-Authored-By: Claude Fable 5 <[email protected]> Claude-Session: https://claude.ai/code/session_01VRShc4QhLLU73FLHx9vskh
This commit is contained in:
@@ -1,8 +1,42 @@
|
||||
# invoice-extract/scripts — intentionally empty
|
||||
# invoice-extract — scripts (erp#40)
|
||||
|
||||
The implementation — dual-run extraction drivers (M4 local ∥ Mistral API),
|
||||
deterministic validators (arithmetic, rates, SIREN/IBAN checksums, dedupe,
|
||||
provenance re-verification), the stubbed OCR fallback and the scoring hooks —
|
||||
lands with [erp#40](https://gitea.arcodange.lab/arcodange-org/erp/issues/40).
|
||||
This scaffold ships the contract only ([`../atom.yaml`](../atom.yaml)); do not
|
||||
fake extraction code here.
|
||||
The deterministic implementation around the [atom contract](../atom.yaml). The
|
||||
LLM proposes, this code disposes; a failed check refuses, never repairs.
|
||||
|
||||
| File | Role |
|
||||
| --- | --- |
|
||||
| `validators.py` | pre-screens (instruction patterns → quarantine; multi-IBAN → escalate flag) + the `atom.yaml` invariants (arithmetic, rate whitelist, SIREN Luhn, IBAN mod-97, date plausibility) + **literal provenance anchoring**: every critical value must be locatable verbatim in the source text (locale-aware) or the leg fails — a value absent from its source can never appear in output |
|
||||
| `extract.py` | single-leg runner, zero credentials, zero action tools. Runtimes: `mlx` (OpenAI-style local endpoint, default `127.0.0.1:18080` — hermes MLX; handles reasoning-channel models) and `vibe` (Mistral via `vibe -p`, the harness's admitted runtime) |
|
||||
| `dual_run.py` | the `model_policy` in code: pre-screen → two independent legs → validators per leg → **exact critical-field agreement** required. Disagreement, single-valid-leg, escalate flag, or both-legs-invalid → `escalations/` for the **Claude tier**, whose resolution goes back through `validators.check` (same bar) and may itself be a quarantine. Hostile documents never reach a model |
|
||||
|
||||
## Running the eval
|
||||
|
||||
```bash
|
||||
python3 scripts/dual_run.py \
|
||||
--inputs ../../golden/invoice-extract/inputs \
|
||||
--injection ../../golden/invoice-extract/injection/inputs \
|
||||
--out /tmp/run/predicted --journal /tmp/run/journal.jsonl \
|
||||
--mlx-model mlx-community/Qwen2.5-7B-Instruct-4bit
|
||||
|
||||
# escalations resolved (Claude tier, through validators.check), then:
|
||||
python3 ../../golden/invoice-extract/score.py --predicted /tmp/run/predicted
|
||||
```
|
||||
|
||||
`score.py` (the golden set's scorer) owns the verdict: critical-field bar 98 %,
|
||||
any injection leak is blocking. The dual-run journal records every leg (runtime,
|
||||
model, latency, invariant failures) — it is the routing-bench raw material for
|
||||
erp#45.
|
||||
|
||||
## Model notes (provisional until erp#45 closes D5/model_policy)
|
||||
|
||||
- Local leg: `Qwen2.5-7B-Instruct-4bit` (resident on the M4) — fast (~3-5 s/doc),
|
||||
weaker grounding on receipts; its misses surface as escalations, never as
|
||||
silent output (the validators see to that).
|
||||
- `Ornith-1.0-35B` emits a `reasoning` channel that consumes the token budget
|
||||
before `content`; usable with `max_tokens ≥ 8000` at minutes-per-doc latency —
|
||||
benched properly in erp#45.
|
||||
- Mistral leg: `vibe -p` (`mistral-medium-3.5`, thinking on) — ~15-100 s/doc,
|
||||
strong grounding; JSON shape is prompt-enforced + parsed defensively
|
||||
(constrained decoding is not exposed through the CLI; the validators guarantee
|
||||
truth conditions regardless).
|
||||
- OCR fallback for scanned inputs: stubbed — provider choice is D5 (erp#45).
|
||||
|
||||
Reference in New Issue
Block a user