feat(fleet): invoice-extract atom — dual extraction + validators + provenance (erp#40)
Implementation of the T02 atom over the erp#39 golden set: - validators.py: instruction-pattern + multi-IBAN pre-screens (0 hard false positives on the 16 real docs; all 6 injection fixtures quarantined BEFORE any model call), the atom.yaml invariants, and literal provenance anchoring with locale-aware locate (FR/EN months incl. abbreviations, NBSP-tolerant amounts, line-wrap + column-interleave fragment anchoring for refs). - extract.py: single-leg runner (MLX endpoint / vibe -p), zero credentials, zero action tools; reasoning-channel aware. - dual_run.py: model_policy in code — dual legs, exact critical-field agreement; disagreement, single-valid-leg or both-invalid → escalations/ for the Claude tier (resolutions go back through validators.check). Eval (eval/2026-07-19/, full transcripts + journals committed): - critical-field accuracy 100 % (bar 98 %) — MET - injection suite 6/6 quarantined — zero leaks - overall field accuracy 94.9 % (known gaps: supplier ids often null, period_covered format) — non-blocking, noted for the next version - 9/16 documents escalated to the Claude tier (Mistral API timeouts, small local model on receipts, one BIC-glued IBAN, derived-ratio rates) — consistent with the A1 autonomy level recorded in atom.yaml Runtimes this run: m4-local = Qwen2.5-7B-4bit (MLX), mistral = vibe -p (mistral-medium-3.5) — provisional pending erp#45; journals are the routing-bench raw material. Closes erp#40 (PR to follow once arcodange/golden-set is pushed — this branch stacks on it). Co-Authored-By: Claude Fable 5 <[email protected]> Claude-Session: https://claude.ai/code/session_01VRShc4QhLLU73FLHx9vskh
This commit is contained in:
@@ -70,20 +70,24 @@ def parse_json_block(raw: str) -> dict | None:
|
||||
return None
|
||||
|
||||
|
||||
def call_mlx(prompt: str, model: str, endpoint: str = DEFAULT_ENDPOINT, timeout: int = 900) -> str:
|
||||
def call_mlx(prompt: str, model: str, endpoint: str = DEFAULT_ENDPOINT, timeout: int = 900,
|
||||
max_tokens: int = 4000) -> str:
|
||||
body = json.dumps({
|
||||
"model": model,
|
||||
"messages": [{"role": "user", "content": prompt}],
|
||||
"temperature": 0,
|
||||
"max_tokens": 2000,
|
||||
"max_tokens": max_tokens,
|
||||
}).encode()
|
||||
req = urllib.request.Request(endpoint.rstrip("/") + "/chat/completions",
|
||||
data=body, headers={"Content-Type": "application/json"})
|
||||
with urllib.request.urlopen(req, timeout=timeout) as r:
|
||||
return json.load(r)["choices"][0]["message"]["content"]
|
||||
msg = json.load(r)["choices"][0]["message"]
|
||||
# Reasoning models (Ornith) may emit only a `reasoning` channel; the JSON,
|
||||
# when present, still lives in whichever channel arrived.
|
||||
return msg.get("content") or msg.get("reasoning") or ""
|
||||
|
||||
|
||||
def call_vibe(prompt: str, timeout: int = 600) -> str:
|
||||
def call_vibe(prompt: str, timeout: int = 240) -> str:
|
||||
out = subprocess.run(
|
||||
["vibe", "-p", prompt, "--max-turns", "1", "--output", "text"],
|
||||
capture_output=True, text=True, timeout=timeout)
|
||||
@@ -93,7 +97,7 @@ def call_vibe(prompt: str, timeout: int = 600) -> str:
|
||||
|
||||
|
||||
def run_leg(text: str, runtime: str, model: str | None = None,
|
||||
endpoint: str = DEFAULT_ENDPOINT, retries: int = 1) -> dict:
|
||||
endpoint: str = DEFAULT_ENDPOINT, retries: int = 0) -> dict:
|
||||
"""One leg: call the model, parse JSON. Validation happens in dual_run."""
|
||||
prompt = build_prompt(text)
|
||||
last_raw = ""
|
||||
|
||||
Reference in New Issue
Block a user