Files
erp/fleet
arcodangeandClaude Fable 5 e9d4a2bcb2 feat(fleet): invoice-extract atom — dual extraction + validators + provenance (erp#40)
Implementation of the T02 atom over the erp#39 golden set:

- validators.py: instruction-pattern + multi-IBAN pre-screens (0 hard false
  positives on the 16 real docs; all 6 injection fixtures quarantined BEFORE
  any model call), the atom.yaml invariants, and literal provenance anchoring
  with locale-aware locate (FR/EN months incl. abbreviations, NBSP-tolerant
  amounts, line-wrap + column-interleave fragment anchoring for refs).
- extract.py: single-leg runner (MLX endpoint / vibe -p), zero credentials,
  zero action tools; reasoning-channel aware.
- dual_run.py: model_policy in code — dual legs, exact critical-field
  agreement; disagreement, single-valid-leg or both-invalid → escalations/
  for the Claude tier (resolutions go back through validators.check).

Eval (eval/2026-07-19/, full transcripts + journals committed):
- critical-field accuracy 100 % (bar 98 %) — MET
- injection suite 6/6 quarantined — zero leaks
- overall field accuracy 94.9 % (known gaps: supplier ids often null,
  period_covered format) — non-blocking, noted for the next version
- 9/16 documents escalated to the Claude tier (Mistral API timeouts, small
  local model on receipts, one BIC-glued IBAN, derived-ratio rates) —
  consistent with the A1 autonomy level recorded in atom.yaml

Runtimes this run: m4-local = Qwen2.5-7B-4bit (MLX), mistral = vibe -p
(mistral-medium-3.5) — provisional pending erp#45; journals are the
routing-bench raw material.

Closes erp#40 (PR to follow once arcodange/golden-set is pushed — this branch
stacks on it).

Co-Authored-By: Claude Fable 5 <[email protected]>
Claude-Session: https://claude.ai/code/session_01VRShc4QhLLU73FLHx9vskh
2026-07-19 00:31:21 +02:00
..

fleet/ — the atom registry

The fleet is Arcodange's AI back-office: narrow agents ("atoms") that operate the Dolibarr ERP's daily admin & accounting under the AI back-office PRD. This directory is the registry — the versioned source of truth for what the fleet may do. An atom absent from the registry does not run. Contract semantics come from the PRD atom contract; file syntax from the PRD document surface.

What an atom is

One narrow capability (classify, extract, validate, record, reconcile, report, remind) with a strict I/O contract and deterministic validators around it. The LLM proposes, code disposes: formats, arithmetic, checksums, dedupe and referential integrity are enforced by validators, and a model output that fails validation is quarantined, never auto-corrected. Workflows are compositions of atoms with explicit gates — never one prompt that "does the accounting".

Each atom lives in fleet/atoms/<atom>/:

File Role
atom.yaml the registry entry — the contract (schema below)
prompt.md thin runtime prompt, ≤ ~40 lines, extends exactly one class skeleton; no business rules (rules live in fleet/profile/ and in validators)
scripts/ the deterministic implementation: runners, validators, scoring hooks

Folder name = atom name = registry name — the house <app> join-key discipline applied to atoms.

Layout

fleet/
├── README.md                  # this file: registry doc + atom.yaml schema
├── classes/                   # the 7 prompt skeletons (PRD agent catalog)
│   ├── sentinel.md
│   ├── extractor.md
│   ├── erp-scribe.md
│   ├── deterministic-controller.md   # no-LLM by design
│   ├── analyst-writer.md
│   ├── researcher.md
│   └── knowledge-archivist.md
├── atoms/
│   └── invoice-extract/       # T02 — the worked example (contract only; implementation = erp#40)
│       ├── atom.yaml
│       ├── prompt.md
│       └── scripts/
├── golden/                    # per-atom golden sets — land with erp#39
└── profile/                   # fiscal.yaml + calendar.yaml + ADC register — land with erp#54

atom.yaml — the contract, field by field

Per the PRD atom contract:

Field Meaning
name, version Identity. Folder name = name. version bumps on any behavioral change (prompt, model, validator) — a bump re-triggers the atom's golden-set evals.
input_schema / output_schema JSON Schema for the atom's I/O; enforced at runtime (constrained decoding where the model tier supports it).
invariants Deterministic post-conditions checked by code after every run (e.g. HT + TVA == TTC ± 0.01). A failed invariant quarantines the output — refuse, never repair.
side_effect_class read · draft · write-sandbox · write-prod · outbound — drives which gates and credentials apply, per the PRD environment posture table.
idempotency_key How a replay is recognized (e.g. supplier + ref_supplier + TTC) — a second run with the same key must be a no-op.
autonomy The earned level (A0A3 on the autonomy ladder) + a pointer to the eval evidence that justifies it.
model_policy Preferred tier, fallbacks, escalation rule, per the PRD model fleet; closed per-atom by routing-bench evidence (erp#45 for the first atoms).
eval_ref Where the golden set + scoring script live (fleet/golden/<atom>/).

Two registry conveniences beyond the PRD contract fields bind the entry to the rest of the surface: class (which fleet/classes/<class>.md skeleton the prompt extends) and task (the PRD task-inventory id the atom serves).

The worked example is atoms/invoice-extract/atom.yaml (T02) — contract only; its implementation is erp#40.

How an atom graduates

Autonomy is earned per atom, never assumed. The levels (A0 manual → A1 prepare → A2 rehearse + gate → A3 autonomous + audit) are defined on the PRD autonomy ladder; promotion and demotion are mechanical, per the PRD autonomy promotion gates (golden-set evals, unedited-approval streaks, incident demotion — the bars live there, not here). The earned level and its evidence are recorded in the atom's autonomy field: a promotion is a PR that changes that field with the evidence linked, verified per the QA strategy's independent-verification rule.

Conventions

  • English for all agent-facing files (house language policy).
  • Same-change freshness: a change to an atom that leaves its atom.yaml / prompt.md stale is an incomplete change.
  • One capability per file; YAML/frontmatter over prose for anything a machine parses.
  • Environment rules (trunk hygiene, read-only prod, sandbox-first writes, promote gate) are the repo-wide ones: AGENTS.md operating rules + dolibarr-sandbox-write SKILL.md.