Files
erp/fleet
arcodangeandClaude Fable 5 bdd3d63b61 feat(fleet): invoice-extract atom — validators, screens, dual-run orchestrator (WIP erp#40)
- validators.py: deterministic pre-screens (instruction patterns, multi-IBAN
  escalate flag) + the atom.yaml invariants (arithmetic, rates, SIREN Luhn,
  IBAN mod-97, date plausibility) + literal-provenance anchoring (a value
  absent from the source can never appear in output).
  Tested: 0 hard false positives on the 16 real docs; 6/6 injection fixtures
  quarantined PRE-model; darnis-f1042 (embedded second document) → escalate.
- extract.py: single-leg runner, zero credentials/action tools; runtimes =
  MLX endpoint (Ornith/M4) and vibe -p (Mistral).
- dual_run.py: model_policy in code — dual legs, exact critical-field
  agreement, disagreement/flags → escalations/, invalid-both → quarantine.

Eval run against the golden set follows in this branch.

Co-Authored-By: Claude Fable 5 <[email protected]>
Claude-Session: https://claude.ai/code/session_01VRShc4QhLLU73FLHx9vskh
2026-07-18 23:41:44 +02:00
..

fleet/ — the atom registry

The fleet is Arcodange's AI back-office: narrow agents ("atoms") that operate the Dolibarr ERP's daily admin & accounting under the AI back-office PRD. This directory is the registry — the versioned source of truth for what the fleet may do. An atom absent from the registry does not run. Contract semantics come from the PRD atom contract; file syntax from the PRD document surface.

What an atom is

One narrow capability (classify, extract, validate, record, reconcile, report, remind) with a strict I/O contract and deterministic validators around it. The LLM proposes, code disposes: formats, arithmetic, checksums, dedupe and referential integrity are enforced by validators, and a model output that fails validation is quarantined, never auto-corrected. Workflows are compositions of atoms with explicit gates — never one prompt that "does the accounting".

Each atom lives in fleet/atoms/<atom>/:

File Role
atom.yaml the registry entry — the contract (schema below)
prompt.md thin runtime prompt, ≤ ~40 lines, extends exactly one class skeleton; no business rules (rules live in fleet/profile/ and in validators)
scripts/ the deterministic implementation: runners, validators, scoring hooks

Folder name = atom name = registry name — the house <app> join-key discipline applied to atoms.

Layout

fleet/
├── README.md                  # this file: registry doc + atom.yaml schema
├── classes/                   # the 7 prompt skeletons (PRD agent catalog)
│   ├── sentinel.md
│   ├── extractor.md
│   ├── erp-scribe.md
│   ├── deterministic-controller.md   # no-LLM by design
│   ├── analyst-writer.md
│   ├── researcher.md
│   └── knowledge-archivist.md
├── atoms/
│   └── invoice-extract/       # T02 — the worked example (contract only; implementation = erp#40)
│       ├── atom.yaml
│       ├── prompt.md
│       └── scripts/
├── golden/                    # per-atom golden sets — land with erp#39
└── profile/                   # fiscal.yaml + calendar.yaml + ADC register — land with erp#54

atom.yaml — the contract, field by field

Per the PRD atom contract:

Field Meaning
name, version Identity. Folder name = name. version bumps on any behavioral change (prompt, model, validator) — a bump re-triggers the atom's golden-set evals.
input_schema / output_schema JSON Schema for the atom's I/O; enforced at runtime (constrained decoding where the model tier supports it).
invariants Deterministic post-conditions checked by code after every run (e.g. HT + TVA == TTC ± 0.01). A failed invariant quarantines the output — refuse, never repair.
side_effect_class read · draft · write-sandbox · write-prod · outbound — drives which gates and credentials apply, per the PRD environment posture table.
idempotency_key How a replay is recognized (e.g. supplier + ref_supplier + TTC) — a second run with the same key must be a no-op.
autonomy The earned level (A0A3 on the autonomy ladder) + a pointer to the eval evidence that justifies it.
model_policy Preferred tier, fallbacks, escalation rule, per the PRD model fleet; closed per-atom by routing-bench evidence (erp#45 for the first atoms).
eval_ref Where the golden set + scoring script live (fleet/golden/<atom>/).

Two registry conveniences beyond the PRD contract fields bind the entry to the rest of the surface: class (which fleet/classes/<class>.md skeleton the prompt extends) and task (the PRD task-inventory id the atom serves).

The worked example is atoms/invoice-extract/atom.yaml (T02) — contract only; its implementation is erp#40.

How an atom graduates

Autonomy is earned per atom, never assumed. The levels (A0 manual → A1 prepare → A2 rehearse + gate → A3 autonomous + audit) are defined on the PRD autonomy ladder; promotion and demotion are mechanical, per the PRD autonomy promotion gates (golden-set evals, unedited-approval streaks, incident demotion — the bars live there, not here). The earned level and its evidence are recorded in the atom's autonomy field: a promotion is a PR that changes that field with the evidence linked, verified per the QA strategy's independent-verification rule.

Conventions

  • English for all agent-facing files (house language policy).
  • Same-change freshness: a change to an atom that leaves its atom.yaml / prompt.md stale is an incomplete change.
  • One capability per file; YAML/frontmatter over prose for anything a machine parses.
  • Environment rules (trunk hygiene, read-only prod, sandbox-first writes, promote gate) are the repo-wide ones: AGENTS.md operating rules + dolibarr-sandbox-write SKILL.md.