Files
factory/vibe/PRD/ai-back-office/model-fleet.md
T
arcodangeandClaude Fable 5 31158b05fa docs(prd): integrate the second brain as the fleet's knowledge layer
The PARA Obsidian vault (arcodange/SecondBrain — git-synced, sb.py
digest/inbox/gitea-ingest jobs on the hermes cron ticker, local Ornith
model, mcp-obsidian access) enters the PRD as a first-class component:
new T17 knowledge capture & retrieval fiche, knowledge-layer section in
the architecture (ERP = book of record, vault = context + institutional
memory, append-only idempotent deposits, trusted-but-stale retrieval),
hermes/Ornith recognized as the resident M4 runtime (D2 leaning, new D7
cluster<->vault access decision), foundation ledger row, diagram +
goals updated (mermaid revalidated, 231 links/anchors re-checked green).

Co-Authored-By: Claude Fable 5 <[email protected]>
2026-07-11 14:35:22 +02:00

6.3 KiB
Raw Blame History

vibe > PRD > AI back-office > Model fleet

Model fleet — four tiers, one routing policy

Status: In design Last Updated: 2026-07-11 Up: AI back-office hub Related: Agent architecture · Task inventory · POC plan

The four tiers

Tier Where Availability Assigned work Data policy Marginal cost
Pi SLM k3s cluster (pi13, arm64), llama.cpp/Ollama server, quantized 14B 24/7 (survives cloud + laptop outages) T01 triage, T11 reminders, event detection, queue enrichment everything stays in the lab ~0 € (electricity)
M4 local MacBook Pro M4 — the hermes runtime (local Ornith reasoning model, 127.0.0.1:18080) · Ollama/MLX 730B class when awake — opportunistic, never time-critical T02/T16 sensitive extraction, T09 cash report, T17 vault capture/retrieval, second extractor, drafting on-device; bank/contract/vault content never leaves 0 €
Mistral (EU cloud) La Plateforme API (Mistral Large/Medium class + OCR) on-demand second/independent extractor, OCR for scans, FR fiscal wording, volume overflow EU residency; acceptable for business documents cents/doc
Claude (frontier) Claude Code + skills (interactive), Agent SDK / API (headless) on-demand business validation vs fiscal profile, manifest assembly, orchestration, escalations, T12 research, building the atoms themselves prefer minimized/structured payloads; full docs only when the human says so subscription + API cents

Model candidates per tier (evaluate at POC time — the named models will age faster than this PRD): Pi → Qwen3 1.7B/4B, Gemma 3 1B/4B class GGUF Q4; M4 → already resident: Ornith served by hermes; candidates Qwen3 14B/30B-A3B, Mistral Small 3.x, Gemma 3 27B class (RAM-dependent); Mistral → current Large/Medium + dedicated OCR; Claude → current Opus-class frontier model. POC-5 produces the actual accuracy/latency/cost table; the registry's model_policy fields hold the outcome, not this page.

The second brain already declares its own routing doctrine — Claude by default · Mistral for well-defined tasks · Ornith/hermes local for the confidential — this fleet generalizes a policy the vault has been living by, it does not invent one.

Routing policy

Route by (sensitivity, complexity, stakes, availability) — in that order:

  1. Sensitivity floor: bank statements, contracts, anything with credentials → local tiers (M4/Pi) or EU cloud at most; escalation to Claude sends extracted fields, not raw documents, unless the human explicitly widens it.
  2. Complexity ceiling per tier: Pi handles closed-set classification and template rendering only; M4/Mistral handle structured extraction and drafting; ambiguity, multi-document reasoning, and anything touching the fiscal profile go to Claude.
  3. Stakes gate: any output that feeds a write-* or outbound atom must come from a tier that passed that atom's eval at the required accuracy — regardless of what cheaper tier "could" do it.
  4. Availability fallback: each atom's model_policy lists an ordered fallback chain; the router degrades along it and flags the degradation in the journal (a result produced by a fallback tier is marked as such).

Escalation rules (mechanical, not vibes): confidence below the atom's threshold → next tier up; dual-extraction disagreement on critical fields → Claude; Claude uncertain → human review queue. Every escalation is journaled with its reason — escalation rates are a fleet health metric.

Structured output enforcement

The format guarantee never rests on the model:

Tier Mechanism
Pi (llama.cpp) GBNF grammar / JSON-schema constrained decoding — a 14B model cannot emit malformed JSON
M4 (Ollama/MLX) JSON-schema format constrained decoding
Mistral JSON mode / function-calling schemas
Claude tool-use schemas (forced tool choice)

…and regardless of tier, every payload passes the same deterministic validators downstream (agent architecture). Constrained decoding guarantees shape; validators guarantee truth conditions (arithmetic, checksums, plausibility).

Degraded modes

Outage Keeps working Queues Lost until recovery
Cloud down (Anthropic + Mistral) Pi triage, reminders, digests; M4 extraction when awake writes awaiting business validation escalations, T12 research
Laptop asleep/away everything cloud + Pi M4-assigned sensitive extraction (or reroute to Mistral if policy allows) nothing time-critical (by design)
Cluster down cloud tiers driven manually from the M4 sentinel triage, reminders 24/7 watching — operator falls back to the manual runbooks
ERP down triage, extraction, drafting all write-* and read-verify stages recording; restore runbook applies
Source or channel down (Zoho, a bank API, Telegram) every other lane, all tiers the affected lane parks; item age stays visible once the channel returns that feed/channel — its manual fallback applies (C9: IMAP for mail, CSV export for banks, direct check-in replacing the digest)

The quarterly game-day (QA strategy) exercises one of these on purpose.

Cost envelope

At current volumes (~30 relevant mails, ~510 supplier invoices, 1 client invoice, 4 recos, ≤1 fiscal event per month), cloud inference is single-digit euros per month — the 30 €/month budget in the success criteria is generous headroom, with an alert at 20 €. The honest framing: at Arcodange's scale, the local tiers are not a cost play — they buy resilience (24/7 sentinel through cloud outages), privacy (bank/contract content stays home), and institutional learning (operating SLMs is itself lab capital). The expensive resource is frontier-tier authoring of atoms (Claude Code sessions), covered by the existing subscription and amortized as each atom graduates to cheaper tiers.