Files
erp/ops/sandbox
arcodangeandClaude Opus 5 6a5b7baa23 chore: réconcilier le trunk avec main, et corriger la note du mode vierge
DEUX CHOSES.

1. RÉCONCILIATION. Le trunk accusait 49 commits de retard, 0 d'avance : tout son
delta vivait dans l'arbre de travail, mélange d'améliorations réelles et d'état
périmé. Il est remis sur origin/main, et son état est préservé sur la branche
trunk-snapshot-20260814. Ce qui méritait d'être porté l'est ici :

  - test/main.ts : chemin Work/ au lieu de Desktop/ (l'ancien n'existe plus) et
    module TIERS activé. L'URL par défaut reste corrigée par la garde d'hôte ;
  - ansible README : la commande du playbook recurrentBackup ;
  - company.json : siteWeb arcodange.fr et email @arcodange.fr — le trunk était
    à jour sur ces deux champs quand main portait encore duckdns et gmail.

NE SONT PAS PORTÉS, délibérément :
  - le .gitignore du trunk, qui avait PERDU les règles replay-packs ;
  - company.json SAS / 6201Z, périmé — main porte SARL / 6201Z depuis le PR #77 ;
  - un console.log de débogage dans forms.ts ;
  - createEphemeralAgent + permissions.ts + orchestratorExample.ts : travail
    cohérent mais qui DOUBLE le modèle de scopes durables (scopes.ts +
    provisionAiUser.ts) effectivement livré. Deux conceptions concurrentes de la
    même chose ne se portent pas en silence — arbitrage opérateur.

2. LA NOTE DU MODE VIERGE ÉTAIT FAUSSE. Elle annonçait qu'il manquait un « golden
empty ». Rien ne manquait : la chaîne existe en entier dans test/, et c'est
main.ts qui pilote l'installeur. J'avais proposé de construire ce qui était déjà
là, faute d'avoir regardé — exactement le contrôle inscrit dans AGENTS.md le
matin même. La note décrit désormais la chaîne réelle :

  blank --yes  ->  main.ts  ->  install.lock  ->  provisionSandbox.ts

Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
2026-08-14 22:04:43 +02:00
..

erp-sandbox lifecycle ops

Tooling to make erp-sandbox iso-prod and to reset it, implementing ADR-0003 (sandbox state lifecycle). The sandbox exists so AI agents can rehearse Dolibarr write operations against a faithful copy of prod, with a structural guarantee that the rehearsal path can never mutate prod.

The prod-integrity guarantee (why this is safe)

Layer Enforcement
prod is read-only during a refresh pg_dump runs with default_transaction_read_only=on
the restore can only write the sandbox it uses the sandbox's own dynamic creds — a member of erp_sandbox_role, which owns only erp-sandbox
no database is dropped/created wipe is DROP OWNED BY erp_sandbox_role CASCADE; reload is pg_restore (no CREATEDB, no superuser)
prod is structurally undroppable DROP DATABASE needs ownership; erp_sandbox_role does not own prod erp (owned by erp_role)

The only prod-capable credential on the platform is the superuser=true provider in factory postgres/iac/providers.tf, used only in the human-gated postgres.yaml CI. This tooling never touches it.

Usage

./sandbox-lifecycle.sh refresh-from-prod   # clone prod DB (data + config) into erp-sandbox
./sandbox-lifecycle.sh sync-documents      # copy mycompany/ uploads (company logo, PDFs)
./sandbox-lifecycle.sh refresh             # both, in order

refresh-from-prod pauses ArgoCD self-heal on the erp-sandbox Application (else self-heal reverts the scale-down within seconds and the seed would run with the app still connected), scales the sandbox pod to 0, dumps the full prod public schema (read-only), wipes the sandbox's app objects, restores, scales back up and re-arms self-heal. App restore and self-heal restoration are guarded by an EXIT trap, so an interrupt can't leave the sandbox scaled to 0 with self-heal off. It dumps the whole schema (not just llx_*) so app helper functions and their triggers (e.g. update_modified_column_tms()) come over; it filters out the provisioner-owned user_lookup pgbouncer function from the restore TOC because that object already exists per-environment and is not app data.

Note: pg_restore runs with an explicit -U (the sandbox role) — without it the postgres image connects as its OS user root and auth-fails. Its exit code is not trusted (it returns non-zero on the harmless "schema public already exists" notice); success is verified by counting restored llx_* tables.

Heads-up: a full refresh is iso-prod, so it overwrites llx_user with prod's and wipes the ai_agent_sandbox write user + its API key (and resets DOLI_INSTANCE_UNIQUE_ID to prod's, invalidating any prior key). After a refresh, re-run test/provisionSandbox.ts to recreate the agent (it re-grants its rights, incl. banque lire) and refresh the dolibarr-sandbox-write skill .env from the new key file.

Two fidelity caveats (by design — see ADR-0003)

  1. Encryption. Dolibarr ties some encrypted fields (notably API keys) to DOLI_INSTANCE_UNIQUE_ID. The sandbox has its own uuid, so prod-encrypted values won't decrypt there. This is why the write-scoped ai_agent_sandbox API key must be generated inside the sandbox (see ../../test/ POC), not copied from prod. Most data is plaintext and unaffected.
  2. Uploaded files live on the PVC, not the DB. A DB refresh copies the logo const (MAIN_INFO_SOCIETE_LOGO) but not the image; sync-documents copies the documents/mycompany tree so the logo + attachments actually render.

BDD reset loop (E4)

For repeated rehearsals, refresh-from-prod is the "reset to prod state". A faster checkpoint/reset that avoids re-reading prod each time (cache a golden dump on a small PVC, then DROP OWNED + pg_restore from it) is the documented next optimization — see ADR-0003 §Decision/Consequences.

Hardening backlog

  • Replace the transient copy of prod's read+write creds with a dedicated read-only Postgres role (issued via a Vault dynamic role) so the dump path is least-privilege by construction, not just by default_transaction_read_only.
  • Provision a golden-cache PVC for fast BDD resets.