DEUX CHOSES, dont une inachevée et dite comme telle. 1. LA SAUVEGARDE N'AVAIT AUCUNE GARDE DE CLUSTER. 24 appels kubectl nus, aucun contexte épinglé — alors que le script lit des secrets, crée des Jobs et sait RESTAURER une base. Lancé sur le contexte courant d'une station de travail, il serait parti chercher les secrets Longhorn du cluster d'un client. Même patron que ops/sandbox : ERP_KUBE_CONTEXT, wrapper K(), empreinte positive vérifiée avant tout dispatch. Testé contre le cluster client et contre un contexte inexistant : il refuse les deux. Sauvegarde de production passée dans la foulée. La base a été dédupliquée à juste titre — rien n'avait été écrit depuis la sauvegarde automatique de 01:00. Au passage, le CronJob quotidien tourne bien depuis juillet ; il était encore noté comme à faire. 2. LE MODE VIERGE PURGE MAIS NE RECONSTRUIT PAS. `blank --yes` vide la base du bac à sable (295 tables -> 0, vérifié) avec trois protections empilées : refus sans --yes, garde de cluster, et une relecture de current_database() DANS le Job lui-même — une purge sur la mauvaise base ne se rattrape pas par un refresh, contrairement à la bonne. Mais l'instance ne se reconstruit pas. Premier essai : install.lock vit sur le volume documents, que la purge ne touche pas, si bien que Dolibarr servait un login sur un schéma inexistant. Correctif appliqué — retrait du verrou puis redémarrage. Second essai : le verrou reste absent, et le schéma reste à ZÉRO table. L'entrypoint de l'image ne lance aucune installation automatique. CE QU'IL MANQUE est donc nommé dans le script : un « golden empty », pg_dump d'une instance fraîchement installée — schéma et données de référence, aucune donnée métier. `blank` le restaurerait au lieu de laisser la base vide, comme refresh-from-prod restaure le dump de production. Seule la source change. En l'état `blank` laisse le bac à sable inutilisable, et le script le dit. Le bac à sable a été remis iso-prod avant de rendre la main. Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
erp-sandbox lifecycle ops
Tooling to make erp-sandbox iso-prod and to reset it, implementing
ADR-0003
(sandbox state lifecycle). The sandbox exists so AI agents can rehearse Dolibarr
write operations against a faithful copy of prod, with a structural guarantee
that the rehearsal path can never mutate prod.
The prod-integrity guarantee (why this is safe)
| Layer | Enforcement |
|---|---|
| prod is read-only during a refresh | pg_dump runs with default_transaction_read_only=on |
| the restore can only write the sandbox | it uses the sandbox's own dynamic creds — a member of erp_sandbox_role, which owns only erp-sandbox |
| no database is dropped/created | wipe is DROP OWNED BY erp_sandbox_role CASCADE; reload is pg_restore (no CREATEDB, no superuser) |
| prod is structurally undroppable | DROP DATABASE needs ownership; erp_sandbox_role does not own prod erp (owned by erp_role) |
The only prod-capable credential on the platform is the superuser=true provider
in factory postgres/iac/providers.tf, used only in the human-gated
postgres.yaml CI. This tooling never touches it.
Usage
./sandbox-lifecycle.sh refresh-from-prod # clone prod DB (data + config) into erp-sandbox
./sandbox-lifecycle.sh sync-documents # copy mycompany/ uploads (company logo, PDFs)
./sandbox-lifecycle.sh refresh # both, in order
refresh-from-prod pauses ArgoCD self-heal on the erp-sandbox Application (else
self-heal reverts the scale-down within seconds and the seed would run with the app
still connected), scales the sandbox pod to 0, dumps the full prod public schema
(read-only), wipes the sandbox's app objects, restores, scales back up and re-arms
self-heal. App restore and self-heal restoration are guarded by an EXIT trap, so an
interrupt can't leave the sandbox scaled to 0 with self-heal off. It dumps the
whole schema (not just llx_*) so app helper functions and their triggers (e.g.
update_modified_column_tms()) come over; it filters out the provisioner-owned
user_lookup pgbouncer function from the restore TOC because that object already
exists per-environment and is not app data.
Note:
pg_restoreruns with an explicit-U(the sandbox role) — without it the postgres image connects as its OS userrootand auth-fails. Its exit code is not trusted (it returns non-zero on the harmless "schema public already exists" notice); success is verified by counting restoredllx_*tables.
Heads-up: a full refresh is iso-prod, so it overwrites
llx_userwith prod's and wipes theai_agent_sandboxwrite user + its API key (and resetsDOLI_INSTANCE_UNIQUE_IDto prod's, invalidating any prior key). After a refresh, re-runtest/provisionSandbox.tsto recreate the agent (it re-grants its rights, incl.banque lire) and refresh thedolibarr-sandbox-writeskill.envfrom the new key file.
Two fidelity caveats (by design — see ADR-0003)
- Encryption. Dolibarr ties some encrypted fields (notably API keys) to
DOLI_INSTANCE_UNIQUE_ID. The sandbox has its own uuid, so prod-encrypted values won't decrypt there. This is why the write-scopedai_agent_sandboxAPI key must be generated inside the sandbox (see../../test/POC), not copied from prod. Most data is plaintext and unaffected. - Uploaded files live on the PVC, not the DB. A DB refresh copies the logo
const (
MAIN_INFO_SOCIETE_LOGO) but not the image;sync-documentscopies thedocuments/mycompanytree so the logo + attachments actually render.
BDD reset loop (E4)
For repeated rehearsals, refresh-from-prod is the "reset to prod state". A
faster checkpoint/reset that avoids re-reading prod each time (cache a golden
dump on a small PVC, then DROP OWNED + pg_restore from it) is the documented
next optimization — see ADR-0003 §Decision/Consequences.
Hardening backlog
- Replace the transient copy of prod's read+write creds with a dedicated
read-only Postgres role (issued via a Vault dynamic role) so the dump path is
least-privilege by construction, not just by
default_transaction_read_only. - Provision a golden-cache PVC for fast BDD resets.