sandbox-lifecycle.sh scales deployments to zero, patches the ArgoCD Application and runs DROP OWNED ... CASCADE. Every one of those ran against whatever kube-context happened to be current. This workstation also carries a CLIENT production cluster. On 2026-07-25 a `checkpoint refresh` was issued while the current context was do-nyc3-kissmetrics-prod-k8s-cluster: the script patched the ArgoCD Application, scaled `erp-sandbox` to zero and copied a prod secret — all against the client's cluster. Nothing was damaged only because that cluster has no `application` CRD and no erp/erp-sandbox namespaces, so each call failed silently under `|| true`. That is luck, not a control. - ERP_KUBE_CONTEXT (default: "default") pins the target; every kubectl call now goes through K(), so nothing inherits the ambient context. - assert_arcodange_cluster() proves the target by positive fingerprint — the erp, erp-sandbox and argocd namespaces AND the erp-sandbox ArgoCD Application. A client cluster cannot match all four by accident. Wired into all three entry points, before any mutation. Verified: refuses the client context, refuses an unknown context, passes on the homelab and completes normally. Co-Authored-By: Claude Opus 5 (1M context) <[email protected]> Claude-Session: https://claude.ai/code/session_01VRShc4QhLLU73FLHx9vskh
erp-sandbox lifecycle ops
Tooling to make erp-sandbox iso-prod and to reset it, implementing
ADR-0003
(sandbox state lifecycle). The sandbox exists so AI agents can rehearse Dolibarr
write operations against a faithful copy of prod, with a structural guarantee
that the rehearsal path can never mutate prod.
The prod-integrity guarantee (why this is safe)
| Layer | Enforcement |
|---|---|
| prod is read-only during a refresh | pg_dump runs with default_transaction_read_only=on |
| the restore can only write the sandbox | it uses the sandbox's own dynamic creds — a member of erp_sandbox_role, which owns only erp-sandbox |
| no database is dropped/created | wipe is DROP OWNED BY erp_sandbox_role CASCADE; reload is pg_restore (no CREATEDB, no superuser) |
| prod is structurally undroppable | DROP DATABASE needs ownership; erp_sandbox_role does not own prod erp (owned by erp_role) |
The only prod-capable credential on the platform is the superuser=true provider
in factory postgres/iac/providers.tf, used only in the human-gated
postgres.yaml CI. This tooling never touches it.
Usage
./sandbox-lifecycle.sh refresh-from-prod # clone prod DB (data + config) into erp-sandbox
./sandbox-lifecycle.sh sync-documents # copy mycompany/ uploads (company logo, PDFs)
./sandbox-lifecycle.sh refresh # both, in order
refresh-from-prod pauses ArgoCD self-heal on the erp-sandbox Application (else
self-heal reverts the scale-down within seconds and the seed would run with the app
still connected), scales the sandbox pod to 0, dumps the full prod public schema
(read-only), wipes the sandbox's app objects, restores, scales back up and re-arms
self-heal. App restore and self-heal restoration are guarded by an EXIT trap, so an
interrupt can't leave the sandbox scaled to 0 with self-heal off. It dumps the
whole schema (not just llx_*) so app helper functions and their triggers (e.g.
update_modified_column_tms()) come over; it filters out the provisioner-owned
user_lookup pgbouncer function from the restore TOC because that object already
exists per-environment and is not app data.
Note:
pg_restoreruns with an explicit-U(the sandbox role) — without it the postgres image connects as its OS userrootand auth-fails. Its exit code is not trusted (it returns non-zero on the harmless "schema public already exists" notice); success is verified by counting restoredllx_*tables.
Heads-up: a full refresh is iso-prod, so it overwrites
llx_userwith prod's and wipes theai_agent_sandboxwrite user + its API key (and resetsDOLI_INSTANCE_UNIQUE_IDto prod's, invalidating any prior key). After a refresh, re-runtest/provisionSandbox.tsto recreate the agent (it re-grants its rights, incl.banque lire) and refresh thedolibarr-sandbox-writeskill.envfrom the new key file.
Two fidelity caveats (by design — see ADR-0003)
- Encryption. Dolibarr ties some encrypted fields (notably API keys) to
DOLI_INSTANCE_UNIQUE_ID. The sandbox has its own uuid, so prod-encrypted values won't decrypt there. This is why the write-scopedai_agent_sandboxAPI key must be generated inside the sandbox (see../../test/POC), not copied from prod. Most data is plaintext and unaffected. - Uploaded files live on the PVC, not the DB. A DB refresh copies the logo
const (
MAIN_INFO_SOCIETE_LOGO) but not the image;sync-documentscopies thedocuments/mycompanytree so the logo + attachments actually render.
BDD reset loop (E4)
For repeated rehearsals, refresh-from-prod is the "reset to prod state". A
faster checkpoint/reset that avoids re-reading prod each time (cache a golden
dump on a small PVC, then DROP OWNED + pg_restore from it) is the documented
next optimization — see ADR-0003 §Decision/Consequences.
Hardening backlog
- Replace the transient copy of prod's read+write creds with a dedicated
read-only Postgres role (issued via a Vault dynamic role) so the dump path is
least-privilege by construction, not just by
default_transaction_read_only. - Provision a golden-cache PVC for fast BDD resets.