DEUX CHOSES, dont une inachevée et dite comme telle.
1. LA SAUVEGARDE N'AVAIT AUCUNE GARDE DE CLUSTER. 24 appels kubectl nus, aucun
contexte épinglé — alors que le script lit des secrets, crée des Jobs et sait
RESTAURER une base. Lancé sur le contexte courant d'une station de travail, il
serait parti chercher les secrets Longhorn du cluster d'un client. Même patron
que ops/sandbox : ERP_KUBE_CONTEXT, wrapper K(), empreinte positive vérifiée
avant tout dispatch. Testé contre le cluster client et contre un contexte
inexistant : il refuse les deux.
Sauvegarde de production passée dans la foulée. La base a été dédupliquée à
juste titre — rien n'avait été écrit depuis la sauvegarde automatique de 01:00.
Au passage, le CronJob quotidien tourne bien depuis juillet ; il était encore
noté comme à faire.
2. LE MODE VIERGE PURGE MAIS NE RECONSTRUIT PAS. `blank --yes` vide la base du
bac à sable (295 tables -> 0, vérifié) avec trois protections empilées : refus
sans --yes, garde de cluster, et une relecture de current_database() DANS le Job
lui-même — une purge sur la mauvaise base ne se rattrape pas par un refresh,
contrairement à la bonne.
Mais l'instance ne se reconstruit pas. Premier essai : install.lock vit sur le
volume documents, que la purge ne touche pas, si bien que Dolibarr servait un
login sur un schéma inexistant. Correctif appliqué — retrait du verrou puis
redémarrage. Second essai : le verrou reste absent, et le schéma reste à ZÉRO
table. L'entrypoint de l'image ne lance aucune installation automatique.
CE QU'IL MANQUE est donc nommé dans le script : un « golden empty », pg_dump
d'une instance fraîchement installée — schéma et données de référence, aucune
donnée métier. `blank` le restaurerait au lieu de laisser la base vide, comme
refresh-from-prod restaure le dump de production. Seule la source change.
En l'état `blank` laisse le bac à sable inutilisable, et le script le dit. Le
bac à sable a été remis iso-prod avant de rendre la main.
Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
sandbox-lifecycle.sh scales deployments to zero, patches the ArgoCD Application
and runs DROP OWNED ... CASCADE. Every one of those ran against whatever
kube-context happened to be current.
This workstation also carries a CLIENT production cluster. On 2026-07-25 a
`checkpoint refresh` was issued while the current context was
do-nyc3-kissmetrics-prod-k8s-cluster: the script patched the ArgoCD Application,
scaled `erp-sandbox` to zero and copied a prod secret — all against the client's
cluster. Nothing was damaged only because that cluster has no `application` CRD
and no erp/erp-sandbox namespaces, so each call failed silently under `|| true`.
That is luck, not a control.
- ERP_KUBE_CONTEXT (default: "default") pins the target; every kubectl call now
goes through K(), so nothing inherits the ambient context.
- assert_arcodange_cluster() proves the target by positive fingerprint — the
erp, erp-sandbox and argocd namespaces AND the erp-sandbox ArgoCD Application.
A client cluster cannot match all four by accident. Wired into all three
entry points, before any mutation.
Verified: refuses the client context, refuses an unknown context, passes on the
homelab and completes normally.
Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
Claude-Session: https://claude.ai/code/session_01VRShc4QhLLU73FLHx9vskh
After an iso-prod refresh the instance unique-id changes, so an api_key encrypted
with the OLD id can't be decrypted — Dolibarr renders non-UTF-8 bytes in the field.
The POC's generateApiKey reused any non-empty value, so it copied that garbage into
test/.ai_agent_sandbox.key (corrupt key, 401s). Now it reuses ONLY a clean key
(^[A-Za-z0-9_-]{24,}$); otherwise it clears the field and regenerates. So
`checkpoint provision` after a refresh yields a fresh, working key.
Also documents the open PLATFORM follow-ups in ops/backup/README.md (easy to find
when revisiting ERP backups): the orphaned Longhorn `default` recurring-job group
(other cluster volumes have no offsite backup), and verifying the factory
pg_dumpall host cron.
Co-Authored-By: Claude Opus 4.7 (1M context) <[email protected]>
dolibarr-backup.sh restore --db|--docs <ts> --env <e> --yes — the recovery half of
the dedicated backup. DESTRUCTIVE (gated by --yes + explicit --env): scales the app
to 0, then
- --db: DROP OWNED BY <owner_role> CASCADE + pg_restore --no-owner --role (same
mechanics as sandbox-lifecycle.sh), from s3://.../erp/<env>/db/<ts>.dump;
- --docs: clears /var/www/documents and untars s3://.../erp/<env>/docs/<ts>.tar.gz;
then scales the app back to 1. OWNER_ROLE per env (erp_role / erp_sandbox_role).
The key is the bare <ts> filename from `list`; --db/--docs selects the subpath.
Proven on the sandbox: backup → mutate MAIN_INFO_SOCIETE_NOM to a sentinel →
restore --db → the value reverted to the backup's ('Arcodange'). (First run caught
a path bug — the fetch missed the db/ subdir — now fixed.)
Co-Authored-By: Claude Opus 4.7 (1M context) <[email protected]>
Builds on the dedicated backup (erp#31).
Skip-if-unchanged: each half (DB / documents) carries a content fingerprint at
erp/<env>/.fp-{db,docs} and is dumped+uploaded only if it differs from the last
run — a quiet ERP day re-uploads nothing. Fingerprint = durable BUSINESS content
only: DB = count+max(tms) over tms tables EXCEPT volatile churn (llx_const,
llx_user, session/cron); docs EXCLUDE */temp/* (Dolibarr stats cache) — from both
the fingerprint and the tar. Proven live: 1st run uploads both, immediate 2nd run
skips both (uploaded=0).
Automation: the in-container logic moves to chart/files/backup-job.sh (single
source of truth, read by the orchestrator AND the chart). New
chart/templates/backup-cronjob.yaml renders a daily CronJob + ConfigMap +
VaultStaticSecret, gated by backup.enabled (default false). Helm-verified: off by
default (0 CronJobs), on renders correctly, env-aware (PREFIX erp/prod vs
erp/sandbox), script embedded.
Activation (documented): store GCS HMAC creds at kvv2/<backup.vaultS3Path>
(default erp/backup), grant the erp `auth` Vault role read on it (tools change),
set backup.enabled=true. Until then the orchestrator runs on demand.
Co-Authored-By: Claude Opus 4.7 (1M context) <[email protected]>
The accounting data + issued documents are legally retained 10 years and warrant a
backup dedicated to Dolibarr. An audit found the generic Longhorn external backup
NEVER covered the erp volume (its Longhorn volume sits in the orphaned `default`
recurring-job group; the only job has groups=[] → serves nothing; lastBackupAt=never).
So /var/www/documents (invoice PDFs, supplier pieces, contracts, ECM) had zero
offsite copy — only in-cluster replicas.
ops/backup/dolibarr-backup.sh (orchestrator) + ops/backup/backup-job.sh (in-container
logic, env-driven, single source of truth):
- pg_dump -Fc of the DB + tar of the documents PVC (RWX, read-only mount) ->
s3://arcodange-backup/erp/<env>/{db,docs}/<ts>, then tiered prune (daily 30d /
monthly 12m / yearly 10y).
- prod is READ-only (dump+tar read; writes go only to the backup bucket); the DB is
read with the env's own dynamic creds; the GCS HMAC secret is copied transiently
(base64, deleted on exit) and never printed; the whole script ships base64.
- fixes the aws-cli v2.23+ default-checksum incompatibility with GCS/S3-compat
(SignatureDoesNotMatch) via AWS_*_CHECKSUM_*=when_required.
Proven live: sandbox end-to-end (dump+tar+upload+prune, verified in GCS, cleaned up)
and retention logic unit-tested (1100 daily -> 46 kept). The FIRST real prod backup
was taken (erp/prod/db 1.2 MB + erp/prod/docs 12.5 MB) — closing the gap now.
Automation (recurring CronJob in the chart + a dedicated erp Vault policy for its
own S3 creds) is the documented next step; the orchestrator works today on demand.
Co-Authored-By: Claude Opus 4.7 (1M context) <[email protected]>
refresh-from-prod was structurally broken and silently no-op'd the restore:
1. pg_restore lacked -U, so the postgres image connected as its OS user `root`
and auth-failed. The failure was swallowed by `|| echo "ignorable warnings"`,
so the script reported success while the DROP OWNED had already emptied the DB.
E2's original seed was a manual process, so this path had never really run.
Fix: pass `-h $PGHOST -U $SB_PGUSER`; don't trust pg_restore's exit code (it
returns non-zero on the harmless "schema public already exists" notice) — verify
by counting restored llx_* tables and FAIL the Job if < 250.
2. erp-sandbox is ArgoCD-managed with self-heal ON, which reverts the
`kubectl scale --replicas=0` within seconds — so the seed ran with Dolibarr
still connected. Fix: pause self-heal for the duration, re-arm it after; app
restore + self-heal restoration + secret cleanup are guarded by an EXIT trap so
an interrupt can't strand the sandbox at replicas=0 / self-heal off.
Validated end-to-end on the live sandbox: 295 llx tables, company=Arcodange,
owner=erp_sandbox_role, self-heal re-armed, pod 1/1. README documents the self-heal
pause and the iso-prod consequence (ai_agent_sandbox is wiped → re-provision).
Co-Authored-By: Claude Opus 4.7 (1M context) <[email protected]>
Productionizes the sandbox state-lifecycle mechanisms validated live against
erp-sandbox. `ops/sandbox/sandbox-lifecycle.sh`:
- refresh-from-prod: read-only pg_dump of prod erp (default_transaction_read_only)
-> DROP OWNED BY erp_sandbox_role CASCADE -> pg_restore into erp-sandbox, using
the sandbox's own membership creds (no DROP/CREATE DATABASE, no CREATEDB, no
superuser). Dumps the full public schema (so app helper functions + triggers
come over) and filters the provisioner-owned pgbouncer user_lookup function
from the restore TOC. Scales the pod to 0 for exclusive access; copies prod
creds into a transient secret that is deleted on exit.
- sync-documents: tar-pipe the documents/mycompany tree (company logo + uploads)
prod -> sandbox, since uploaded files live on the PVC, not the DB.
Prod integrity is structural: prod is read-only during dump; the restore can only
write erp-sandbox (erp_sandbox_role owns only the sandbox DB and cannot drop prod
erp/erp_role); the platform's only prod-capable superuser stays behind the
human-gated postgres.yaml CI and is never used here.
README documents the integrity guarantee, the encryption + PVC fidelity caveats,
the BDD reset loop, and the hardening backlog (dedicated read-only dump role,
golden-cache PVC).
Refs ADR-0003 (factory#19). Chart owner-role fix = erp#13.
Co-Authored-By: Claude Opus 4.7 (1M context) <[email protected]>