Commit Graph
5 Commits
Author SHA1 Message Date
arcodangeandClaude Opus 5 a3540982a9 feat(ops): garde de cluster sur la sauvegarde, et mode vierge (incomplet)
DEUX CHOSES, dont une inachevée et dite comme telle.

1. LA SAUVEGARDE N'AVAIT AUCUNE GARDE DE CLUSTER. 24 appels kubectl nus, aucun
contexte épinglé — alors que le script lit des secrets, crée des Jobs et sait
RESTAURER une base. Lancé sur le contexte courant d'une station de travail, il
serait parti chercher les secrets Longhorn du cluster d'un client. Même patron
que ops/sandbox : ERP_KUBE_CONTEXT, wrapper K(), empreinte positive vérifiée
avant tout dispatch. Testé contre le cluster client et contre un contexte
inexistant : il refuse les deux.

Sauvegarde de production passée dans la foulée. La base a été dédupliquée à
juste titre — rien n'avait été écrit depuis la sauvegarde automatique de 01:00.
Au passage, le CronJob quotidien tourne bien depuis juillet ; il était encore
noté comme à faire.

2. LE MODE VIERGE PURGE MAIS NE RECONSTRUIT PAS. `blank --yes` vide la base du
bac à sable (295 tables -> 0, vérifié) avec trois protections empilées : refus
sans --yes, garde de cluster, et une relecture de current_database() DANS le Job
lui-même — une purge sur la mauvaise base ne se rattrape pas par un refresh,
contrairement à la bonne.

Mais l'instance ne se reconstruit pas. Premier essai : install.lock vit sur le
volume documents, que la purge ne touche pas, si bien que Dolibarr servait un
login sur un schéma inexistant. Correctif appliqué — retrait du verrou puis
redémarrage. Second essai : le verrou reste absent, et le schéma reste à ZÉRO
table. L'entrypoint de l'image ne lance aucune installation automatique.

CE QU'IL MANQUE est donc nommé dans le script : un « golden empty », pg_dump
d'une instance fraîchement installée — schéma et données de référence, aucune
donnée métier. `blank` le restaurerait au lieu de laisser la base vide, comme
refresh-from-prod restaure le dump de production. Seule la source change.

En l'état `blank` laisse le bac à sable inutilisable, et le script le dit. Le
bac à sable a été remis iso-prod avant de rendre la main.

Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
2026-08-14 19:11:18 +02:00
arcodangeandClaude Opus 4.7 7dcd982448 fix(sandbox-poc): regenerate api_key when the existing one is garbage; note platform backup gaps
After an iso-prod refresh the instance unique-id changes, so an api_key encrypted
with the OLD id can't be decrypted — Dolibarr renders non-UTF-8 bytes in the field.
The POC's generateApiKey reused any non-empty value, so it copied that garbage into
test/.ai_agent_sandbox.key (corrupt key, 401s). Now it reuses ONLY a clean key
(^[A-Za-z0-9_-]{24,}$); otherwise it clears the field and regenerates. So
`checkpoint provision` after a refresh yields a fresh, working key.

Also documents the open PLATFORM follow-ups in ops/backup/README.md (easy to find
when revisiting ERP backups): the orphaned Longhorn `default` recurring-job group
(other cluster volumes have no offsite backup), and verifying the factory
pg_dumpall host cron.

Co-Authored-By: Claude Opus 4.7 (1M context) <[email protected]>
2026-06-30 22:44:40 +02:00
arcodangeandClaude Opus 4.7 a8b80f17e4 feat(backup): restore subcommand (db + documents), proven on the sandbox
dolibarr-backup.sh restore --db|--docs <ts> --env <e> --yes — the recovery half of
the dedicated backup. DESTRUCTIVE (gated by --yes + explicit --env): scales the app
to 0, then
- --db: DROP OWNED BY <owner_role> CASCADE + pg_restore --no-owner --role (same
  mechanics as sandbox-lifecycle.sh), from s3://.../erp/<env>/db/<ts>.dump;
- --docs: clears /var/www/documents and untars s3://.../erp/<env>/docs/<ts>.tar.gz;
then scales the app back to 1. OWNER_ROLE per env (erp_role / erp_sandbox_role).
The key is the bare <ts> filename from `list`; --db/--docs selects the subpath.

Proven on the sandbox: backup → mutate MAIN_INFO_SOCIETE_NOM to a sentinel →
restore --db → the value reverted to the backup's ('Arcodange'). (First run caught
a path bug — the fetch missed the db/ subdir — now fixed.)

Co-Authored-By: Claude Opus 4.7 (1M context) <[email protected]>
2026-06-30 18:05:19 +02:00
arcodangeandClaude Opus 4.7 a3f0586c77 feat(backup): skip-if-unchanged + scheduled CronJob in the chart
Builds on the dedicated backup (erp#31).

Skip-if-unchanged: each half (DB / documents) carries a content fingerprint at
erp/<env>/.fp-{db,docs} and is dumped+uploaded only if it differs from the last
run — a quiet ERP day re-uploads nothing. Fingerprint = durable BUSINESS content
only: DB = count+max(tms) over tms tables EXCEPT volatile churn (llx_const,
llx_user, session/cron); docs EXCLUDE */temp/* (Dolibarr stats cache) — from both
the fingerprint and the tar. Proven live: 1st run uploads both, immediate 2nd run
skips both (uploaded=0).

Automation: the in-container logic moves to chart/files/backup-job.sh (single
source of truth, read by the orchestrator AND the chart). New
chart/templates/backup-cronjob.yaml renders a daily CronJob + ConfigMap +
VaultStaticSecret, gated by backup.enabled (default false). Helm-verified: off by
default (0 CronJobs), on renders correctly, env-aware (PREFIX erp/prod vs
erp/sandbox), script embedded.

Activation (documented): store GCS HMAC creds at kvv2/<backup.vaultS3Path>
(default erp/backup), grant the erp `auth` Vault role read on it (tools change),
set backup.enabled=true. Until then the orchestrator runs on demand.

Co-Authored-By: Claude Opus 4.7 (1M context) <[email protected]>
2026-06-30 15:53:13 +02:00
arcodangeandClaude Opus 4.7 8ec8fde67e feat(ops): dedicated Dolibarr backup (DB + documents → offsite GCS, 10y retention)
The accounting data + issued documents are legally retained 10 years and warrant a
backup dedicated to Dolibarr. An audit found the generic Longhorn external backup
NEVER covered the erp volume (its Longhorn volume sits in the orphaned `default`
recurring-job group; the only job has groups=[] → serves nothing; lastBackupAt=never).
So /var/www/documents (invoice PDFs, supplier pieces, contracts, ECM) had zero
offsite copy — only in-cluster replicas.

ops/backup/dolibarr-backup.sh (orchestrator) + ops/backup/backup-job.sh (in-container
logic, env-driven, single source of truth):
- pg_dump -Fc of the DB + tar of the documents PVC (RWX, read-only mount) ->
  s3://arcodange-backup/erp/<env>/{db,docs}/<ts>, then tiered prune (daily 30d /
  monthly 12m / yearly 10y).
- prod is READ-only (dump+tar read; writes go only to the backup bucket); the DB is
  read with the env's own dynamic creds; the GCS HMAC secret is copied transiently
  (base64, deleted on exit) and never printed; the whole script ships base64.
- fixes the aws-cli v2.23+ default-checksum incompatibility with GCS/S3-compat
  (SignatureDoesNotMatch) via AWS_*_CHECKSUM_*=when_required.

Proven live: sandbox end-to-end (dump+tar+upload+prune, verified in GCS, cleaned up)
and retention logic unit-tested (1100 daily -> 46 kept). The FIRST real prod backup
was taken (erp/prod/db 1.2 MB + erp/prod/docs 12.5 MB) — closing the gap now.

Automation (recurring CronJob in the chart + a dedicated erp Vault policy for its
own S3 creds) is the documented next step; the orchestrator works today on demand.

Co-Authored-By: Claude Opus 4.7 (1M context) <[email protected]>
2026-06-30 15:32:36 +02:00