Files
erp/ops/backup/README.md
T
arcodangeandClaude Opus 4.7 7dcd982448 fix(sandbox-poc): regenerate api_key when the existing one is garbage; note platform backup gaps
After an iso-prod refresh the instance unique-id changes, so an api_key encrypted
with the OLD id can't be decrypted — Dolibarr renders non-UTF-8 bytes in the field.
The POC's generateApiKey reused any non-empty value, so it copied that garbage into
test/.ai_agent_sandbox.key (corrupt key, 401s). Now it reuses ONLY a clean key
(^[A-Za-z0-9_-]{24,}$); otherwise it clears the field and regenerates. So
`checkpoint provision` after a refresh yields a fresh, working key.

Also documents the open PLATFORM follow-ups in ops/backup/README.md (easy to find
when revisiting ERP backups): the orphaned Longhorn `default` recurring-job group
(other cluster volumes have no offsite backup), and verifying the factory
pg_dumpall host cron.

Co-Authored-By: Claude Opus 4.7 (1M context) <[email protected]>
2026-06-30 22:44:40 +02:00

106 lines
5.4 KiB
Markdown

# Dolibarr dedicated backup
A backup strategy **dedicated to Dolibarr**, because the accounting data and the
issued documents are critical and legally retained **10 years** — they warrant more
than the generic platform backup.
## Why this exists (the gap it closes)
On 2026-06-30 an audit of the Longhorn external backup found that **the erp documents
volume had never been backed up offsite** (`lastBackupAt = never`): its Longhorn
volume is enrolled only in the `default` recurring-job group, but the single backup
job (`thrice-a-month-backup`) has `groups=[]`, so it serves *no* group — the erp
volume (and erp-sandbox) fell through the crack. Only in-cluster Longhorn replicas
protected `/var/www/documents` (issued invoice PDFs, supplier pieces, contracts, ECM)
— which does not survive a cluster loss / corruption / power-cut.
This tool backs up **both halves** of Dolibarr state to the existing object store
(`s3://arcodange-backup`, GCS via the S3-compatible API), under `erp/<env>/`:
| half | how | key |
|---|---|---|
| Postgres DB | `pg_dump -Fc` (restorable) | `erp/<env>/db/<ts>.dump` |
| documents PVC | `tar -czf` of `/var/www/documents` (RWX, mounted read-only) | `erp/<env>/docs/<ts>.tar.gz` |
then prunes to a **tiered retention**: daily for 30 days, monthly for 12 months,
yearly for ~10 years.
**Skip-if-unchanged:** each half carries a content fingerprint at `erp/<env>/.fp-{db,docs}`
and is dumped+uploaded only if it **differs** from the last run — so a quiet ERP day
re-uploads nothing. The fingerprint is over **durable business content only**: the DB
side is `count + max(tms)` over every `tms` table *except* volatile ones (`llx_const`,
`llx_user`, sessions/cron), and the documents side excludes `*/temp/*` (Dolibarr's
constantly-regenerated stats cache) — from both the fingerprint *and* the tar.
## Safety (mirrors `ops/sandbox/sandbox-lifecycle.sh`)
- **prod is read-only**: `pg_dump` and `tar` only read; the only writes go to the
backup bucket, never to prod. The DB is read with the env's *own* dynamic creds
(`vso-db-credentials`); prod and sandbox never cross.
- **S3 creds are never exposed**: the GCS HMAC secret is copied into a *transient*
secret in the app namespace (values stay base64), deleted on exit. The whole
in-container script is shipped base64 — no secret is ever printed.
## Usage
```sh
# one-shot backup + prune (run from anywhere; needs kubectl on the lab cluster)
ops/backup/dolibarr-backup.sh backup --env prod
ops/backup/dolibarr-backup.sh backup --env sandbox
# what's in the store
ops/backup/dolibarr-backup.sh list --env prod
```
`chart/files/backup-job.sh` is the in-container logic (env-driven: `BUCKET PREFIX
DB PGHOST` + the mounted DB/S3 creds) — the single source of truth shared by this
orchestrator and the scheduled CronJob (see "Automation" below).
**Status:** the first real prod backup was taken 2026-06-30
(`erp/prod/db/…` 1.2 MB, `erp/prod/docs/…` 12.5 MB). Proven end-to-end live on the
sandbox (dump + tar + GCS upload + retention prune).
## Restore
```sh
ops/backup/dolibarr-backup.sh list --env <e> # find the <ts> key
ops/backup/dolibarr-backup.sh restore --db <ts>.dump --env <e> --yes
ops/backup/dolibarr-backup.sh restore --docs <ts>.tar.gz --env <e> --yes
```
**DESTRUCTIVE** (requires `--yes` + an explicit `--env`): scales the app to 0, then
- **`--db`**: `DROP OWNED BY <owner_role> CASCADE` + `pg_restore --no-owner --role`
(same reset mechanics as `ops/sandbox/sandbox-lifecycle.sh`), then scales back;
- **`--docs`**: clears `/var/www/documents` and untars the archive, then scales back.
The key is the bare `<ts>` filename from `list`; `--db`/`--docs` selects the
`db/` or `docs/` subpath. Proven on the sandbox: a mutated `MAIN_INFO_SOCIETE_NOM`
was reverted to the backup's value by `restore --db`.
## Automation — the CronJob (gated on creds)
The recurring form ships in the chart (`chart/templates/backup-cronjob.yaml`,
`backup.enabled=false` by default): a daily **CronJob** (ConfigMap-mounted
`backup-job.sh`) with its **own** S3 creds via a `VaultStaticSecret` — no
cross-namespace borrowing of the Longhorn secret.
**Status: LIVE (2026-06-30).** tools#5 granted the erp prod Vault policy read on
`kvv2/data/longhorn/gcs-backup` (via the module's `kv_read_paths`), and the chart
sets `backup.enabled: true` + `vaultS3Path: longhorn/gcs-backup`. The CronJob runs
daily at 03:00 UTC; a manual run was verified (db + docs → GCS). Sandbox keeps it off.
## Platform follow-ups — NOT erp (open; revisit when we look at ERP backups)
1. **Longhorn `default` group is orphaned → other volumes have NO offsite backup.**
The cluster's only Longhorn recurring-backup job (`thrice-a-month-backup`) has
`groups=[]`, so it serves no group — yet volumes are enrolled in the `default`
group. Result found 2026-06-30: erp + most volumes had `lastBackupAt=never`. This
dedicated backup covers Dolibarr, but every OTHER cluster volume is still exposed.
Fix at the platform layer (factory): attach a recurring job to the `default` group
(or set the storageclass `recurringJobSelector`).
2. **Verify the factory `pg_dumpall` host cron actually runs.** `factory` ships
`ansible/.../playbooks/backup/postgres.yml` that `pg_dumpall`s daily on the
Postgres host (192.168.1.202) — unverified from the k8s side. Confirm it's live
and where it ships; it's the platform-level DB safety net beneath this app-level
one.