Commit Graph
10 Commits
Author SHA1 Message Date
arcodangeandClaude Opus 5 6a5b7baa23 chore: réconcilier le trunk avec main, et corriger la note du mode vierge
DEUX CHOSES.

1. RÉCONCILIATION. Le trunk accusait 49 commits de retard, 0 d'avance : tout son
delta vivait dans l'arbre de travail, mélange d'améliorations réelles et d'état
périmé. Il est remis sur origin/main, et son état est préservé sur la branche
trunk-snapshot-20260814. Ce qui méritait d'être porté l'est ici :

  - test/main.ts : chemin Work/ au lieu de Desktop/ (l'ancien n'existe plus) et
    module TIERS activé. L'URL par défaut reste corrigée par la garde d'hôte ;
  - ansible README : la commande du playbook recurrentBackup ;
  - company.json : siteWeb arcodange.fr et email @arcodange.fr — le trunk était
    à jour sur ces deux champs quand main portait encore duckdns et gmail.

NE SONT PAS PORTÉS, délibérément :
  - le .gitignore du trunk, qui avait PERDU les règles replay-packs ;
  - company.json SAS / 6201Z, périmé — main porte SARL / 6201Z depuis le PR #77 ;
  - un console.log de débogage dans forms.ts ;
  - createEphemeralAgent + permissions.ts + orchestratorExample.ts : travail
    cohérent mais qui DOUBLE le modèle de scopes durables (scopes.ts +
    provisionAiUser.ts) effectivement livré. Deux conceptions concurrentes de la
    même chose ne se portent pas en silence — arbitrage opérateur.

2. LA NOTE DU MODE VIERGE ÉTAIT FAUSSE. Elle annonçait qu'il manquait un « golden
empty ». Rien ne manquait : la chaîne existe en entier dans test/, et c'est
main.ts qui pilote l'installeur. J'avais proposé de construire ce qui était déjà
là, faute d'avoir regardé — exactement le contrôle inscrit dans AGENTS.md le
matin même. La note décrit désormais la chaîne réelle :

  blank --yes  ->  main.ts  ->  install.lock  ->  provisionSandbox.ts

Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
2026-08-14 22:04:43 +02:00
arcodangeandClaude Opus 5 bc5e0603d1 docs(ops): ne jamais éditer le script de cycle de vie pendant son exécution
Bash lit ses scripts au fil de l'eau. Éditer le fichier en cours de route décale
les offsets et corrompt l'exécution : un refresh-from-prod en cours s'est mis à
afficher les messages de blank_sandbox, puis a fini sur « syntax error ». Le
travail avait bien abouti — sandbox restaurée, 295 tables, 10 factures, 12 tiers
— mais la trace mentait sur la fonction qui tournait. Sur un script qui purge des
bases, une trace qui ment est inacceptable.

Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
2026-08-14 19:13:57 +02:00
arcodangeandClaude Opus 5 a3540982a9 feat(ops): garde de cluster sur la sauvegarde, et mode vierge (incomplet)
DEUX CHOSES, dont une inachevée et dite comme telle.

1. LA SAUVEGARDE N'AVAIT AUCUNE GARDE DE CLUSTER. 24 appels kubectl nus, aucun
contexte épinglé — alors que le script lit des secrets, crée des Jobs et sait
RESTAURER une base. Lancé sur le contexte courant d'une station de travail, il
serait parti chercher les secrets Longhorn du cluster d'un client. Même patron
que ops/sandbox : ERP_KUBE_CONTEXT, wrapper K(), empreinte positive vérifiée
avant tout dispatch. Testé contre le cluster client et contre un contexte
inexistant : il refuse les deux.

Sauvegarde de production passée dans la foulée. La base a été dédupliquée à
juste titre — rien n'avait été écrit depuis la sauvegarde automatique de 01:00.
Au passage, le CronJob quotidien tourne bien depuis juillet ; il était encore
noté comme à faire.

2. LE MODE VIERGE PURGE MAIS NE RECONSTRUIT PAS. `blank --yes` vide la base du
bac à sable (295 tables -> 0, vérifié) avec trois protections empilées : refus
sans --yes, garde de cluster, et une relecture de current_database() DANS le Job
lui-même — une purge sur la mauvaise base ne se rattrape pas par un refresh,
contrairement à la bonne.

Mais l'instance ne se reconstruit pas. Premier essai : install.lock vit sur le
volume documents, que la purge ne touche pas, si bien que Dolibarr servait un
login sur un schéma inexistant. Correctif appliqué — retrait du verrou puis
redémarrage. Second essai : le verrou reste absent, et le schéma reste à ZÉRO
table. L'entrypoint de l'image ne lance aucune installation automatique.

CE QU'IL MANQUE est donc nommé dans le script : un « golden empty », pg_dump
d'une instance fraîchement installée — schéma et données de référence, aucune
donnée métier. `blank` le restaurerait au lieu de laisser la base vide, comme
refresh-from-prod restaure le dump de production. Seule la source change.

En l'état `blank` laisse le bac à sable inutilisable, et le script le dit. Le
bac à sable a été remis iso-prod avant de rendre la main.

Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
2026-08-14 19:11:18 +02:00
arcodangeandClaude Opus 5 613f8b0a8f fix(ops): pin the kube-context — never run destructive steps on the ambient one
sandbox-lifecycle.sh scales deployments to zero, patches the ArgoCD Application
and runs DROP OWNED ... CASCADE. Every one of those ran against whatever
kube-context happened to be current.

This workstation also carries a CLIENT production cluster. On 2026-07-25 a
`checkpoint refresh` was issued while the current context was
do-nyc3-kissmetrics-prod-k8s-cluster: the script patched the ArgoCD Application,
scaled `erp-sandbox` to zero and copied a prod secret — all against the client's
cluster. Nothing was damaged only because that cluster has no `application` CRD
and no erp/erp-sandbox namespaces, so each call failed silently under `|| true`.
That is luck, not a control.

- ERP_KUBE_CONTEXT (default: "default") pins the target; every kubectl call now
  goes through K(), so nothing inherits the ambient context.
- assert_arcodange_cluster() proves the target by positive fingerprint — the
  erp, erp-sandbox and argocd namespaces AND the erp-sandbox ArgoCD Application.
  A client cluster cannot match all four by accident. Wired into all three
  entry points, before any mutation.

Verified: refuses the client context, refuses an unknown context, passes on the
homelab and completes normally.

Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
Claude-Session: https://claude.ai/code/session_01VRShc4QhLLU73FLHx9vskh
2026-07-25 23:43:26 +02:00
arcodangeandClaude Opus 4.7 7dcd982448 fix(sandbox-poc): regenerate api_key when the existing one is garbage; note platform backup gaps
After an iso-prod refresh the instance unique-id changes, so an api_key encrypted
with the OLD id can't be decrypted — Dolibarr renders non-UTF-8 bytes in the field.
The POC's generateApiKey reused any non-empty value, so it copied that garbage into
test/.ai_agent_sandbox.key (corrupt key, 401s). Now it reuses ONLY a clean key
(^[A-Za-z0-9_-]{24,}$); otherwise it clears the field and regenerates. So
`checkpoint provision` after a refresh yields a fresh, working key.

Also documents the open PLATFORM follow-ups in ops/backup/README.md (easy to find
when revisiting ERP backups): the orphaned Longhorn `default` recurring-job group
(other cluster volumes have no offsite backup), and verifying the factory
pg_dumpall host cron.

Co-Authored-By: Claude Opus 4.7 (1M context) <[email protected]>
2026-06-30 22:44:40 +02:00
arcodangeandClaude Opus 4.7 a8b80f17e4 feat(backup): restore subcommand (db + documents), proven on the sandbox
dolibarr-backup.sh restore --db|--docs <ts> --env <e> --yes — the recovery half of
the dedicated backup. DESTRUCTIVE (gated by --yes + explicit --env): scales the app
to 0, then
- --db: DROP OWNED BY <owner_role> CASCADE + pg_restore --no-owner --role (same
  mechanics as sandbox-lifecycle.sh), from s3://.../erp/<env>/db/<ts>.dump;
- --docs: clears /var/www/documents and untars s3://.../erp/<env>/docs/<ts>.tar.gz;
then scales the app back to 1. OWNER_ROLE per env (erp_role / erp_sandbox_role).
The key is the bare <ts> filename from `list`; --db/--docs selects the subpath.

Proven on the sandbox: backup → mutate MAIN_INFO_SOCIETE_NOM to a sentinel →
restore --db → the value reverted to the backup's ('Arcodange'). (First run caught
a path bug — the fetch missed the db/ subdir — now fixed.)

Co-Authored-By: Claude Opus 4.7 (1M context) <[email protected]>
2026-06-30 18:05:19 +02:00
arcodangeandClaude Opus 4.7 a3f0586c77 feat(backup): skip-if-unchanged + scheduled CronJob in the chart
Builds on the dedicated backup (erp#31).

Skip-if-unchanged: each half (DB / documents) carries a content fingerprint at
erp/<env>/.fp-{db,docs} and is dumped+uploaded only if it differs from the last
run — a quiet ERP day re-uploads nothing. Fingerprint = durable BUSINESS content
only: DB = count+max(tms) over tms tables EXCEPT volatile churn (llx_const,
llx_user, session/cron); docs EXCLUDE */temp/* (Dolibarr stats cache) — from both
the fingerprint and the tar. Proven live: 1st run uploads both, immediate 2nd run
skips both (uploaded=0).

Automation: the in-container logic moves to chart/files/backup-job.sh (single
source of truth, read by the orchestrator AND the chart). New
chart/templates/backup-cronjob.yaml renders a daily CronJob + ConfigMap +
VaultStaticSecret, gated by backup.enabled (default false). Helm-verified: off by
default (0 CronJobs), on renders correctly, env-aware (PREFIX erp/prod vs
erp/sandbox), script embedded.

Activation (documented): store GCS HMAC creds at kvv2/<backup.vaultS3Path>
(default erp/backup), grant the erp `auth` Vault role read on it (tools change),
set backup.enabled=true. Until then the orchestrator runs on demand.

Co-Authored-By: Claude Opus 4.7 (1M context) <[email protected]>
2026-06-30 15:53:13 +02:00
arcodangeandClaude Opus 4.7 8ec8fde67e feat(ops): dedicated Dolibarr backup (DB + documents → offsite GCS, 10y retention)
The accounting data + issued documents are legally retained 10 years and warrant a
backup dedicated to Dolibarr. An audit found the generic Longhorn external backup
NEVER covered the erp volume (its Longhorn volume sits in the orphaned `default`
recurring-job group; the only job has groups=[] → serves nothing; lastBackupAt=never).
So /var/www/documents (invoice PDFs, supplier pieces, contracts, ECM) had zero
offsite copy — only in-cluster replicas.

ops/backup/dolibarr-backup.sh (orchestrator) + ops/backup/backup-job.sh (in-container
logic, env-driven, single source of truth):
- pg_dump -Fc of the DB + tar of the documents PVC (RWX, read-only mount) ->
  s3://arcodange-backup/erp/<env>/{db,docs}/<ts>, then tiered prune (daily 30d /
  monthly 12m / yearly 10y).
- prod is READ-only (dump+tar read; writes go only to the backup bucket); the DB is
  read with the env's own dynamic creds; the GCS HMAC secret is copied transiently
  (base64, deleted on exit) and never printed; the whole script ships base64.
- fixes the aws-cli v2.23+ default-checksum incompatibility with GCS/S3-compat
  (SignatureDoesNotMatch) via AWS_*_CHECKSUM_*=when_required.

Proven live: sandbox end-to-end (dump+tar+upload+prune, verified in GCS, cleaned up)
and retention logic unit-tested (1100 daily -> 46 kept). The FIRST real prod backup
was taken (erp/prod/db 1.2 MB + erp/prod/docs 12.5 MB) — closing the gap now.

Automation (recurring CronJob in the chart + a dedicated erp Vault policy for its
own S3 creds) is the documented next step; the orchestrator works today on demand.

Co-Authored-By: Claude Opus 4.7 (1M context) <[email protected]>
2026-06-30 15:32:36 +02:00
arcodangeandClaude Opus 4.7 434be7488d fix(ops): sandbox refresh-from-prod actually restores now (pg_restore -U + self-heal pause)
refresh-from-prod was structurally broken and silently no-op'd the restore:

1. pg_restore lacked -U, so the postgres image connected as its OS user `root`
   and auth-failed. The failure was swallowed by `|| echo "ignorable warnings"`,
   so the script reported success while the DROP OWNED had already emptied the DB.
   E2's original seed was a manual process, so this path had never really run.
   Fix: pass `-h $PGHOST -U $SB_PGUSER`; don't trust pg_restore's exit code (it
   returns non-zero on the harmless "schema public already exists" notice) — verify
   by counting restored llx_* tables and FAIL the Job if < 250.

2. erp-sandbox is ArgoCD-managed with self-heal ON, which reverts the
   `kubectl scale --replicas=0` within seconds — so the seed ran with Dolibarr
   still connected. Fix: pause self-heal for the duration, re-arm it after; app
   restore + self-heal restoration + secret cleanup are guarded by an EXIT trap so
   an interrupt can't strand the sandbox at replicas=0 / self-heal off.

Validated end-to-end on the live sandbox: 295 llx tables, company=Arcodange,
owner=erp_sandbox_role, self-heal re-armed, pod 1/1. README documents the self-heal
pause and the iso-prod consequence (ai_agent_sandbox is wiped → re-provision).

Co-Authored-By: Claude Opus 4.7 (1M context) <[email protected]>
2026-06-30 06:59:39 +02:00
arcodangeandClaude Opus 4.7 7264f00ed4 feat(ops): erp-sandbox iso-prod seed + documents sync tooling (ADR-0003 E2)
Productionizes the sandbox state-lifecycle mechanisms validated live against
erp-sandbox. `ops/sandbox/sandbox-lifecycle.sh`:
  - refresh-from-prod: read-only pg_dump of prod erp (default_transaction_read_only)
    -> DROP OWNED BY erp_sandbox_role CASCADE -> pg_restore into erp-sandbox, using
    the sandbox's own membership creds (no DROP/CREATE DATABASE, no CREATEDB, no
    superuser). Dumps the full public schema (so app helper functions + triggers
    come over) and filters the provisioner-owned pgbouncer user_lookup function
    from the restore TOC. Scales the pod to 0 for exclusive access; copies prod
    creds into a transient secret that is deleted on exit.
  - sync-documents: tar-pipe the documents/mycompany tree (company logo + uploads)
    prod -> sandbox, since uploaded files live on the PVC, not the DB.

Prod integrity is structural: prod is read-only during dump; the restore can only
write erp-sandbox (erp_sandbox_role owns only the sandbox DB and cannot drop prod
erp/erp_role); the platform's only prod-capable superuser stays behind the
human-gated postgres.yaml CI and is never used here.

README documents the integrity guarantee, the encryption + PVC fidelity caveats,
the BDD reset loop, and the hardening backlog (dedicated read-only dump role,
golden-cache PVC).

Refs ADR-0003 (factory#19). Chart owner-role fix = erp#13.

Co-Authored-By: Claude Opus 4.7 (1M context) <[email protected]>
2026-06-29 07:42:00 +02:00