feat(k3s) — ramasse-miettes d'images du kubelet à 65 % → 55 % : il nettoyait APRÈS la ligne de Longhorn
Le kubelet ne supprimait les images inutilisées qu'à 85 % d'occupation du
disque des images, et s'arrêtait à 80 %. Ce disque est le disque externe
(/mnt/arcodange), que Longhorn retire de l'ordonnancement dès 75 %
occupés : le nettoyage laissait donc le disque en permanence au-dessus de
la ligne de Longhorn. Relevé le 2026-09-25 sur pi3 : 82 %, disque
Longhorn Schedulable=False, 272 anciennes images du front Kadans.
Ajout de --kubelet-arg=image-gc-high-threshold=65 et
--kubelet-arg=image-gc-low-threshold=55 au serveur (pi1) ET aux agents
(pi2, pi3). Le mécanisme s'applique bien au runtime Docker : k3s --docker
passe par cri-dockerd, qui rend au kubelet le système de fichiers des
images et exécute ses suppressions (371 images déjà supprimées sur pi3,
kubelet_image_garbage_collected_total{reason="space"}).
Indépendant du disque externe : sans lui, les images tombent sur la carte
SD, et le même seuil la protège.
Co-Authored-By: Claude Opus 5.5 <[email protected]>
This commit is contained in:
@@ -42,14 +42,40 @@
|
||||
# redémarré depuis avril, donc jamais régénéré par la version actuelle du
|
||||
# script), révélé par le premier restart forcé par ce playbook. La forme
|
||||
# SANS guillemets (`--kubelet-arg=k=v`) traverse la génération intacte.
|
||||
#
|
||||
# Ramasse-miettes d'images du kubelet : haut 85 → 65 %, bas 80 → 55 %.
|
||||
# Les images Docker (data-root /mnt/arcodange/docker) partagent le disque
|
||||
# externe avec Longhorn (/mnt/arcodange/longhorn), qui retire le disque de
|
||||
# l'ordonnancement sous 25 % libres, soit dès 75 % occupés. Au défaut, le
|
||||
# kubelet ne nettoie qu'à 85 % et s'arrête à 80 % : il laisse le disque
|
||||
# EN PERMANENCE au-dessus de la ligne de Longhorn. Relevé le 2026-09-25 :
|
||||
# pi3 à 82 % (491 Go), disque Longhorn « Schedulable=False », 272
|
||||
# anciennes images du front Kadans (~450 Mo chacune) jamais retirées.
|
||||
# 65 % laisse 10 points de marge sous la ligne de Longhorn.
|
||||
#
|
||||
# ⚠ Ça s'applique BIEN au runtime Docker : `--docker` fait passer le
|
||||
# kubelet par cri-dockerd (containerRuntimeEndpoint
|
||||
# unix:///run/k3s/cri-dockerd/cri-dockerd.sock), qui lui rend le système
|
||||
# de fichiers des images (491 Go = le disque externe) et exécute ses
|
||||
# suppressions — `kubelet_image_garbage_collected_total{reason="space"}`
|
||||
# valait 371 sur pi3 le 2026-09-25. Une image référencée par un conteneur,
|
||||
# même arrêté (épingle de ci-node-playwright, runner, Gitea et Postgres en
|
||||
# compose), est refusée par Docker : le kubelet passe à la suivante.
|
||||
#
|
||||
# Ne dépend PAS du disque externe : sans lui, les images tombent sur la
|
||||
# carte SD, et le même seuil la protège.
|
||||
extra_server_args: >-
|
||||
--docker --disable traefik
|
||||
--kubelet-arg=container-log-max-files=5
|
||||
--kubelet-arg=container-log-max-size=10Mi
|
||||
--kubelet-arg=image-gc-high-threshold=65
|
||||
--kubelet-arg=image-gc-low-threshold=55
|
||||
extra_agent_args: >-
|
||||
--docker
|
||||
--kubelet-arg=container-log-max-files=5
|
||||
--kubelet-arg=container-log-max-size=10Mi
|
||||
--kubelet-arg=image-gc-high-threshold=65
|
||||
--kubelet-arg=image-gc-low-threshold=55
|
||||
{{ kubelet_reserved_args | default('') }}
|
||||
api_endpoint: "{{ hostvars[groups['server'][0]]['ansible_host'] | default(groups['server'][0]) }}"
|
||||
|
||||
|
||||
@@ -3,7 +3,7 @@
|
||||
# 01 · System — base OS, Docker, K3s, Longhorn, DNS, SSL
|
||||
|
||||
> [!NOTE]
|
||||
> **Status:** ✅ active · **Last Updated:** 2026-06-23
|
||||
> **Status:** ✅ active · **Last Updated:** 2026-09-25
|
||||
> **Upstream:** [Ansible sub-hub](README.md) · [Factory provisioning hub](../README.md)
|
||||
> **Downstream:** [02 · Setup](02-setup.md) · [03 · CI/CD](03-cicd.md)
|
||||
> **Related:** [Storage & recovery](../../lab-ecosystem/storage-and-recovery.md) · [Secrets & Vault](../../lab-ecosystem/secrets-and-vault.md) · [Naming conventions](../../lab-ecosystem/naming-conventions.md) · [ADR-0001 safe prod-like environment](../../../ADR/0001-safe-prod-like-environment.md)
|
||||
@@ -24,7 +24,7 @@ All host-facing plays target `raspberries:&local` — the intersection of the `r
|
||||
| 4 | [`system/prepare_disks.yml`](../../../../ansible/arcodange/factory/playbooks/system/prepare_disks.yml) | Auto-detect the largest external (non-`mmcblk0`) USB partition, format it **ext4 with label `arcodange_500`**, mount at `/mnt/arcodange`, and persist in `fstab`. Skips format if the label already exists. **`pause` confirm before any format.** | `mount_point: /mnt/arcodange`, `disk_label: arcodange_500` |
|
||||
| 5 | [`system/system_docker.yml`](../../../../ansible/arcodange/factory/playbooks/system/system_docker.yml) | Install Docker via `geerlingguy.docker`; write `daemon.json` with **json-file logging** (`max-size 10m`, `max-file 5`) and **`data-root: /mnt/arcodange/docker`** (only when the external disk is mounted). | `tags: never`; `storage-driver: overlay2` |
|
||||
| 6 | [`system/iscsi_longhorn.yml`](../../../../ansible/arcodange/factory/playbooks/system/iscsi_longhorn.yml) | Install `open-iscsi` (+ enable `iscsid`) and `cryptsetup`, and load the **`dm_crypt`** kernel module (persisted in `/etc/modules`) — Longhorn's encrypted-volume prerequisites. Creates `/mnt/arcodange/longhorn`. | module `dm_crypt` |
|
||||
| 7 | [`system/system_k3s.yml`](../../../../ansible/arcodange/factory/playbooks/system/system_k3s.yml) | Build the K3s inventory dynamically (first sorted host → `server`, rest → `agent`), install the `k3s-ansible` content, run `k3s.orchestration.site`, then **fetch the kubeconfig** to `~/.kube/config` (rewriting `127.0.0.1` → server IP). | **k3s `v1.34.3+k3s1`**; server args `--docker --disable traefik` |
|
||||
| 7 | [`system/system_k3s.yml`](../../../../ansible/arcodange/factory/playbooks/system/system_k3s.yml) | Build the K3s inventory dynamically (first sorted host → `server`, rest → `agent`), install the `k3s-ansible` content, run `k3s.orchestration.site`, then **fetch the kubeconfig** to `~/.kube/config` (rewriting `127.0.0.1` → server IP). | **k3s `v1.34.3+k3s1`**; server args `--docker --disable traefik`; kubelet args on server **and** agents: container logs capped (`5 × 10Mi`), **image GC at 65 % → 55 %** (below Longhorn's 75 % disk-pressure line) |
|
||||
| 8 | [`system/k3s_dns.yml`](../../../../ansible/arcodange/factory/playbooks/system/k3s_dns.yml) | Create the **`coredns-custom`** ConfigMap so cluster DNS forwards `arcodange.lab:53` to the Pi-hole IPs; also patch the main CoreDNS Corefile to forward to the same HA Pi-holes. | `pihole_ips` (extracted from hostvars) |
|
||||
| 9 | [`system/k3s_ssl.yml`](../../../../ansible/arcodange/factory/playbooks/system/k3s_ssl.yml) | Deploy **cert-manager** + **step-issuer** as k3s static HelmCharts; create the `StepClusterIssuer` `step-ca` wired to the JWK provisioner and root CA. | cert-manager `v1.19.2`, step-issuer `1.9.11`, `caUrl: https://ssl-ca.arcodange.lab:8443`, **ARM64 `kube-rbac-proxy` override** |
|
||||
| 10 | [`system/k3s_config.yml`](../../../../ansible/arcodange/factory/playbooks/system/k3s_config.yml) | Deploy **Longhorn** + **Traefik** as HelmCharts; issue the wildcard cert, set the default `TLSStore`, wire Gitea, the IP-allow-list middleware, and the CrowdSec bouncer plugin; then **delete the old Traefik** to force a redeploy. | Longhorn `v1.9.1`, Traefik `v37.4.0` (see detail below) |
|
||||
@@ -58,7 +58,7 @@ flowchart TD
|
||||
4. **`prepare_disks.yml`** formats and mounts the external USB disk at `/mnt/arcodange` (with a confirmation pause).
|
||||
5. **Docker** installs with its data-root pointed at that disk and capped logging.
|
||||
6. **iSCSI + dm_crypt** prerequisites land so Longhorn can attach (and encrypt) volumes.
|
||||
7. **K3s** installs with the first host as server, Docker as the container runtime, and Traefik disabled.
|
||||
7. **K3s** installs with the first host as server, Docker as the container runtime (through k3s's embedded cri-dockerd), and Traefik disabled; the kubelet garbage-collects unused images from 65 % disk use down to 55 %, before Longhorn stops scheduling on the shared external disk at 75 %.
|
||||
8. **CoreDNS** is reconfigured to forward `arcodange.lab` to the Pi-holes.
|
||||
9. **cert-manager + step-issuer** wire the in-cluster issuer to step-ca.
|
||||
10. **`k3s_config.yml`** deploys Longhorn and a fully-customized Traefik, then deletes the old Traefik so the helm-controller redeploys with the new config.
|
||||
|
||||
Reference in New Issue
Block a user