feat(harness): gated promote pipeline — rehearse, judge, human gate, apply, judge #79
@@ -0,0 +1,72 @@
|
|||||||
|
# fleet/harness/promote/ — the gated pipeline
|
||||||
|
|
||||||
|
Rehearse on the sandbox → an independent agent judges → **a human decides** →
|
||||||
|
production → a second agent verifies what actually landed.
|
||||||
|
|
||||||
|
## Why this exists as code
|
||||||
|
|
||||||
|
The discipline was already written down — [ADR-0003](https://gitea.arcodange.lab/arcodange-org/factory/src/branch/main/vibe/ADR/0003-sandbox-state-lifecycle.md),
|
||||||
|
the operating rules in [`AGENTS.md`](../../../AGENTS.md), the promote flow — and it
|
||||||
|
still depended on whoever was driving choosing to follow it. On 2026-07-25 an
|
||||||
|
agent session wrote five documents into the production ledger through direct API
|
||||||
|
calls, bypassing the promote flow entirely. Nothing was wrong with the result;
|
||||||
|
everything was wrong with the path. A rule an operator can skip is a
|
||||||
|
recommendation.
|
||||||
|
|
||||||
|
So the stages here are **chained by artefacts**, not by good intentions. Each
|
||||||
|
stage refuses to run until the previous one has produced its file, and the file
|
||||||
|
has to say what the stage needs to hear:
|
||||||
|
|
||||||
|
| Stage | Produces | Refuses unless |
|
||||||
|
| --- | --- | --- |
|
||||||
|
| 1 `rehearse` | `01-rehearsal.json` | the target is the sandbox (host-guarded) |
|
||||||
|
| 2 `judge --pre` | `02-pre-verdict.json` | a rehearsal exists and its writes succeeded |
|
||||||
|
| 3 `gate` | `03-gate.json` | a pre-verdict exists; **a human types the decision** |
|
||||||
|
| 4 `apply` | `04-applied.json` | the gate says `approved`, by a named human, for *this* manifest |
|
||||||
|
| 5 `judge --post` | `05-post-verdict.json` | production was applied |
|
||||||
|
|
||||||
|
The gate binds to a **manifest digest**: approving a change-set approves *that*
|
||||||
|
change-set. Edit one amount afterwards and stage 4 refuses — the approval no
|
||||||
|
longer matches what is about to be written.
|
||||||
|
|
||||||
|
## The two judges
|
||||||
|
|
||||||
|
Both are context-free: they receive the manifest and the evidence, never the
|
||||||
|
conversation that produced them. Per the PRD
|
||||||
|
[cross-family rule](https://gitea.arcodange.lab/arcodange-org/factory/src/branch/main/vibe/PRD/ai-back-office/qa-strategy.md#independent-verification--no-self-grading),
|
||||||
|
a judge SHOULD be a different model family than whoever built the change-set —
|
||||||
|
the admitted runtimes are Mistral (`vibe -p`) and Ornith 35B (hermes MLX), both
|
||||||
|
proven at verdict parity in [erp#63](https://gitea.arcodange.lab/arcodange-org/erp/issues/63).
|
||||||
|
|
||||||
|
- **Pre-gate** — prompted to *refuse*: find why this change-set is not safe to
|
||||||
|
promote. It reads the rehearsal evidence, not a description of it. Its verdict
|
||||||
|
goes to the human as an opinion, not a veto: a `BLOCK` still lets the operator
|
||||||
|
approve, and the override is recorded in the gate file.
|
||||||
|
- **Post-gate** — prompted to *doubt the success*: compare what production now
|
||||||
|
holds against what the sandbox rehearsal predicted, and report drift. It runs
|
||||||
|
after the writes, so it cannot prevent them — it exists so a silent
|
||||||
|
discrepancy becomes a recorded finding instead of a surprise months later.
|
||||||
|
|
||||||
|
Judges are advisory by design. The blocking authority is the human gate and the
|
||||||
|
host guards; an LLM verdict never silently stops or starts a production write.
|
||||||
|
|
||||||
|
## Usage
|
||||||
|
|
||||||
|
```bash
|
||||||
|
P=fleet/harness/promote/pipeline.py
|
||||||
|
python3 $P rehearse --manifest changeset.json --run-dir runs/2026-07-24-m3-deferred
|
||||||
|
python3 $P judge --run-dir runs/... --stage pre --runtime mistral
|
||||||
|
python3 $P gate --run-dir runs/... # interactive; records who and when
|
||||||
|
python3 $P apply --run-dir runs/... # needs ARCO_PROD_CONFIRM
|
||||||
|
python3 $P judge --run-dir runs/... --stage post --runtime ornith
|
||||||
|
```
|
||||||
|
|
||||||
|
Every stage appends to `journal.jsonl`. The run directory is the evidence pack:
|
||||||
|
it is what you keep, and what an auditor reads.
|
||||||
|
|
||||||
|
## What this does not do
|
||||||
|
|
||||||
|
It does not replace the host guards (`dol-write.sh` refusing non-sandbox hosts,
|
||||||
|
`guard.ts` requiring an explicit production opt-in, the chronology guard in
|
||||||
|
`invoice-create.sh`). Those are structural and stay underneath. This pipeline
|
||||||
|
adds sequence and evidence on top of them.
|
||||||
@@ -0,0 +1,48 @@
|
|||||||
|
{
|
||||||
|
"title": "M3 deferred part — USD 3,000 due 2026-10-23 (issued at D-60)",
|
||||||
|
"observe": [
|
||||||
|
"/invoices?sortfield=t.rowid&sortorder=DESC&limit=3&thirdparty_ids=1"
|
||||||
|
],
|
||||||
|
"ops": [
|
||||||
|
{
|
||||||
|
"label": "invoice: M3 deferred, USD 3000",
|
||||||
|
"api": {
|
||||||
|
"method": "POST",
|
||||||
|
"path": "/invoices",
|
||||||
|
"body": {
|
||||||
|
"socid": 1,
|
||||||
|
"type": 0,
|
||||||
|
"date": 1787529600,
|
||||||
|
"multicurrency_code": "USD",
|
||||||
|
"note_public": "PART DIFFÉRÉE DU CYCLE M3 — contrat cadre du 23/04/2026 (art. 6) et son avenant (art. 2 et 4).\nMontant dû : 3 000,00 USD, réglé en euros au taux de référence EUR/USD du jour du paiement.\nÉmise le 24/08/2026 pour une échéance au 23/10/2026, soit 60 jours — conforme au plafond de l'art. L.441-10 I.\nPÉNALITÉS DE RETARD — En cas de retard de paiement, sont automatiquement dues, sans rappel préalable : (i) des pénalités calculées au taux BCE de refinancement majoré de 10 points (art. L.441-10 II du Code de commerce) — soit 12,15 % l'an au 1er semestre 2026 et 12,40 % l'an au 2e semestre 2026 ; (ii) une indemnité forfaitaire de recouvrement de quarante euros (40 €) par facture impayée (art. L.441-10 III et décret n° 2012-1115) ; (iii) une indemnisation complémentaire sur justification. Aucun escompte pour paiement anticipé.",
|
||||||
|
"lines": [
|
||||||
|
{
|
||||||
|
"desc": "Conseil et accompagnement infrastructure cloud — Cycle M3 (part différée) — période d'exécution du 23/06/2026 au 23/07/2026. Échéance 23/10/2026. Montant contractuel : 3 000,00 USD.\nPrestation de services au sens des articles 259, 1° et 283-2 du CGI. TVA française non applicable — preneur assujetti établi hors de l'Union européenne (États-Unis).",
|
||||||
|
"multicurrency_subprice": "3000",
|
||||||
|
"subprice": "2622.01",
|
||||||
|
"qty": "1",
|
||||||
|
"tva_tx": "0",
|
||||||
|
"product_type": "1"
|
||||||
|
}
|
||||||
|
]
|
||||||
|
}
|
||||||
|
},
|
||||||
|
"then": [
|
||||||
|
{
|
||||||
|
"label": "validate",
|
||||||
|
"method": "POST",
|
||||||
|
"path": "/invoices/{id}/validate",
|
||||||
|
"body": {}
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"label": "set due date 2026-10-23 (D-60)",
|
||||||
|
"method": "PUT",
|
||||||
|
"path": "/invoices/{id}",
|
||||||
|
"body": {
|
||||||
|
"date_lim_reglement": 1792749600
|
||||||
|
}
|
||||||
|
}
|
||||||
|
]
|
||||||
|
}
|
||||||
|
]
|
||||||
|
}
|
||||||
@@ -0,0 +1,40 @@
|
|||||||
|
You are an independent reviewer. A change-set was rehearsed on a sandbox, a human
|
||||||
|
approved it, and it has now been applied to the **production** accounting ledger.
|
||||||
|
The writes already happened — you cannot prevent them. Your job is to make any
|
||||||
|
discrepancy a recorded finding instead of a surprise discovered months later.
|
||||||
|
|
||||||
|
You did not build or approve this change and have no conversation history about
|
||||||
|
it. Judge only the evidence below.
|
||||||
|
|
||||||
|
**Assume production drifted from what the rehearsal predicted, and look for the
|
||||||
|
drift.**
|
||||||
|
|
||||||
|
Compare `sandbox_after` (what the rehearsal produced) with `production_after`
|
||||||
|
(what production now holds), keeping in mind that these are two different systems:
|
||||||
|
internal ids, timestamps and document references legitimately differ. What must
|
||||||
|
match is the **substance**:
|
||||||
|
|
||||||
|
1. **Same objects, same count.** Everything the manifest intended exists in
|
||||||
|
production — and nothing extra appeared.
|
||||||
|
2. **Same amounts.** Totals, currencies, tax treatment.
|
||||||
|
3. **Same dates and states.** Invoice dates, due dates, paid/unpaid, validated
|
||||||
|
or draft.
|
||||||
|
4. **Every op reported ok.** Check `apply_results`; a failed op mid-run may have
|
||||||
|
left the ledger partially written, which matters more than anything else here.
|
||||||
|
5. **Chronology.** Document numbering in production stayed in date order.
|
||||||
|
|
||||||
|
Then answer in this shape, nothing else:
|
||||||
|
|
||||||
|
```
|
||||||
|
VERDICT: PASS (or) VERDICT: BLOCK
|
||||||
|
REASON: one sentence.
|
||||||
|
DRIFT:
|
||||||
|
- one line per substantive difference between rehearsal and production.
|
||||||
|
(write "none" if the substance matches)
|
||||||
|
FOLLOW-UP:
|
||||||
|
- what a human must now do, if anything. (write "none" if nothing)
|
||||||
|
```
|
||||||
|
|
||||||
|
Start your reply with the VERDICT line. `BLOCK` here does not undo anything — it
|
||||||
|
means a human must act. Differences in ids, refs or timestamps are expected and
|
||||||
|
are not drift; do not report them.
|
||||||
@@ -0,0 +1,40 @@
|
|||||||
|
You are an independent reviewer. A change-set has been rehearsed on a sandbox
|
||||||
|
copy of a French company's accounting system (Dolibarr ERP) and is about to be
|
||||||
|
written to the **production ledger**, which is append-only: a wrong entry cannot
|
||||||
|
be deleted, only offset by a credit note.
|
||||||
|
|
||||||
|
You did not build this change-set and you have no conversation history about it.
|
||||||
|
Judge only the evidence below.
|
||||||
|
|
||||||
|
**Your job is to find why this should NOT be promoted.** Assume it is flawed and
|
||||||
|
look for the flaw. Concede only if you cannot find one.
|
||||||
|
|
||||||
|
Check, in this order:
|
||||||
|
|
||||||
|
1. **Did the rehearsal actually work?** `all_writes_succeeded`, and every op's
|
||||||
|
return code and stderr. A failed or half-applied rehearsal is not evidence.
|
||||||
|
2. **Do the observed states support the claim?** Compare `observed_before` and
|
||||||
|
`observed_after`. Did the intended objects appear, with the intended amounts,
|
||||||
|
dates and states? Did anything else change that nobody asked for?
|
||||||
|
3. **Arithmetic and dates.** Totals consistent (HT + VAT = TTC), due dates
|
||||||
|
consistent with the stated payment terms, no date in the future, no invoice
|
||||||
|
dated before one already issued (numbering must stay chronological —
|
||||||
|
French CGI art. 289).
|
||||||
|
4. **Duplication.** Would applying this create a second copy of something that
|
||||||
|
already exists in the observed state?
|
||||||
|
5. **Scope.** Does the change-set do exactly what its title says — no more?
|
||||||
|
|
||||||
|
Then answer in this shape, nothing else:
|
||||||
|
|
||||||
|
```
|
||||||
|
VERDICT: PASS (or) VERDICT: BLOCK
|
||||||
|
REASON: one sentence.
|
||||||
|
FINDINGS:
|
||||||
|
- one line per concrete problem, with the field or object it concerns.
|
||||||
|
(write "none" if you found none)
|
||||||
|
RESIDUAL RISK:
|
||||||
|
- what a human should look at before approving, even if you passed it.
|
||||||
|
```
|
||||||
|
|
||||||
|
Start your reply with the VERDICT line. Be terse. A finding you cannot ground in
|
||||||
|
the evidence below is noise — do not invent one to look thorough.
|
||||||
@@ -0,0 +1,281 @@
|
|||||||
|
#!/usr/bin/env python3
|
||||||
|
"""Gated promote pipeline: rehearse → judge → human gate → apply → judge.
|
||||||
|
|
||||||
|
Stages are chained by artefacts on disk, not by discipline: each one refuses to
|
||||||
|
run until the previous produced its file and that file says what this stage
|
||||||
|
needs. See README.md for why.
|
||||||
|
|
||||||
|
Stdlib only.
|
||||||
|
"""
|
||||||
|
from __future__ import annotations
|
||||||
|
|
||||||
|
import argparse
|
||||||
|
import getpass
|
||||||
|
import hashlib
|
||||||
|
import json
|
||||||
|
import os
|
||||||
|
import subprocess
|
||||||
|
import sys
|
||||||
|
import urllib.error
|
||||||
|
import urllib.request
|
||||||
|
from datetime import datetime, timezone
|
||||||
|
|
||||||
|
REPO = os.path.realpath(os.path.join(os.path.dirname(os.path.abspath(__file__)), "..", "..", ".."))
|
||||||
|
SKILL = os.path.join(REPO, ".claude", "skills")
|
||||||
|
SANDBOX_HOST = "erp-sandbox.arcodange.lab"
|
||||||
|
STAGES = {"rehearsal": "01-rehearsal.json", "pre": "02-pre-verdict.json", "gate": "03-gate.json",
|
||||||
|
"applied": "04-applied.json", "post": "05-post-verdict.json"}
|
||||||
|
|
||||||
|
|
||||||
|
def die(msg: str) -> None:
|
||||||
|
sys.exit(f"pipeline: REFUSED — {msg}")
|
||||||
|
|
||||||
|
|
||||||
|
def now() -> str:
|
||||||
|
return datetime.now(timezone.utc).isoformat(timespec="seconds")
|
||||||
|
|
||||||
|
|
||||||
|
def digest(obj) -> str:
|
||||||
|
return hashlib.sha256(json.dumps(obj, sort_keys=True, ensure_ascii=False).encode()).hexdigest()[:16]
|
||||||
|
|
||||||
|
|
||||||
|
def art(run_dir: str, stage: str) -> str:
|
||||||
|
return os.path.join(run_dir, STAGES[stage])
|
||||||
|
|
||||||
|
|
||||||
|
def read_stage(run_dir: str, stage: str, why: str) -> dict:
|
||||||
|
p = art(run_dir, stage)
|
||||||
|
if not os.path.exists(p):
|
||||||
|
die(f"{why}\n missing: {p}\n run the earlier stage first.")
|
||||||
|
return json.load(open(p))
|
||||||
|
|
||||||
|
|
||||||
|
def write_stage(run_dir: str, stage: str, payload: dict) -> str:
|
||||||
|
p = art(run_dir, stage)
|
||||||
|
json.dump(payload, open(p, "w"), indent=2, ensure_ascii=False)
|
||||||
|
with open(os.path.join(run_dir, "journal.jsonl"), "a") as j:
|
||||||
|
j.write(json.dumps({"at": now(), "stage": stage, "file": os.path.basename(p)},
|
||||||
|
ensure_ascii=False) + "\n")
|
||||||
|
return p
|
||||||
|
|
||||||
|
|
||||||
|
def env_from(path: str) -> dict:
|
||||||
|
cfg = {}
|
||||||
|
for line in open(path):
|
||||||
|
if "=" in line and not line.strip().startswith("#"):
|
||||||
|
k, v = line.strip().split("=", 1)
|
||||||
|
cfg[k] = v.strip().strip('"').strip("'")
|
||||||
|
return cfg
|
||||||
|
|
||||||
|
|
||||||
|
def api(base: str, key: str, method: str, path: str, body=None):
|
||||||
|
url = f"{base.rstrip('/')}/api/index.php{path}"
|
||||||
|
data = json.dumps(body).encode() if body is not None else None
|
||||||
|
req = urllib.request.Request(url, data=data, method=method,
|
||||||
|
headers={"DOLAPIKEY": key, "Content-Type": "application/json",
|
||||||
|
"Accept": "application/json"})
|
||||||
|
try:
|
||||||
|
with urllib.request.urlopen(req, timeout=120) as r:
|
||||||
|
raw = r.read().decode()
|
||||||
|
except urllib.error.HTTPError as e:
|
||||||
|
return {"_error": e.code, "_body": e.read().decode()[:300]}
|
||||||
|
try:
|
||||||
|
return json.loads(raw)
|
||||||
|
except json.JSONDecodeError:
|
||||||
|
return raw.strip().strip('"')
|
||||||
|
|
||||||
|
|
||||||
|
# --- stage 1: rehearse -------------------------------------------------------
|
||||||
|
|
||||||
|
def stage_rehearse(args) -> None:
|
||||||
|
manifest = json.load(open(args.manifest))
|
||||||
|
os.makedirs(args.run_dir, exist_ok=True)
|
||||||
|
cfg = env_from(os.path.join(SKILL, "dolibarr-sandbox-write", ".env"))
|
||||||
|
# the write skill namespaces its keys (DOLIBARR_SANDBOX_*), the read skill does not
|
||||||
|
base = cfg.get("DOLIBARR_SANDBOX_URL") or cfg.get("DOLIBARR_URL", "")
|
||||||
|
skey = cfg.get("DOLIBARR_SANDBOX_API_KEY") or cfg.get("DOLIBARR_API_KEY", "")
|
||||||
|
if SANDBOX_HOST not in base:
|
||||||
|
die(f"the write .env points at {base!r}, not the sandbox — rehearsal must target {SANDBOX_HOST}")
|
||||||
|
|
||||||
|
ops, before = manifest.get("ops", []), {}
|
||||||
|
for probe in manifest.get("observe", []):
|
||||||
|
before[probe] = api(base, skey, "GET", probe)
|
||||||
|
|
||||||
|
results = []
|
||||||
|
for op in ops:
|
||||||
|
call = op.get("api")
|
||||||
|
if not call:
|
||||||
|
die(f"op {op.get('label')!r} has no 'api' block. An op is defined ONCE and replayed on both\n"
|
||||||
|
" targets — two descriptions of the same write are two things that can disagree.")
|
||||||
|
r = api(base, skey, call["method"], call["path"], call.get("body"))
|
||||||
|
failed = isinstance(r, dict) and "_error" in r
|
||||||
|
results.append({"label": op.get("label", ""), "method": call["method"], "path": call["path"],
|
||||||
|
"rc": 1 if failed else 0, "result": r})
|
||||||
|
print(f" [{op.get('label','op')}] {call['method']} {call['path']} -> {'ok' if not failed else r}")
|
||||||
|
for follow in op.get("then", []):
|
||||||
|
fr = api(base, skey, follow["method"], follow["path"].replace("{id}", str(r)), follow.get("body"))
|
||||||
|
ffailed = isinstance(fr, dict) and "_error" in fr
|
||||||
|
results.append({"label": f"{op.get('label','')} :: {follow.get('label','follow-up')}",
|
||||||
|
"method": follow["method"], "path": follow["path"],
|
||||||
|
"rc": 1 if ffailed else 0, "result": fr})
|
||||||
|
print(f" └ {follow.get('label','follow-up')} -> {'ok' if not ffailed else fr}")
|
||||||
|
|
||||||
|
after = {p: api(base, skey, "GET", p) for p in manifest.get("observe", [])}
|
||||||
|
ok = all(r["rc"] == 0 for r in results)
|
||||||
|
payload = {"at": now(), "target": base, "manifest_file": os.path.abspath(args.manifest),
|
||||||
|
"manifest_digest": digest(manifest), "manifest": manifest,
|
||||||
|
"results": results, "observed_before": before, "observed_after": after,
|
||||||
|
"all_writes_succeeded": ok}
|
||||||
|
print(f"→ {write_stage(args.run_dir, 'rehearsal', payload)} (writes ok: {ok})")
|
||||||
|
if not ok:
|
||||||
|
print(" some ops failed — the pre-gate judge will see that, and apply stays blocked.")
|
||||||
|
|
||||||
|
|
||||||
|
# --- stage 2 & 5: judges -----------------------------------------------------
|
||||||
|
|
||||||
|
def call_runtime(runtime: str, prompt: str, model: str | None) -> tuple[str, str]:
|
||||||
|
if runtime == "mistral":
|
||||||
|
r = subprocess.run(["vibe", "-p", prompt, "--max-turns", "1", "--output", "text"],
|
||||||
|
capture_output=True, text=True, timeout=600)
|
||||||
|
if r.returncode != 0:
|
||||||
|
die(f"vibe failed: {r.stderr[-300:]}")
|
||||||
|
return r.stdout.strip(), "vibe -p (mistral)"
|
||||||
|
if runtime in ("ornith", "mlx"):
|
||||||
|
model = model or "leonsarmiento/Ornith-1.0-35B-5bit-mlx"
|
||||||
|
body = json.dumps({"model": model, "messages": [{"role": "user", "content": prompt}],
|
||||||
|
"temperature": 0, "max_tokens": 3000}).encode()
|
||||||
|
req = urllib.request.Request("http://127.0.0.1:18080/v1/chat/completions", data=body,
|
||||||
|
headers={"Content-Type": "application/json"})
|
||||||
|
with urllib.request.urlopen(req, timeout=900) as r:
|
||||||
|
m = json.load(r)["choices"][0]["message"]
|
||||||
|
return (m.get("content") or m.get("reasoning") or ""), model
|
||||||
|
die(f"unknown runtime {runtime!r} (admitted: mistral, ornith, mlx)")
|
||||||
|
raise AssertionError
|
||||||
|
|
||||||
|
|
||||||
|
def stage_judge(args) -> None:
|
||||||
|
here = os.path.dirname(os.path.abspath(__file__))
|
||||||
|
if args.stage == "pre":
|
||||||
|
reh = read_stage(args.run_dir, "rehearsal", "a pre-gate judge needs a rehearsal to judge")
|
||||||
|
tmpl = open(os.path.join(here, "judges", "pre-gate.md")).read()
|
||||||
|
evidence = {k: reh[k] for k in ("manifest", "results", "observed_before", "observed_after",
|
||||||
|
"all_writes_succeeded")}
|
||||||
|
else:
|
||||||
|
applied = read_stage(args.run_dir, "applied", "a post-gate judge needs a production apply to verify")
|
||||||
|
reh = read_stage(args.run_dir, "rehearsal", "the post-gate judge compares production to the rehearsal")
|
||||||
|
tmpl = open(os.path.join(here, "judges", "post-gate.md")).read()
|
||||||
|
evidence = {"manifest": reh["manifest"], "sandbox_after": reh["observed_after"],
|
||||||
|
"production_after": applied.get("observed_after"), "apply_results": applied.get("results")}
|
||||||
|
|
||||||
|
prompt = f"{tmpl}\n\n--- EVIDENCE (JSON) ---\n{json.dumps(evidence, indent=2, ensure_ascii=False)}\n--- END ---"
|
||||||
|
text, model = call_runtime(args.runtime, prompt, args.model)
|
||||||
|
verdict = "BLOCK" if "BLOCK" in text.upper()[:400] else ("PASS" if "PASS" in text.upper()[:400] else "UNCLEAR")
|
||||||
|
payload = {"at": now(), "stage": args.stage, "runtime": args.runtime, "model": model,
|
||||||
|
"prompt_sha256": hashlib.sha256(prompt.encode()).hexdigest(), "verdict": verdict,
|
||||||
|
"response": text}
|
||||||
|
print(f"→ {write_stage(args.run_dir, args.stage, payload)} verdict={verdict} ({model})")
|
||||||
|
print(text[:1200])
|
||||||
|
|
||||||
|
|
||||||
|
# --- stage 3: the human gate -------------------------------------------------
|
||||||
|
|
||||||
|
def stage_gate(args) -> None:
|
||||||
|
reh = read_stage(args.run_dir, "rehearsal", "nothing to approve — rehearse first")
|
||||||
|
pre = read_stage(args.run_dir, "pre", "the human decides WITH a judge's opinion, not without one")
|
||||||
|
m = reh["manifest"]
|
||||||
|
print("=" * 72)
|
||||||
|
print("HUMAN GATE — this approves writing to the PRODUCTION ledger.")
|
||||||
|
print("=" * 72)
|
||||||
|
print(f"change-set : {m.get('title', '(untitled)')}")
|
||||||
|
print(f"digest : {reh['manifest_digest']} (approval binds to this exact change-set)")
|
||||||
|
print(f"ops : {len(m.get('ops', []))}")
|
||||||
|
for op in m.get("ops", []):
|
||||||
|
print(f" - {op.get('label') or op['script']}")
|
||||||
|
print(f"rehearsal : writes ok = {reh['all_writes_succeeded']} on {reh['target']}")
|
||||||
|
print(f"judge ({pre['runtime']}) : {pre['verdict']}")
|
||||||
|
for line in pre["response"].strip().splitlines()[:8]:
|
||||||
|
print(f" | {line[:100]}")
|
||||||
|
print("=" * 72)
|
||||||
|
if not reh["all_writes_succeeded"]:
|
||||||
|
print("NOTE: the rehearsal had failing ops. Approving anyway is your call, and is recorded.")
|
||||||
|
if pre["verdict"] == "BLOCK":
|
||||||
|
print("NOTE: the judge says BLOCK. You may still approve — the override is recorded.")
|
||||||
|
|
||||||
|
if args.decision:
|
||||||
|
decision, who = args.decision, args.by or getpass.getuser()
|
||||||
|
else:
|
||||||
|
decision = input("\ntype 'approve' to promote, anything else to abort: ").strip().lower()
|
||||||
|
who = input("your name (recorded in the evidence pack): ").strip() or getpass.getuser()
|
||||||
|
|
||||||
|
payload = {"at": now(), "decision": "approved" if decision == "approve" else "rejected",
|
||||||
|
"by": who, "manifest_digest": reh["manifest_digest"],
|
||||||
|
"judge_verdict": pre["verdict"],
|
||||||
|
"override_of_judge": pre["verdict"] == "BLOCK" and decision == "approve",
|
||||||
|
"override_of_failed_rehearsal": (not reh["all_writes_succeeded"]) and decision == "approve"}
|
||||||
|
print(f"→ {write_stage(args.run_dir, 'gate', payload)} decision={payload['decision']} by {who}")
|
||||||
|
|
||||||
|
|
||||||
|
# --- stage 4: apply to production -------------------------------------------
|
||||||
|
|
||||||
|
def stage_apply(args) -> None:
|
||||||
|
reh = read_stage(args.run_dir, "rehearsal", "nothing rehearsed")
|
||||||
|
gate = read_stage(args.run_dir, "gate", "production requires a human decision on record")
|
||||||
|
if gate["decision"] != "approved":
|
||||||
|
die(f"the gate recorded '{gate['decision']}' by {gate['by']} — not approved.")
|
||||||
|
if gate["manifest_digest"] != reh["manifest_digest"]:
|
||||||
|
die("the approved change-set is not the one about to be applied "
|
||||||
|
f"(approved {gate['manifest_digest']}, current {reh['manifest_digest']}).\n"
|
||||||
|
" Re-run the gate on the current change-set.")
|
||||||
|
if os.environ.get("ARCO_PROD_CONFIRM") != "I-UNDERSTAND-THIS-WRITES-PROD":
|
||||||
|
die("set ARCO_PROD_CONFIRM=I-UNDERSTAND-THIS-WRITES-PROD to apply to production")
|
||||||
|
|
||||||
|
cfg = env_from(os.path.join(SKILL, "dolibarr", ".env"))
|
||||||
|
base, key = cfg["DOLIBARR_URL"], cfg["DOLIBARR_API_KEY"]
|
||||||
|
if SANDBOX_HOST in base:
|
||||||
|
die(f"the prod .env points at the sandbox ({base}) — nothing to promote to")
|
||||||
|
print(f"*** PRODUCTION: {base} — approved by {gate['by']} at {gate['at']} ***")
|
||||||
|
|
||||||
|
m = reh["manifest"]
|
||||||
|
results = []
|
||||||
|
for op in m.get("ops", []):
|
||||||
|
call = op["api"]
|
||||||
|
r = api(base, key, call["method"], call["path"], call.get("body"))
|
||||||
|
failed = isinstance(r, dict) and "_error" in r
|
||||||
|
results.append({"label": op.get("label"), "method": call["method"], "path": call["path"],
|
||||||
|
"result": r, "ok": not failed})
|
||||||
|
print(f" [{op.get('label')}] {call['method']} {call['path']} -> {'ok' if not failed else r}")
|
||||||
|
if failed and not args.keep_going:
|
||||||
|
break
|
||||||
|
for follow in op.get("then", []):
|
||||||
|
fr = api(base, key, follow["method"], follow["path"].replace("{id}", str(r)), follow.get("body"))
|
||||||
|
ffailed = isinstance(fr, dict) and "_error" in fr
|
||||||
|
results.append({"label": f"{op.get('label')} :: {follow.get('label','follow-up')}",
|
||||||
|
"method": follow["method"], "path": follow["path"],
|
||||||
|
"result": fr, "ok": not ffailed})
|
||||||
|
print(f" └ {follow.get('label','follow-up')} -> {'ok' if not ffailed else fr}")
|
||||||
|
if ffailed and not args.keep_going:
|
||||||
|
break
|
||||||
|
after = {p: api(base, key, "GET", p) for p in m.get("observe", [])}
|
||||||
|
payload = {"at": now(), "target": base, "approved_by": gate["by"],
|
||||||
|
"manifest_digest": reh["manifest_digest"], "results": results, "observed_after": after,
|
||||||
|
"all_ok": all(r["ok"] for r in results)}
|
||||||
|
print(f"→ {write_stage(args.run_dir, 'applied', payload)} (all ok: {payload['all_ok']})")
|
||||||
|
|
||||||
|
|
||||||
|
def main() -> int:
|
||||||
|
ap = argparse.ArgumentParser(description="gated promote pipeline")
|
||||||
|
sub = ap.add_subparsers(dest="cmd", required=True)
|
||||||
|
r = sub.add_parser("rehearse"); r.add_argument("--manifest", required=True); r.add_argument("--run-dir", required=True)
|
||||||
|
j = sub.add_parser("judge"); j.add_argument("--run-dir", required=True)
|
||||||
|
j.add_argument("--stage", choices=["pre", "post"], required=True)
|
||||||
|
j.add_argument("--runtime", default="mistral"); j.add_argument("--model")
|
||||||
|
g = sub.add_parser("gate"); g.add_argument("--run-dir", required=True)
|
||||||
|
g.add_argument("--decision", choices=["approve", "reject"]); g.add_argument("--by")
|
||||||
|
a = sub.add_parser("apply"); a.add_argument("--run-dir", required=True); a.add_argument("--keep-going", action="store_true")
|
||||||
|
args = ap.parse_args()
|
||||||
|
{"rehearse": stage_rehearse, "judge": stage_judge, "gate": stage_gate, "apply": stage_apply}[args.cmd](args)
|
||||||
|
return 0
|
||||||
|
|
||||||
|
|
||||||
|
if __name__ == "__main__":
|
||||||
|
sys.exit(main())
|
||||||
Reference in New Issue
Block a user