L'extraction de champs vivait dans un heredoc à l'intérieur d'email-inspect.sh : impossible à exécuter isolément, donc jamais mesurée, donc fausse sans que personne puisse le voir. Sur la facture Darnis F1048 elle renvoyait le numéro de TVA d'Arcodange comme référence de facture, et aucune date. - extract_fields.py : l'extraction sort du shell et devient un module. - test_extract.py : régression contre les 16 factures hand-vérifiées de fleet/golden/invoice-extract/. Score par champ, et une valeur FAUSSE pèse plus qu'une valeur absente — un humain recopie ce qui s'affiche. Valeurs fausses : 4 → 0. Exactitude ref 62,5 → 75 %, date 62,5 → 75 %, HT 68,8 → 75 %, TTC 81,2 → 93,8 %. Cinq bugs réels, dont trois invisibles sans test : - « Nº » sur les factures françaises est U+00BA (ordinal masculin), pas le signe degré. La classe [°o] le rate, le motif principal échoue, et le repli attrape le premier jeton ref-shaped du document — très souvent un numéro de TVA. - Le filtre anti-TVA rejetait « FR73261832 », qui est la vraie référence OVH : un numéro FR fait exactement 11 caractères après le préfixe. - « Montant total (HT) » était lu comme un TTC. - Une référence coupée par la colonne (« 06-01-26- » / « payment-366753 ») était renvoyée amputée : le recollage doit précéder le scan, sinon la queue seule est trouvée en premier. - Un `\b` après `€` ne peut jamais matcher en fin de ligne (€ n'est pas un caractère de mot) — la TVA n'était jamais extraite. adc-008 : une facture fournisseur s'enregistre à SA date, même future, tant que l'exercice (année civile) ne bascule pas. Le document fait foi ; altérer sa date ferait diverger l'écriture de sa pièce justificative (CGI art. 289 VII). Registre validé : 8 règles, 8 ADC, 0 erreur. scopes.ts : 1232 (factures fournisseur) ajouté à prod-write — oubli initial, révélé par un 403 en production sur F1048. Le pipeline s'est arrêté sans écrire. Appliqué en production via le pipeline gated : FAF2026014 (Darnis F1048), 218,50 HT + 43,70 TVA = 262,20 TTC, validée, non réglée. Co-Authored-By: Claude Opus 5 (1M context) <[email protected]> Claude-Session: https://claude.ai/code/session_01VRShc4QhLLU73FLHx9vskh
233 lines
9.0 KiB
Bash
Executable File
233 lines
9.0 KiB
Bash
Executable File
#!/usr/bin/env bash
|
|
# Inspect one email by id and propose a Dolibarr supplier-invoice draft.
|
|
#
|
|
# Usage:
|
|
# email-inspect.sh <messageId> [--folder PATH] # default folder: /Inbox/books
|
|
# [--save-pdf DIR] # save PDF attachments under DIR/
|
|
# [--json] # emit a single JSON object on stdout
|
|
#
|
|
# Pipeline (read-only):
|
|
# 1. Find the message (in the given folder, default /Inbox/books).
|
|
# 2. List attachments via /attachmentinfo.
|
|
# 3. For each PDF attachment: download, run pdftotext, extract supplier-side
|
|
# heuristics (name, totals, dates, ref).
|
|
# 4. Emit a draft "Dolibarr-ready" record per attachment so the operator can
|
|
# hand-create the supplier invoice in the Dolibarr UI.
|
|
#
|
|
# This skill DOES NOT write to Dolibarr. Auto-creation of supplier invoices is
|
|
# V9 candidate.
|
|
|
|
set -euo pipefail
|
|
|
|
SCRIPT_DIR="$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)"
|
|
ZOHO_CURL="${SCRIPT_DIR}/zoho-curl.sh"
|
|
|
|
if [[ $# -lt 1 ]]; then
|
|
echo "email-inspect.sh: missing <messageId>" >&2
|
|
echo " Hint: bin/arcodange email list to see candidate ids." >&2
|
|
exit 2
|
|
fi
|
|
MID="$1"; shift || true
|
|
FOLDER="/Inbox/books"; SAVE_PDF_DIR=""; FMT="text"
|
|
while [[ $# -gt 0 ]]; do
|
|
case "$1" in
|
|
--folder) FOLDER="$2"; shift 2 ;;
|
|
--save-pdf) SAVE_PDF_DIR="$2"; shift 2 ;;
|
|
--json) FMT="json"; shift ;;
|
|
-h|--help) sed -n '2,18p' "$0" | sed 's/^# \{0,1\}//'; exit 0 ;;
|
|
*) echo "email-inspect.sh: unknown arg: $1" >&2; exit 2 ;;
|
|
esac
|
|
done
|
|
|
|
command -v pdftotext >/dev/null || { echo "email-inspect.sh: pdftotext not found (brew install poppler)" >&2; exit 2; }
|
|
|
|
WORK="$(mktemp -d -t emailinspect.XXXXXX)"
|
|
trap 'rm -rf "${WORK}"' EXIT
|
|
|
|
# 1. accountId + folderId
|
|
"${ZOHO_CURL}" /accounts > "${WORK}/accounts.json"
|
|
AID=$(python3 -c "import json,sys; d=json.load(open(sys.argv[1])); print((d.get('data') or [{}])[0].get('accountId',''))" "${WORK}/accounts.json")
|
|
"${ZOHO_CURL}" "/accounts/${AID}/folders" > "${WORK}/folders.json"
|
|
FID=$(python3 -c "
|
|
import json, sys
|
|
d = json.load(open(sys.argv[1]))
|
|
target = sys.argv[2]
|
|
for f in (d.get('data') or []):
|
|
if f.get('path') == target:
|
|
print(f.get('folderId')); break" "${WORK}/folders.json" "${FOLDER}")
|
|
[[ -z "${FID}" ]] && { echo "email-inspect.sh: folder '${FOLDER}' not found" >&2; exit 2; }
|
|
|
|
# 2. Find the message in the folder listing (to grab metadata: subject, from, date)
|
|
"${ZOHO_CURL}" "/accounts/${AID}/messages/view?folderId=${FID}&limit=100&sortorder=false&start=1" > "${WORK}/folder_msgs.json"
|
|
python3 - "${WORK}/folder_msgs.json" "${MID}" > "${WORK}/meta.json" <<'PY'
|
|
import json, sys
|
|
d = json.load(open(sys.argv[1]))
|
|
mid = sys.argv[2]
|
|
for m in (d.get("data") or []):
|
|
if str(m.get("messageId")) == mid:
|
|
json.dump(m, sys.stdout); sys.exit(0)
|
|
sys.exit(f"messageId {mid} not found in this folder")
|
|
PY
|
|
|
|
# 3. Attachment metadata
|
|
"${ZOHO_CURL}" "/accounts/${AID}/folders/${FID}/messages/${MID}/attachmentinfo" > "${WORK}/attachinfo.json"
|
|
|
|
# 4. Download each attachment — needs raw bytes (Accept: */*), not the JSON
|
|
# wrapper's default. We bypass zoho-curl.sh for the attachment download but
|
|
# reuse the cached access_token it wrote.
|
|
set -a; source "${SCRIPT_DIR}/../../dolibarr/.env"; set +a
|
|
: "${ZOHO_DC:=eu}"
|
|
TOKEN_CACHE="${TMPDIR:-/tmp}/zoho-access-$(whoami)"
|
|
if [[ ! -s "${TOKEN_CACHE}" ]]; then
|
|
echo "email-inspect.sh: missing access token cache — run any zoho-curl call first to populate it" >&2
|
|
exit 2
|
|
fi
|
|
ACCESS_TOKEN=$(cat "${TOKEN_CACHE}")
|
|
MAIL_BASE="https://mail.zoho.${ZOHO_DC}/api"
|
|
|
|
mkdir -p "${WORK}/atts" "${WORK}/text"
|
|
ATT_IDS=$(python3 -c "
|
|
import json, sys
|
|
d = json.load(open(sys.argv[1]))
|
|
data = d.get('data') or {}
|
|
for a in (data.get('attachments') or []):
|
|
print(f\"{a.get('attachmentId')}|{a.get('attachmentName','-')}\")" "${WORK}/attachinfo.json")
|
|
while IFS='|' read -r aid aname; do
|
|
[[ -z "${aid}" ]] && continue
|
|
outpath="${WORK}/atts/${aname}"
|
|
curl -sS \
|
|
-H "Authorization: Zoho-oauthtoken ${ACCESS_TOKEN}" \
|
|
-H "Accept: */*" \
|
|
--max-time 60 \
|
|
-o "${outpath}" \
|
|
"${MAIL_BASE}/accounts/${AID}/folders/${FID}/messages/${MID}/attachments/${aid}" || true
|
|
# If pdf, extract text (bash 3.2 compatible — no ${var,,})
|
|
aname_lc=$(echo "${aname}" | tr '[:upper:]' '[:lower:]')
|
|
if [[ "${aname_lc}" == *.pdf ]]; then
|
|
pdftotext -layout "${outpath}" "${WORK}/text/${aname%.pdf}.txt" 2>/dev/null || true
|
|
fi
|
|
done <<< "${ATT_IDS}"
|
|
|
|
# Optional save
|
|
if [[ -n "${SAVE_PDF_DIR}" ]]; then
|
|
mkdir -p "${SAVE_PDF_DIR}"
|
|
cp "${WORK}/atts/"*.pdf "${SAVE_PDF_DIR}/" 2>/dev/null || true
|
|
fi
|
|
|
|
# 5. Heuristic extract + render
|
|
EXTRACT_DIR="${SCRIPT_DIR}" python3 - "${WORK}" "${FMT}" <<'PY'
|
|
import json, sys, os, re, datetime, glob
|
|
work, fmt = sys.argv[1:3]
|
|
|
|
meta = json.load(open(os.path.join(work,"meta.json")))
|
|
ts = int(meta.get("sentDateInGMT") or meta.get("receivedTime") or 0) // 1000
|
|
mail_date = datetime.datetime.fromtimestamp(ts).strftime("%Y-%m-%d") if ts else None
|
|
mail_from = (meta.get("fromAddress") or meta.get("sender") or "-").replace("<","<").replace(">",">").replace("<","").replace(">","")
|
|
mail_subject = meta.get("subject") or "-"
|
|
|
|
# Heuristics on PDF text
|
|
def extract(text):
|
|
out = {}
|
|
# First non-empty line is often the supplier name (or the address block first line)
|
|
lines = [l.strip() for l in text.splitlines() if l.strip()]
|
|
out["pdf_top_line"] = lines[0] if lines else None
|
|
|
|
# Total TTC / HT / TVA — try multiple French/English patterns
|
|
def first_match(*patterns):
|
|
for p in patterns:
|
|
for line in lines:
|
|
m = re.search(p, line, re.IGNORECASE)
|
|
if m: return m.group(1).replace(",", ".").replace(" ", "")
|
|
return None
|
|
|
|
def parse_amount(s):
|
|
if not s: return None
|
|
clean = s.replace(",", ".").replace(" ", "")
|
|
try:
|
|
v = float(clean)
|
|
# Money amounts < 1M EUR; filters out VAT-number false positives (FR12345678901)
|
|
return v if 0 <= v < 1_000_000 else None
|
|
except: return None
|
|
|
|
def first_amount(*patterns):
|
|
for p in patterns:
|
|
for line in lines:
|
|
m = re.search(p, line, re.IGNORECASE)
|
|
if m:
|
|
v = parse_amount(m.group(1))
|
|
if v is not None: return f"{v:.2f}"
|
|
return None
|
|
|
|
# Extraction déléguée au module testé (extract_fields.py), pinné par
|
|
# test_extract.py contre les 16 factures du golden set erp#39. L'ancienne
|
|
# version inline n'était pas exécutable isolément, donc jamais mesurée :
|
|
# elle renvoyait le n° de TVA d'Arcodange comme référence de facture.
|
|
# `python3 -` reads from stdin, so __file__ does not exist here: the shell
|
|
# passes the module's directory in EXTRACT_DIR.
|
|
sys.path.insert(0, os.environ["EXTRACT_DIR"])
|
|
from extract_fields import extract as _extract
|
|
out.update(_extract(text))
|
|
return out
|
|
|
|
|
|
pdfs = []
|
|
for pdf in sorted(glob.glob(os.path.join(work,"atts","*.pdf")) +
|
|
glob.glob(os.path.join(work,"atts","*.PDF"))):
|
|
name = os.path.basename(pdf)
|
|
txt_path = os.path.join(work,"text", os.path.splitext(name)[0] + ".txt")
|
|
text = open(txt_path).read() if os.path.isfile(txt_path) else ""
|
|
h = extract(text)
|
|
h["attachment_name"] = name
|
|
h["pdf_size_bytes"] = os.path.getsize(pdf)
|
|
h["pdf_text_len"] = len(text)
|
|
pdfs.append(h)
|
|
|
|
result = {
|
|
"email": {
|
|
"messageId": meta.get("messageId"),
|
|
"subject": mail_subject,
|
|
"from": mail_from,
|
|
"date": mail_date,
|
|
"hasAttachment": str(meta.get("hasAttachment","")) == "1",
|
|
},
|
|
"attachments": pdfs,
|
|
"dolibarr_draft_suggestions": [
|
|
{
|
|
"supplier_hint": p.get("pdf_top_line"),
|
|
"invoice_ref": p.get("invoice_ref"),
|
|
"invoice_date": p.get("invoice_date_raw"),
|
|
"total_ht": p.get("total_ht"),
|
|
"total_tva": p.get("total_tva"),
|
|
"total_ttc": p.get("total_ttc"),
|
|
"vat_rate_pct": p.get("vat_rate_pct"),
|
|
"source_email": meta.get("messageId"),
|
|
"source_attachment": p.get("attachment_name"),
|
|
} for p in pdfs
|
|
]
|
|
}
|
|
|
|
if fmt == "json":
|
|
print(json.dumps(result, indent=2, ensure_ascii=False))
|
|
sys.exit(0)
|
|
|
|
print("=" * 80)
|
|
print(f" Email {meta.get('messageId')}")
|
|
print("=" * 80)
|
|
print(f" subject : {mail_subject}")
|
|
print(f" from : {mail_from}")
|
|
print(f" date : {mail_date}")
|
|
print(f" attached : {result['email']['hasAttachment']}")
|
|
print()
|
|
if not pdfs:
|
|
print(" (no PDF attachments — try inspecting body or other types)")
|
|
for i, p in enumerate(pdfs, 1):
|
|
print(f" -- Attachment {i}: {p['attachment_name']} ({p['pdf_size_bytes']} bytes, {p['pdf_text_len']} chars extracted) --")
|
|
for k in ("pdf_top_line","invoice_ref","invoice_date_raw","total_ht","total_tva","total_ttc","vat_rate_pct"):
|
|
v = p.get(k)
|
|
print(f" {k:<16} = {v!r}")
|
|
print()
|
|
|
|
print(" Suggested Dolibarr supplier-invoice draft entries:")
|
|
print(json.dumps(result["dolibarr_draft_suggestions"], indent=4, ensure_ascii=False))
|
|
PY
|