- runs/2026-07-18/: the 8 sha256-pinned verifier transcripts (4 runtimes × 2 tests), blind-judging verdicts (2 independent judges per cell, unanimous), the erp#56 builder-bench journal + prompt + caps, and the evidence README with the parity table. - run-verifier.sh: mistral runtime drops the tool-filter flag (--enabled-tools with a no-match pattern hangs vibe 2.21.0); plain -p with --max-turns 1. Verdicts: Mistral (vibe -p, mistral-medium-3.5) and Ornith 35B (hermes MLX) reach verdict parity with the Claude baseline on both tests → admitted to verifier duty. Qwen2.5-7B-4bit fails both → the honest small-model floor. Builder bench: erp#56 completed by the Mistral runtime, 0 code corrections, 261 s, acceptance run clean (0 bank-UNKNOWN) → merged as PR #68. Closes #63 (with the paired factory qa-strategy PR). Co-Authored-By: Claude Fable 5 <[email protected]> Claude-Session: https://claude.ai/code/session_01VRShc4QhLLU73FLHx9vskh
17 lines
711 B
JSON
17 lines
711 B
JSON
{
|
|
"test": "locate",
|
|
"runtime": "mistral",
|
|
"model": "vibe-active-model",
|
|
"endpoint": "vibe -p",
|
|
"timestamp": "20260718T194754",
|
|
"latency_s": 14,
|
|
"prompt_sha256": "b22405e7d0db8915c4aae6eddbc2a003f600be41f605572356fe79a7c1f51c00",
|
|
"inputs": {
|
|
"AGENTS.md": {
|
|
"path": "/Users/gabrielradureau/Work/Arcodange/erp/.claude/worktrees/harness-portability/AGENTS.md",
|
|
"sha256": "a77d356e7804d2792b605bef2e9daba0233d93ef791675acda021d2dac9b02a6"
|
|
}
|
|
},
|
|
"response": "- **Atom registry**: `fleet/README.md`\n- **Class skeletons**: `fleet/classes/`\n- **Environment rules**: the [operating rules](#operating-rules-for-agents) section + `.claude/skills/dolibarr-sandbox-write/SKILL.md`"
|
|
}
|