Table of Contents generated with DocToc
Test whether changes to AGENTS.md break or improve agent decisions. The harness compares the main branch version against your working tree, running the same cases against both and reporting the diff.
Each arm is a git worktree of the real repo — the agent sees actual source files (pyproject.toml, directory structure, etc.). The only difference between arms is which AGENTS.md is present.
The agent reads guidance through the CLAUDE.md → AGENTS.md symlink, so the harness verifies that symlink exists (in the working tree and on the base branch) before running and aborts if it is broken — a regular-file CLAUDE.md would make every arm read identical guidance.
claude /login) — Pro/Max subscriptionANTHROPIC_API_KEY environment variable — API creditsThat's it — prek provisions Node, promptfoo, and the Claude Agent SDK automatically.
# Run the eval (single pass over all cases — satisfies the hash gate). # Stage your changes first — prek stashes unstaged edits: prek run run-skill-eval --hook-stage manual --all-files # Repeat each case to reduce nondeterminism: EVAL_REPEAT=3 prek run run-skill-eval --hook-stage manual --all-files # Add a baseline arm (no AGENTS.md) to measure raw model capability: EVAL_FULL=1 prek run run-skill-eval --hook-stage manual --all-files # Use a cheaper model for fast iteration: MODEL=claude-haiku-4-5-20251001 prek run run-skill-eval --hook-stage manual --all-files # Test a skill alongside AGENTS.md (not combinable with EVAL_FULL — the # baseline arm has no skill, so the skill-used assertion would always fail): SKILL_NAME=airflow-contribution prek run run-skill-eval --hook-stage manual --all-files # View results in browser (Ctrl-C to stop the server): prek run view-skill-eval --hook-stage manual --all-files
Other promptfoo flags (--filter*, --no-cache) are argv-only — prek run can't forward arguments, so wire them as fixed entry args on a hook variant if a preset is needed.
A run with --filter* flags covers only a subset of cases, so it does not update the hash file.
Each run also writes a JSON report to files/skill-evals/results.json (per the files/ output convention) — handy for pasting results into a PR.
Everything the harness stores lives inside the repo — nothing is left in your home directory:
rm -rf .build/promptfoo # eval history and cache prek clean # prek-provisioned node envs
Cases live in cases/*.yaml. Add entries to an existing file or create a new one — no config changes needed.
- description: "Scheduler bugfix (#64322): no newsfragment" vars: request: | I fixed the scheduler to skip asset-triggered Dags that don't have a SerializedDagModel yet. Should I create a newsfragment? assert: - type: javascript value: 'output.should_create === false'
The agent returns structured JSON ({should_create, type, rationale}). Use output.should_create directly in assertions.
main's AGENTS.md, one with your working tree version. Both are full repo checkouts.anthropic:claude-agent-sdk provider and output_format: json_schema for structured output.After every completed run, eval.py records a hash of its inputs (AGENTS.md + cases/*.yaml) in last-eval-hash.txt. The check-eval-hash prek hook — enforced locally and in CI — recomputes the hash and fails when guidance changed without a re-run. Commit the updated hash file together with your change.
main — that is signal, not a defect).SKIP=check-eval-hash git commit ... — CI stays red until the eval is re-run.SKILL.md files are not covered by the gate yet — there are no skill cases, so a hash over them would prove nothing. Extend compute_guidance_hash in scripts/ci/prek/check_eval_hash.py when per-skill cases land.