Authoring evals
You can create a suite in four ways: run arc-skill-eval create, run arc-skill-eval create --guided, use the bundled arc-creating-evals skill, or write evals/evals.json yourself. Each method creates an evals/ directory next to SKILL.md, with evals.json and any fixture files.
Choose deterministic or guided create
Section titled “Choose deterministic or guided create”Start with deterministic create when the skill has obvious mechanical outcomes:
arc-skill-eval create ./skills/my-skill --dry-run --summaryarc-skill-eval create ./skills/my-skillThis mode does not call a model. It reads SKILL.md, creates trigger, execution, and adjacent-negative cases, infers fixture inputs, and adds deterministic assertions when the skill names files, JSON, or text patterns.
Use guided create when the eval design itself needs judgment:
arc-skill-eval create ./skills/my-skill --guided --interactiveUse guided mode for conceptual, planning, review, and routing skills. Review its proposal before writing files: remove weak cases, edit vague assertions, and keep deterministic checks for concrete outputs.
Conceptual example: grill-me
Section titled “Conceptual example: grill-me”A grill-me skill interviews a user about a plan. It may not produce a report.json or plan.md file. Its suite can check:
- explicit trigger cases such as “grill me on this launch plan”
- adjacent negatives such as “summarize this launch plan” or “rewrite this plan more clearly”
- judge assertions that the assistant asks pointed follow-up questions, challenges assumptions, explores tradeoffs, and avoids solving the plan for the user too early
Do not invent file assertions for files the skill does not create. Use deterministic assertions for concrete outputs and judge assertions for tone, interview behavior, and other prose properties.
Prefer behavior-focused assertions
Section titled “Prefer behavior-focused assertions”Check the skill’s stated behavior: created files, JSON shape, avoided commands, tool use, routing decisions, and response content. A correct response should not fail because it uses different wording.
Brittle examples:
"The response says exactly: Phase 1 — detection complete.""The response uses the heading ## Recommended Fix.""The response mentions the arc-conventional-commits skill name."Checks for observable results:
{ "type": "file-exists", "path": ".releaserc.json" }{ "type": "regex-match", "pattern": "conventionalcommits", "target": { "file": ".releaserc.json" } }"The response names semantic-release and describes configuring release automation for this repository rather than giving generic advice."Exact wording is still valid when the words are the feature. Commit messages, CLI output, public copy, email subjects, and mandatory safety text may all need literal checks. In those cases, say so in expected_output and prefer deterministic assertions:
{ "expected_output": "The assistant returns exactly one conventional commit message beginning with fix:", "assertions": [ { "type": "regex-match", "pattern": "^fix(\\(.+\\))?:", "flags": "m", "target": "assistant-text" } ]}Use arc-creating-evals
Section titled “Use arc-creating-evals”arc-creating-evals ships in this repository. Install it in your tool’s skills directory, such as .claude/skills/ for Claude Code. See skills/README.md. Then use prompts such as:
- “Write evals for this skill.”
- “Create evals.json for
arc-conventional-commits.” - “Make this skill testable.”
It walks five phases and stops at each boundary for confirmation:
- Locate and summarize the target skill. The skill reads
SKILL.mdand extracts its name, description, phases, expected tools, and file changes. It asks, “Did I capture the skill’s behavior correctly?” - Define success across outcome, process, style, and efficiency.
- Outcome: what file state, response content, or artifact shows that the skill worked?
- Process: which tools must or must not run?
- Style: does tone, structure, or formatting matter?
- Efficiency: is there a limit on tool calls or duration?
- Draft six to ten cases in three groups:
- Three to five trigger cases, including explicit and implicit requests plus an adjacent negative.
- One to three execution cases, including a fixture-backed main path and materially different alternate paths.
- One or two edge or negative cases for malformed input, ambiguity, or conflicting state.
- Add two to five assertions per case. Start with scripts, then add judged strings for prose that scripts cannot check.
- Write and validate the files, run one case, and summarize the suite.
The skill pauses after the success criteria and case list. It does not write files until you approve the cases.
Write evals/evals.json by hand
Section titled “Write evals/evals.json by hand”The shape (also documented under Concepts → Eval cases and Concepts → Assertions):
{ "skill_name": "arc-conventional-commits", "evals": [ { "id": 1, "prompt": "Set up semantic-release in this repo.", "expected_output": "semantic-release configured with the Conventional Commits preset.", "files": ["files/clean-repo"], "assertions": [ { "type": "file-exists", "path": ".releaserc.json" }, { "type": "regex-match", "pattern": "conventionalcommits", "target": { "file": ".releaserc.json" } }, "The response summarizes the semantic-release plugins it installed." ] } ]}Place the file at <your-skill>/evals/evals.json. Fixtures, if any, go under <your-skill>/evals/files/<fixture-name>/.
Quality rules
Section titled “Quality rules”Follow these rules for every authoring method:
- Do not add cases for behavior the skill does not support.
- Test string assertions against a real run before keeping them.
- Do not copy
SKILL.mdinstructions into assertions. Check the effect, such as valid JSON containingconventionalcommits. - Prefer one execution case with strong script assertions to several cases with weak text checks.
- Use lowercase kebab-case IDs or numbers.
- Ask the user when expected behavior is unclear.
Pinning a model and dry-running
Section titled “Pinning a model and dry-running”evals.json doesn’t have a top-level model field; the skill’s assistant uses whatever your Pi default is. If your global default is quota-capped (for example, openai-codex on a ChatGPT Plus account), set a judgeModel override at the CLI layer or update ~/.pi/agent/settings.json.
Dry-run one case against your cheapest available model before committing the suite:
arc-skill-eval run <your-skill> --case <first-case-id>If the judge’s evidence reads like it paraphrased the assertion instead of citing assistant text, tighten the assertion (more literal, more specific) or swap it for a script assertion.
Add regression cases
Section titled “Add regression cases”Add a case when you fix a regression. The new case should reproduce the failure and prevent it from returning.