Skip to content

Authoring evals

You can create a suite in four ways: run arc-skill-eval create, run arc-skill-eval create --guided, use the bundled arc-creating-evals skill, or write evals/evals.json yourself. Each method creates an evals/ directory next to SKILL.md, with evals.json and any fixture files.

Start with deterministic create when the skill has obvious mechanical outcomes:

Terminal window
arc-skill-eval create ./skills/my-skill --dry-run --summary
arc-skill-eval create ./skills/my-skill

This mode does not call a model. It reads SKILL.md, creates trigger, execution, and adjacent-negative cases, infers fixture inputs, and adds deterministic assertions when the skill names files, JSON, or text patterns.

Use guided create when the eval design itself needs judgment:

Terminal window
arc-skill-eval create ./skills/my-skill --guided --interactive

Use guided mode for conceptual, planning, review, and routing skills. Review its proposal before writing files: remove weak cases, edit vague assertions, and keep deterministic checks for concrete outputs.

A grill-me skill interviews a user about a plan. It may not produce a report.json or plan.md file. Its suite can check:

  • explicit trigger cases such as “grill me on this launch plan”
  • adjacent negatives such as “summarize this launch plan” or “rewrite this plan more clearly”
  • judge assertions that the assistant asks pointed follow-up questions, challenges assumptions, explores tradeoffs, and avoids solving the plan for the user too early

Do not invent file assertions for files the skill does not create. Use deterministic assertions for concrete outputs and judge assertions for tone, interview behavior, and other prose properties.

Check the skill’s stated behavior: created files, JSON shape, avoided commands, tool use, routing decisions, and response content. A correct response should not fail because it uses different wording.

Brittle examples:

"The response says exactly: Phase 1 — detection complete."
"The response uses the heading ## Recommended Fix."
"The response mentions the arc-conventional-commits skill name."

Checks for observable results:

{ "type": "file-exists", "path": ".releaserc.json" }
{ "type": "regex-match", "pattern": "conventionalcommits", "target": { "file": ".releaserc.json" } }
"The response names semantic-release and describes configuring release automation for this repository rather than giving generic advice."

Exact wording is still valid when the words are the feature. Commit messages, CLI output, public copy, email subjects, and mandatory safety text may all need literal checks. In those cases, say so in expected_output and prefer deterministic assertions:

{
"expected_output": "The assistant returns exactly one conventional commit message beginning with fix:",
"assertions": [
{ "type": "regex-match", "pattern": "^fix(\\(.+\\))?:", "flags": "m", "target": "assistant-text" }
]
}

arc-creating-evals ships in this repository. Install it in your tool’s skills directory, such as .claude/skills/ for Claude Code. See skills/README.md. Then use prompts such as:

  • “Write evals for this skill.”
  • “Create evals.json for arc-conventional-commits.”
  • “Make this skill testable.”

It walks five phases and stops at each boundary for confirmation:

  1. Locate and summarize the target skill. The skill reads SKILL.md and extracts its name, description, phases, expected tools, and file changes. It asks, “Did I capture the skill’s behavior correctly?”
  2. Define success across outcome, process, style, and efficiency.
    • Outcome: what file state, response content, or artifact shows that the skill worked?
    • Process: which tools must or must not run?
    • Style: does tone, structure, or formatting matter?
    • Efficiency: is there a limit on tool calls or duration?
  3. Draft six to ten cases in three groups:
    • Three to five trigger cases, including explicit and implicit requests plus an adjacent negative.
    • One to three execution cases, including a fixture-backed main path and materially different alternate paths.
    • One or two edge or negative cases for malformed input, ambiguity, or conflicting state.
  4. Add two to five assertions per case. Start with scripts, then add judged strings for prose that scripts cannot check.
  5. Write and validate the files, run one case, and summarize the suite.

The skill pauses after the success criteria and case list. It does not write files until you approve the cases.

The shape (also documented under Concepts → Eval cases and Concepts → Assertions):

{
"skill_name": "arc-conventional-commits",
"evals": [
{
"id": 1,
"prompt": "Set up semantic-release in this repo.",
"expected_output": "semantic-release configured with the Conventional Commits preset.",
"files": ["files/clean-repo"],
"assertions": [
{ "type": "file-exists", "path": ".releaserc.json" },
{ "type": "regex-match", "pattern": "conventionalcommits", "target": { "file": ".releaserc.json" } },
"The response summarizes the semantic-release plugins it installed."
]
}
]
}

Place the file at <your-skill>/evals/evals.json. Fixtures, if any, go under <your-skill>/evals/files/<fixture-name>/.

Follow these rules for every authoring method:

  • Do not add cases for behavior the skill does not support.
  • Test string assertions against a real run before keeping them.
  • Do not copy SKILL.md instructions into assertions. Check the effect, such as valid JSON containing conventionalcommits.
  • Prefer one execution case with strong script assertions to several cases with weak text checks.
  • Use lowercase kebab-case IDs or numbers.
  • Ask the user when expected behavior is unclear.

evals.json doesn’t have a top-level model field; the skill’s assistant uses whatever your Pi default is. If your global default is quota-capped (for example, openai-codex on a ChatGPT Plus account), set a judgeModel override at the CLI layer or update ~/.pi/agent/settings.json.

Dry-run one case against your cheapest available model before committing the suite:

Terminal window
arc-skill-eval run <your-skill> --case <first-case-id>

If the judge’s evidence reads like it paraphrased the assertion instead of citing assistant text, tighten the assertion (more literal, more specific) or swap it for a script assertion.

Add a case when you fix a regression. The new case should reproduce the failure and prevent it from returning.