Skip to content

Quickstart

This page gets you from zero to a green eval run on the bundled hello-world skill in about five minutes. Everything here uses Anthropic’s evals/evals.json standard; nothing about the format is Skeval-specific.

  • Node.js ≥ 20.
  • Pi installed and configured with at least one provider API key (Anthropic, OpenAI, Google/Gemini, Mistral, xAI, etc.). The skill’s assistant runs via @mariozechner/pi-coding-agent.
Terminal window
npm install --global arc-skill-eval
arc-skill-eval --help

From a local checkout:

Terminal window
git clone https://github.com/andysolomon/arc-skill-eval
cd arc-skill-eval
npm install
npm run build
npm link
arc-skill-eval --help

The package ships a deterministic reference skill. Resolve its path with bundled so you never need to know where npm installed the package:

Terminal window
arc-skill-eval run "$(arc-skill-eval bundled hello-world)"

From a local checkout you can also use the repo-relative path:

Terminal window
arc-skill-eval run ./skills/hello-world

You’ll see one stdout line per case (PASS/FAIL plus the case id), and a summary at the end with the total pass rate and total wall-clock time.

Every run writes a per-case artifact tree:

skills/hello-world/evals-runs/<runId>/
├── eval-default-world/
│ ├── assistant.md
│ ├── outputs/greeting.txt
│ ├── timing.json
│ ├── grading.json
│ ├── trace.json
│ ├── tool-summary.json
│ └── context-manifest.json
├── eval-named-ada/
│ └── ...
└── eval-assistant-names-file/
└── ...

Open grading.json for any case to see each assertion’s verdict and evidence in Anthropic’s published format:

{
"case_id": "default-world",
"assertion_results": [
{ "text": "file-exists: greeting.txt", "passed": true, "evidence": "Found greeting.txt (15 bytes)" },
{ "text": "regex-match: Hello, world!", "passed": true, "evidence": "matched in greeting.txt" },
{ "text": "The response mentions the file `greeting.txt` by name", "passed": true, "evidence": "\"I wrote greeting.txt with the requested message.\"" }
],
"summary": { "passed": 3, "failed": 0, "total": 3, "pass_rate": 1.0 }
}

By default, Skeval inherits Pi’s configured provider/model from your Pi settings. For dogfood runs, CI, or provider experiments, pin the runner and judge explicitly:

Terminal window
arc-skill-eval run ./skills/hello-world \
--model openai-codex/gpt-5.5:medium \
--judge-model openai-codex/gpt-5.5:medium

For a low-cost Ollama Cloud smoke test:

Terminal window
arc-skill-eval run ./skills/hello-world \
--model ollama-cloud/gpt-oss:20b \
--judge-model ollama-cloud/gpt-oss:20b

See Runtime & Models for provider setup and Pi default behavior.

For a new skill, start with deterministic create when the success signal is mechanical: a file exists, JSON is valid, a command output matches a regex, or a fixture should be changed in a predictable way.

Terminal window
arc-skill-eval create ./skills/my-skill --dry-run --summary
arc-skill-eval create ./skills/my-skill

Use guided create when the skill is more conceptual and the hard part is designing cases and assertions:

Terminal window
arc-skill-eval create ./skills/my-skill --guided --interactive

A grill-me interview skill is a good example. It probably should not assert that a file exists. Instead, it needs judge assertions for behavior like asking direct follow-up questions, challenging assumptions, and staying in interview mode, plus adjacent negatives for nearby requests like “summarize this plan” that should not trigger the skill.

The pass-rate difference shows whether the skill helped. Add --compare to run each case once with the skill and once without it. The command also writes a top-level benchmark.json:

Terminal window
arc-skill-eval run ./skills/hello-world --compare

You now get sibling with_skill/ and without_skill/ directories under each eval-<id>/, and the run root contains benchmark.json with per-case deltas plus an overall delta. See With/without skill for why this comparison matters more than any absolute score, and Dogfooding & Authoring Loop for a practical iteration workflow.

Generate a static report after a run:

Terminal window
arc-skill-eval review ./skills/hello-world/evals-runs/<runId>

Open review.html to inspect assistant output, deterministic assertion evidence, judge evidence, tool summaries, and compare deltas. Add human notes to feedback.json, then use the feedback-driven improvement workflow when you are ready to plan changes:

Terminal window
arc-skill-eval improve ./skills/hello-world \
--from-feedback ./skills/hello-world/evals-runs/<runId>/feedback.json

The usual loop is to create a suite, run one case, run with --compare, inspect review.html, record feedback, and change the skill or evals. Start a new --iteration and repeat.