Quickstart
This page gets you from zero to a green eval run on the bundled hello-world skill in about five minutes. Everything here uses Anthropic’s evals/evals.json standard; nothing about the format is Skeval-specific.
Requirements
Section titled “Requirements”- Node.js ≥ 20.
- Pi installed and configured with at least one provider API key (Anthropic, OpenAI, Google/Gemini, Mistral, xAI, etc.). The skill’s assistant runs via
@mariozechner/pi-coding-agent.
Install
Section titled “Install”npm install --global arc-skill-evalarc-skill-eval --helpFrom a local checkout:
git clone https://github.com/andysolomon/arc-skill-evalcd arc-skill-evalnpm installnpm run buildnpm linkarc-skill-eval --helpRun the bundled hello-world skill
Section titled “Run the bundled hello-world skill”The package ships a deterministic reference skill. Resolve its path with bundled so you never need to know where npm installed the package:
arc-skill-eval run "$(arc-skill-eval bundled hello-world)"From a local checkout you can also use the repo-relative path:
arc-skill-eval run ./skills/hello-worldYou’ll see one stdout line per case (PASS/FAIL plus the case id), and a summary at the end with the total pass rate and total wall-clock time.
Inspect the artifacts
Section titled “Inspect the artifacts”Every run writes a per-case artifact tree:
skills/hello-world/evals-runs/<runId>/├── eval-default-world/│ ├── assistant.md│ ├── outputs/greeting.txt│ ├── timing.json│ ├── grading.json│ ├── trace.json│ ├── tool-summary.json│ └── context-manifest.json├── eval-named-ada/│ └── ...└── eval-assistant-names-file/ └── ...Open grading.json for any case to see each assertion’s verdict and evidence in Anthropic’s published format:
{ "case_id": "default-world", "assertion_results": [ { "text": "file-exists: greeting.txt", "passed": true, "evidence": "Found greeting.txt (15 bytes)" }, { "text": "regex-match: Hello, world!", "passed": true, "evidence": "matched in greeting.txt" }, { "text": "The response mentions the file `greeting.txt` by name", "passed": true, "evidence": "\"I wrote greeting.txt with the requested message.\"" } ], "summary": { "passed": 3, "failed": 0, "total": 3, "pass_rate": 1.0 }}Pin a model when you need reproducibility
Section titled “Pin a model when you need reproducibility”By default, Skeval inherits Pi’s configured provider/model from your Pi settings. For dogfood runs, CI, or provider experiments, pin the runner and judge explicitly:
arc-skill-eval run ./skills/hello-world \ --model openai-codex/gpt-5.5:medium \ --judge-model openai-codex/gpt-5.5:mediumFor a low-cost Ollama Cloud smoke test:
arc-skill-eval run ./skills/hello-world \ --model ollama-cloud/gpt-oss:20b \ --judge-model ollama-cloud/gpt-oss:20bSee Runtime & Models for provider setup and Pi default behavior.
Scaffold evals for your own skill
Section titled “Scaffold evals for your own skill”For a new skill, start with deterministic create when the success signal is mechanical: a file exists, JSON is valid, a command output matches a regex, or a fixture should be changed in a predictable way.
arc-skill-eval create ./skills/my-skill --dry-run --summaryarc-skill-eval create ./skills/my-skillUse guided create when the skill is more conceptual and the hard part is designing cases and assertions:
arc-skill-eval create ./skills/my-skill --guided --interactiveA grill-me interview skill is a good example. It probably should not assert that a file exists. Instead, it needs judge assertions for behavior like asking direct follow-up questions, challenging assumptions, and staying in interview mode, plus adjacent negatives for nearby requests like “summarize this plan” that should not trigger the skill.
Compare against a no-skill baseline
Section titled “Compare against a no-skill baseline”The pass-rate difference shows whether the skill helped. Add --compare to run each case once with the skill and once without it. The command also writes a top-level benchmark.json:
arc-skill-eval run ./skills/hello-world --compareYou now get sibling with_skill/ and without_skill/ directories under each eval-<id>/, and the run root contains benchmark.json with per-case deltas plus an overall delta. See With/without skill for why this comparison matters more than any absolute score, and Dogfooding & Authoring Loop for a practical iteration workflow.
Review and improve
Section titled “Review and improve”Generate a static report after a run:
arc-skill-eval review ./skills/hello-world/evals-runs/<runId>Open review.html to inspect assistant output, deterministic assertion evidence, judge evidence, tool summaries, and compare deltas. Add human notes to feedback.json, then use the feedback-driven improvement workflow when you are ready to plan changes:
arc-skill-eval improve ./skills/hello-world \ --from-feedback ./skills/hello-world/evals-runs/<runId>/feedback.jsonThe usual loop is to create a suite, run one case, run with --compare, inspect review.html, record feedback, and change the skill or evals. Start a new --iteration and repeat.
Where next
Section titled “Where next”- CLI reference. See every
arc-skill-eval runflag. - Runtime and models. Configure model pinning, Pi defaults, GPT 5.5, and Ollama Cloud.
- Authoring evals. Write
evals/evals.jsonfor your own skill. - Dogfooding and authoring loop. Iterate with
--compareand run artifacts. - Hello-world example. Walk through the skill you just ran.