Skip to content

Skeval

Run and compare skill evals in Anthropic's published format.

Skeval documents arc-skill-eval, a Pi-native CLI for running skill evals. Put evals/evals.json next to SKILL.md. The CLI finds the suite, runs each case with the skill attached, grades the output, and writes artifacts you can inspect or compare.

The grading workflow comes from OpenAI’s Testing Agent Skills Systematically with Evals: start with a small suite, add cases from real failures, and compare the same prompts with and without the skill. Files follow Anthropic’s published format, including evals/evals.json, grading.json, and benchmark.json. The runtime design draws on Ampcode’s How to Build an Agent and Mihail Eric’s The Emperor Has No Clothes. Inspiration and credits has the full attribution.

The brand is Skeval. The package and the CLI binary are still arc-skill-eval.

Skills are markdown files with optional scripts. They can change model behavior without making the change easy to measure. Skeval runs the same prompt with and without the skill, then grades both runs against the same assertions. The pass-rate difference shows whether the skill helped.

  • evals/evals.json as the authoring format, using Anthropic’s published shape.
  • LLM-judged assertions for prose claims, deterministic script assertions for file presence, regex, and JSON validity.
  • Opt-in with_skill vs without_skill dual runs and a benchmark.json aggregate.
  • Per-case observability: assistant.md, outputs/, timing.json, trace.json, tool-summary.json, context-manifest.json.
  • Model pinning for runner and judge models, including GPT 5.5 and Ollama Cloud through Pi.
  • An authoring skill, arc-creating-evals, that turns success criteria into cases, assertions, and fixtures.
  • A public roadmap for runtime setup, review reports, feedback-driven changes, and description testing.
  • Quickstart. Install the CLI, run the hello-world skill, and read grading.json.
  • Concepts. Learn about skills, cases, assertions, grading, artifacts, and the with/without difference.
  • Authoring evals. Write evals/evals.json for your own skill.
  • Dogfooding and authoring loop. Use --compare, iterations, and artifacts to improve a skill.
  • Runtime and models. Configure Pi defaults, pinned models, GPT 5.5, Ollama Cloud, and eval-owned runtimes.
  • Skill creator roadmap. See planned commands and product work.
  • Blog. Read why skill evals matter and what I learned while building this project.
  • Inspiration and credits. Read the sources that shaped the project.