Skip to content

Inspiration and credits

arc-skill-eval, branded Skeval in these docs, combines a grading workflow, file format, and runtime design from earlier work. This page credits those sources.

Testing Agent Skills Systematically with Evals, by Dominik Kundel and Gabriel Chua, OpenAI Developers Blog, January 22, 2026.

The post informed these parts of Skeval:

  • The unit of evaluation: “An eval is: a prompt → a captured run (trace + artifacts) → a small set of checks → a score you can compare over time.” Skeval uses this shape for each case.
  • Layered grading: “Start small, then layer in deeper checks only where they add real confidence.” Skeval runs deterministic assertions and batches prose claims into one judge call.
  • Success criteria across outcome, process, style, and efficiency. arc-creating-evals asks authors about these criteria.
  • Starter suite size: “For a single skill, a small set of 10–20 prompts is enough.” Skeval recommends six to ten cases for an initial suite.
  • Inspectable failures: “Everything is deterministic and debuggable. If a check fails, you can open the JSONL file and see exactly what happened.” Skeval writes grading.json, trace.json, and tool-summary.json for each case.
  • Cases from observed failures: “Every manual fix you make here is a candidate for a future eval.” Authoring evals applies this rule.

If you read one source before authoring an eval, read OpenAI’s.

Anthropic’s documented skill-eval methodology, Anthropic Platform Documentation.

Anthropic documents:

  • evals/evals.json next to SKILL.md as the on-disk shape.
  • EvalCase with id, prompt, expected_output, assertions.
  • Per-case grading.json with assertion_results[] and a summary block.
  • with_skill / without_skill execution as the dual-run comparison model.
  • benchmark.json as the aggregate over a --compare run.

Skeval’s loader, runner, and writers use this format so other compatible tools can read its output.

Anthropic Engineering’s Demystifying evals for AI agents by Mikaela Grace et al., January 9, 2026, describes code, model, and human graders; pass@k and pass^k; capability and regression evals; and layered evaluation. Skeval uses this terminology where it applies.

Runtime design from Ampcode and Mihail Eric

Section titled “Runtime design from Ampcode and Mihail Eric”

Two posts describe a small tool-calling loop as the core of a coding agent. Skeval follows that structure.

Thorsten Ball, Ampcode, April 15, 2025. Ball writes, “It’s an LLM, a loop, and enough tokens. It’s what we’ve been saying on the podcast from the start.” He builds a code-editing agent in Go with the Anthropic SDK and three tools: read_file, list_files, and edit_file. The loop accepts input, sends the conversation to the model, executes requested tools, returns results, and repeats until the model returns text.

What carried into Skeval:

  • Skeval passes a prompt to the model, executes tool calls, returns results, and repeats.
  • Tool definitions combine a name, description, schema, and function. tool-summary.json reports calls to those tools.
  • Fixtures, traces, and manifests support the loop. Ball summarizes the core: “300 lines of code and three tools and now you’re able to talk to an alien intelligence that edits your code.”

Mihail Eric, January 2026. Eric describes a harness of “about 200 lines of straightforward Python.” It uses a tool registry, parses tool: NAME({JSON}) calls, and runs tools until the model responds without another request. The example starts with read, list, and edit tools.

What carried into Skeval:

  • Skeval exposes a flat tool registry during skill runs.
  • The model requests file operations; the runner executes them. tool-summary.json and context-manifest.json record those actions and available tools.
  • Skeval supports more than three tools, but many cases need only read, write, and bash.

If you’re trying to write your first eval, in this order:

  1. Read OpenAI’s eval-skills post for the grading workflow.
  2. Follow Skeval’s Quickstart and hello-world example.
  3. Use Anthropic’s skill-eval docs while authoring.

If you’re trying to extend the framework or build a similar runner, in this order:

  1. Read Ball’s “How to Build an Agent.”
  2. Read Eric’s “The Emperor Has No Clothes.”
  3. Read Skeval’s Concepts section, then src/.

Related sources:

If you’ve published something on this topic that should be cited here, open an issue.