Inspiration and credits
arc-skill-eval, branded Skeval in these docs, combines a grading workflow, file format, and runtime design from earlier work. This page credits those sources.
Grading workflow from OpenAI
Section titled “Grading workflow from OpenAI”Testing Agent Skills Systematically with Evals, by Dominik Kundel and Gabriel Chua, OpenAI Developers Blog, January 22, 2026.
The post informed these parts of Skeval:
- The unit of evaluation: “An eval is: a prompt → a captured run (trace + artifacts) → a small set of checks → a score you can compare over time.” Skeval uses this shape for each case.
- Layered grading: “Start small, then layer in deeper checks only where they add real confidence.” Skeval runs deterministic assertions and batches prose claims into one judge call.
- Success criteria across outcome, process, style, and efficiency.
arc-creating-evalsasks authors about these criteria. - Starter suite size: “For a single skill, a small set of 10–20 prompts is enough.” Skeval recommends six to ten cases for an initial suite.
- Inspectable failures: “Everything is deterministic and debuggable. If a check fails, you can open the JSONL file and see exactly what happened.” Skeval writes
grading.json,trace.json, andtool-summary.jsonfor each case. - Cases from observed failures: “Every manual fix you make here is a candidate for a future eval.” Authoring evals applies this rule.
If you read one source before authoring an eval, read OpenAI’s.
File format from Anthropic
Section titled “File format from Anthropic”Anthropic’s documented skill-eval methodology, Anthropic Platform Documentation.
Anthropic documents:
evals/evals.jsonnext toSKILL.mdas the on-disk shape.EvalCasewithid,prompt,expected_output,assertions.- Per-case
grading.jsonwithassertion_results[]and asummaryblock. with_skill/without_skillexecution as the dual-run comparison model.benchmark.jsonas the aggregate over a--comparerun.
Skeval’s loader, runner, and writers use this format so other compatible tools can read its output.
Anthropic Engineering’s Demystifying evals for AI agents by Mikaela Grace et al., January 9, 2026, describes code, model, and human graders; pass@k and pass^k; capability and regression evals; and layered evaluation. Skeval uses this terminology where it applies.
Runtime design from Ampcode and Mihail Eric
Section titled “Runtime design from Ampcode and Mihail Eric”Two posts describe a small tool-calling loop as the core of a coding agent. Skeval follows that structure.
Thorsten Ball, Ampcode, April 15, 2025. Ball writes, “It’s an LLM, a loop, and enough tokens. It’s what we’ve been saying on the podcast from the start.” He builds a code-editing agent in Go with the Anthropic SDK and three tools: read_file, list_files, and edit_file. The loop accepts input, sends the conversation to the model, executes requested tools, returns results, and repeats until the model returns text.
What carried into Skeval:
- Skeval passes a prompt to the model, executes tool calls, returns results, and repeats.
- Tool definitions combine a name, description, schema, and function.
tool-summary.jsonreports calls to those tools. - Fixtures, traces, and manifests support the loop. Ball summarizes the core: “300 lines of code and three tools and now you’re able to talk to an alien intelligence that edits your code.”
The Emperor Has No Clothes: How to Code Claude Code in 200 Lines of Code
Section titled “The Emperor Has No Clothes: How to Code Claude Code in 200 Lines of Code”Mihail Eric, January 2026. Eric describes a harness of “about 200 lines of straightforward Python.” It uses a tool registry, parses tool: NAME({JSON}) calls, and runs tools until the model responds without another request. The example starts with read, list, and edit tools.
What carried into Skeval:
- Skeval exposes a flat tool registry during skill runs.
- The model requests file operations; the runner executes them.
tool-summary.jsonandcontext-manifest.jsonrecord those actions and available tools. - Skeval supports more than three tools, but many cases need only read, write, and bash.
Suggested reading order
Section titled “Suggested reading order”If you’re trying to write your first eval, in this order:
- Read OpenAI’s eval-skills post for the grading workflow.
- Follow Skeval’s Quickstart and hello-world example.
- Use Anthropic’s skill-eval docs while authoring.
If you’re trying to extend the framework or build a similar runner, in this order:
- Read Ball’s “How to Build an Agent.”
- Read Eric’s “The Emperor Has No Clothes.”
- Read Skeval’s Concepts section, then
src/.
Other writing in the space
Section titled “Other writing in the space”Related sources:
- Anthropic Console docs, Using the Evaluation Tool, describe prompt-level evals with variables and side-by-side comparison.
- Inspiration only, not consulted for design: the agents-v2 Obsidian vault has notes on single-turn evals and multi-turn evals that point at where the single-prompt → captured-run methodology starts to break down. Part 4 of the blog picks up the multi-turn thread.
If you’ve published something on this topic that should be cited here, open an issue.