Skip to content

Concepts

This section is the architectural map of arc-skill-eval. Each sub-page covers one runtime entity and how it shows up in the artifacts a run produces. Read them in order if you’re new; jump straight to the one you need if you’re not.

PageWhat it covers
SkillsWhat counts as a skill, how Skeval discovers them, and the domain types that classify capabilities and policy.
Eval casesThe shape of an EvalCase and the workspace setup options (empty, seeded, fixture) that prepare a case to run.
AssertionsThe discriminated union of LLM-judged strings, legacy script assertions, and intent assertions.
With/without skillThe dual-run comparison and how to read the pass-rate delta.
GradingHow the LLM-judge and deterministic scripts coexist, batched into one judge call where possible.
ArtifactsThe per-case output tree (assistant.md, outputs/, grading.json, timing.json, trace.json, tool-summary.json, context-manifest.json) and the run-level benchmark.json.

The full pipeline (discover, load, materialize, run, grade, write) is summarized at the bottom of the Skills page.