Concepts
This section is the architectural map of arc-skill-eval. Each sub-page covers one runtime entity and how it shows up in the artifacts a run produces. Read them in order if you’re new; jump straight to the one you need if you’re not.
| Page | What it covers |
|---|---|
| Skills | What counts as a skill, how Skeval discovers them, and the domain types that classify capabilities and policy. |
| Eval cases | The shape of an EvalCase and the workspace setup options (empty, seeded, fixture) that prepare a case to run. |
| Assertions | The discriminated union of LLM-judged strings, legacy script assertions, and intent assertions. |
| With/without skill | The dual-run comparison and how to read the pass-rate delta. |
| Grading | How the LLM-judge and deterministic scripts coexist, batched into one judge call where possible. |
| Artifacts | The per-case output tree (assistant.md, outputs/, grading.json, timing.json, trace.json, tool-summary.json, context-manifest.json) and the run-level benchmark.json. |
The full pipeline (discover, load, materialize, run, grade, write) is summarized at the bottom of the Skills page.