Dogfooding and authoring loop
Skeval works best when eval authoring and skill improvement are the same loop: write a small suite, run it, compare with and without the skill, inspect failures, tighten assertions, and repeat.
The loop
Section titled “The loop”- Create or update evals for the target skill. Use deterministic
createfor concrete artifacts; usecreate --guided --interactivefor conceptual or semantic behavior. - Run one case to catch obvious fixture, assertion, or model issues.
- Run with
--compareto measure whether the skill helps relative to the no-skill baseline. - Generate a review report and inspect assistant output, files, grading evidence, traces, and benchmark deltas.
- Capture feedback in
feedback.jsonand use it to improve the skill or evals. - Record the next run under a new iteration.
arc-skill-eval create ./skills/my-skill --guided --interactivearc-skill-eval run ./skills/my-skill --case golden-path
arc-skill-eval run ./skills/my-skill \ --case golden-path \ --compare \ --iteration dogfood-1
arc-skill-eval review ./skills/my-skill/evals-runs/iteration-dogfood-1/<runId>arc-skill-eval improve ./skills/my-skill \ --from-feedback ./skills/my-skill/evals-runs/iteration-dogfood-1/<runId>/feedback.jsonCreating evals with create
Section titled “Creating evals with create”Use deterministic create when the success criteria are obvious and mechanical:
arc-skill-eval create ./skills/my-skill --dry-run --summaryarc-skill-eval create ./skills/my-skillUse guided create when you need help designing the cases:
arc-skill-eval create ./skills/my-skill --guided --interactiveA conceptual skill such as grill-me should usually lean on judge assertions and adjacent negatives. It does not need fake artifact checks; it needs evidence that the assistant asks hard follow-up questions, challenges assumptions, and does not trigger for nearby summarization or editing tasks.
Creating evals with arc-creating-evals
Section titled “Creating evals with arc-creating-evals”The companion arc-creating-evals skill helps with cases that need more judgment. Ask your agent to use it with prompts like:
- “Create evals for this skill.”
- “Make
arc-conventional-commitstestable.” - “Write
evals/evals.jsonand fixtures for this skill.”
It should:
- locate and summarize the target
SKILL.md - confirm success criteria before writing cases
- draft trigger, execution, and negative cases
- prefer deterministic assertions for file and JSON effects
- write
evals/evals.jsonand fixture files - dry-run at least one case before declaring the suite ready
See Authoring evals for the full contract.
Why compare mode matters
Section titled “Why compare mode matters”An absolute pass rate can be misleading. A skill that passes because the base model already knew what to do is less valuable than a skill that creates a measurable improvement over baseline.
--compare runs every case twice:
with_skillincludes the target skill and any explicit extras.without_skillomits the target skill but uses the same prompt, fixtures, and extras.
Skeval writes a benchmark.json with per-case and aggregate deltas.
A positive delta means the skill helped. Keep the case and consider adding nearby cases. A neutral delta often means the base model already passes or the assertion does not distinguish the two runs. A negative delta means the skill hurt the result; inspect the trace for a wrong trigger, an overly strict instruction, or a conflict with another skill.
Make assertions discriminating
Section titled “Make assertions discriminating”A good assertion checks the effect of the skill, not whether the assistant repeated the skill’s instructions.
Prefer this:
{ "type": "file-exists", "path": ".releaserc.json" }and this:
{ "type": "regex-match", "pattern": "conventionalcommits", "target": { "file": ".releaserc.json" }}over a weak prose assertion like:
"The assistant says it used Conventional Commits."Use LLM-judged string assertions for tone, explanation quality, or semantic properties that cannot be checked mechanically.
Review before improving
Section titled “Review before improving”review turns raw run artifacts into a human-readable handoff:
arc-skill-eval review ./skills/my-skill/evals-runs/<runId>Open review.html for case summaries, assistant output, assertion evidence, timing/model/tool metadata, and compare deltas. Use feedback.json for human notes such as:
- the case passed but did not prove the skill helped
- the adjacent negative is too broad or too easy
- a judge assertion is vague or accepts paraphrase without evidence
- a fixture is missing a real-world constraint
Then feed those notes into the improvement workflow:
arc-skill-eval improve ./skills/my-skill \ --from-feedback ./skills/my-skill/evals-runs/<runId>/feedback.jsonThe improvement plan should tell you whether to change the skill description, add fixtures, tighten assertions, or add/remove cases.
Use iteration buckets
Section titled “Use iteration buckets”Group runs by iteration so you can keep evidence from each improvement cycle:
arc-skill-eval run ./skills/my-skill --compare --iteration 1arc-skill-eval run ./skills/my-skill --compare --iteration 2arc-skill-eval run ./skills/my-skill --compare --iteration description-tuningThe artifacts land under:
<skillDir>/evals-runs/iteration-<name>/<runId>/Add distractor skills for conflict testing
Section titled “Add distractor skills for conflict testing”Use --extra-skill to test routing and context conflicts:
arc-skill-eval run ./skills/arc-conventional-commits \ --compare \ --extra-skill ./skills/release-please \ --iteration conflict-1In compare mode, with_skill receives the target plus extras. without_skill receives only the extras. That isolates whether the target skill adds value in a realistic crowded context.
Optimize the description for routing accuracy
Section titled “Optimize the description for routing accuracy”Most measured failures happen during routing. The description misses an implicit request or matches a nearby request that should not trigger the skill. optimize-description measures both errors:
# Generate should-trigger + adjacent near-miss prompts, tagged train/testarc-skill-eval optimize-description ./skills/my-skill --generate-only
# Review evals/description-evals.json by hand.# Edit the near-miss negatives, then score the current description.arc-skill-eval optimize-description ./skills/my-skill \ --eval-set ./skills/my-skill/evals/description-evals.json
# Optimize with a held-out split; nothing touches SKILL.md without --applyarc-skill-eval optimize-description ./skills/my-skill \ --eval-set ./skills/my-skill/evals/description-evals.json \ --max-iterations 5Each probe is a no-tools completion that compares the target description with real sibling-skill descriptions in rotated order. About 100 probes cost less than one agent-session eval case. The optimizer proposes candidates from train failures and chooses a winner by held-out test accuracy, which prevents a description from winning by memorizing train prompts. Apply a candidate only after it improves held-out accuracy. Then rerun the regular suite to check that better routing did not break execution.
Run the bundled dogfood suite
Section titled “Run the bundled dogfood suite”The bundled arc-creating-evals skill has its own suite at skills/arc-creating-evals/evals/. Run it from the repository root:
arc-skill-eval run skills/arc-creating-evalsThe suite gives the assistant a small fixture skill named arc-demo-file-writer and checks that arc-creating-evals produces a valid evals/evals.json. An adjacent-negative case checks that a plain unit-test request does not trigger eval authoring. create --guided uses the same authoring skill, so this suite covers both paths.
When triaging failures, follow the skill’s own Phase 5 rules: Judge error: evidence means the judge infrastructure failed (fix --judge-model or provider auth, not the assertion); paraphrased evidence means the assertion should be tightened or replaced with a script assertion.
Concrete dogfood evidence
Section titled “Concrete dogfood evidence”The arc-skills repo contains a dogfood suite for arc-creating-evals, the meta-skill that authors evals for other skills. A recent golden-path compare run showed a positive +16.7% with-skill delta after the suite was tightened to assert behavior unique to arc-creating-evals.
That is the shape of useful dogfooding:
- start with a real skill
- create a small eval suite
- run with and without the skill
- inspect failures
- tighten assertions so the suite measures the skill’s unique contribution
- repeat
For deeper notes, see docs/skill-eval-authoring-debrief.md.