Skip to content

Dogfooding and authoring loop

Skeval works best when eval authoring and skill improvement are the same loop: write a small suite, run it, compare with and without the skill, inspect failures, tighten assertions, and repeat.

  1. Create or update evals for the target skill. Use deterministic create for concrete artifacts; use create --guided --interactive for conceptual or semantic behavior.
  2. Run one case to catch obvious fixture, assertion, or model issues.
  3. Run with --compare to measure whether the skill helps relative to the no-skill baseline.
  4. Generate a review report and inspect assistant output, files, grading evidence, traces, and benchmark deltas.
  5. Capture feedback in feedback.json and use it to improve the skill or evals.
  6. Record the next run under a new iteration.
Terminal window
arc-skill-eval create ./skills/my-skill --guided --interactive
arc-skill-eval run ./skills/my-skill --case golden-path
arc-skill-eval run ./skills/my-skill \
--case golden-path \
--compare \
--iteration dogfood-1
arc-skill-eval review ./skills/my-skill/evals-runs/iteration-dogfood-1/<runId>
arc-skill-eval improve ./skills/my-skill \
--from-feedback ./skills/my-skill/evals-runs/iteration-dogfood-1/<runId>/feedback.json

Use deterministic create when the success criteria are obvious and mechanical:

Terminal window
arc-skill-eval create ./skills/my-skill --dry-run --summary
arc-skill-eval create ./skills/my-skill

Use guided create when you need help designing the cases:

Terminal window
arc-skill-eval create ./skills/my-skill --guided --interactive

A conceptual skill such as grill-me should usually lean on judge assertions and adjacent negatives. It does not need fake artifact checks; it needs evidence that the assistant asks hard follow-up questions, challenges assumptions, and does not trigger for nearby summarization or editing tasks.

The companion arc-creating-evals skill helps with cases that need more judgment. Ask your agent to use it with prompts like:

  • “Create evals for this skill.”
  • “Make arc-conventional-commits testable.”
  • “Write evals/evals.json and fixtures for this skill.”

It should:

  • locate and summarize the target SKILL.md
  • confirm success criteria before writing cases
  • draft trigger, execution, and negative cases
  • prefer deterministic assertions for file and JSON effects
  • write evals/evals.json and fixture files
  • dry-run at least one case before declaring the suite ready

See Authoring evals for the full contract.

An absolute pass rate can be misleading. A skill that passes because the base model already knew what to do is less valuable than a skill that creates a measurable improvement over baseline.

--compare runs every case twice:

  • with_skill includes the target skill and any explicit extras.
  • without_skill omits the target skill but uses the same prompt, fixtures, and extras.

Skeval writes a benchmark.json with per-case and aggregate deltas.

A positive delta means the skill helped. Keep the case and consider adding nearby cases. A neutral delta often means the base model already passes or the assertion does not distinguish the two runs. A negative delta means the skill hurt the result; inspect the trace for a wrong trigger, an overly strict instruction, or a conflict with another skill.

A good assertion checks the effect of the skill, not whether the assistant repeated the skill’s instructions.

Prefer this:

{ "type": "file-exists", "path": ".releaserc.json" }

and this:

{
"type": "regex-match",
"pattern": "conventionalcommits",
"target": { "file": ".releaserc.json" }
}

over a weak prose assertion like:

"The assistant says it used Conventional Commits."

Use LLM-judged string assertions for tone, explanation quality, or semantic properties that cannot be checked mechanically.

review turns raw run artifacts into a human-readable handoff:

Terminal window
arc-skill-eval review ./skills/my-skill/evals-runs/<runId>

Open review.html for case summaries, assistant output, assertion evidence, timing/model/tool metadata, and compare deltas. Use feedback.json for human notes such as:

  • the case passed but did not prove the skill helped
  • the adjacent negative is too broad or too easy
  • a judge assertion is vague or accepts paraphrase without evidence
  • a fixture is missing a real-world constraint

Then feed those notes into the improvement workflow:

Terminal window
arc-skill-eval improve ./skills/my-skill \
--from-feedback ./skills/my-skill/evals-runs/<runId>/feedback.json

The improvement plan should tell you whether to change the skill description, add fixtures, tighten assertions, or add/remove cases.

Group runs by iteration so you can keep evidence from each improvement cycle:

Terminal window
arc-skill-eval run ./skills/my-skill --compare --iteration 1
arc-skill-eval run ./skills/my-skill --compare --iteration 2
arc-skill-eval run ./skills/my-skill --compare --iteration description-tuning

The artifacts land under:

<skillDir>/evals-runs/iteration-<name>/<runId>/

Add distractor skills for conflict testing

Section titled “Add distractor skills for conflict testing”

Use --extra-skill to test routing and context conflicts:

Terminal window
arc-skill-eval run ./skills/arc-conventional-commits \
--compare \
--extra-skill ./skills/release-please \
--iteration conflict-1

In compare mode, with_skill receives the target plus extras. without_skill receives only the extras. That isolates whether the target skill adds value in a realistic crowded context.

Optimize the description for routing accuracy

Section titled “Optimize the description for routing accuracy”

Most measured failures happen during routing. The description misses an implicit request or matches a nearby request that should not trigger the skill. optimize-description measures both errors:

Terminal window
# Generate should-trigger + adjacent near-miss prompts, tagged train/test
arc-skill-eval optimize-description ./skills/my-skill --generate-only
# Review evals/description-evals.json by hand.
# Edit the near-miss negatives, then score the current description.
arc-skill-eval optimize-description ./skills/my-skill \
--eval-set ./skills/my-skill/evals/description-evals.json
# Optimize with a held-out split; nothing touches SKILL.md without --apply
arc-skill-eval optimize-description ./skills/my-skill \
--eval-set ./skills/my-skill/evals/description-evals.json \
--max-iterations 5

Each probe is a no-tools completion that compares the target description with real sibling-skill descriptions in rotated order. About 100 probes cost less than one agent-session eval case. The optimizer proposes candidates from train failures and chooses a winner by held-out test accuracy, which prevents a description from winning by memorizing train prompts. Apply a candidate only after it improves held-out accuracy. Then rerun the regular suite to check that better routing did not break execution.

The bundled arc-creating-evals skill has its own suite at skills/arc-creating-evals/evals/. Run it from the repository root:

Terminal window
arc-skill-eval run skills/arc-creating-evals

The suite gives the assistant a small fixture skill named arc-demo-file-writer and checks that arc-creating-evals produces a valid evals/evals.json. An adjacent-negative case checks that a plain unit-test request does not trigger eval authoring. create --guided uses the same authoring skill, so this suite covers both paths.

When triaging failures, follow the skill’s own Phase 5 rules: Judge error: evidence means the judge infrastructure failed (fix --judge-model or provider auth, not the assertion); paraphrased evidence means the assertion should be tightened or replaced with a script assertion.

The arc-skills repo contains a dogfood suite for arc-creating-evals, the meta-skill that authors evals for other skills. A recent golden-path compare run showed a positive +16.7% with-skill delta after the suite was tightened to assert behavior unique to arc-creating-evals.

That is the shape of useful dogfooding:

  1. start with a real skill
  2. create a small eval suite
  3. run with and without the skill
  4. inspect failures
  5. tighten assertions so the suite measures the skill’s unique contribution
  6. repeat

For deeper notes, see docs/skill-eval-authoring-debrief.md.