hello-world (worked example)
hello-world is a small reference skill that ships with arc-skill-eval. It exercises discovery, execution, grading, workspace capture, and artifact writing. If its evals pass, those pipeline stages are working.
You can find it in the repo at skills/hello-world/.
What SKILL.md says
Section titled “What SKILL.md says”The skill is intentionally trivial. From its SKILL.md:
When invoked, follow these steps exactly:
- Read the prompt and extract the intended greeter name.
- If the prompt names a specific recipient (“a greeting for Ada”, “greet Grace”), use that name verbatim.
- If the prompt is generic (“create a greeting”), use
world.- Use the Write tool to create a file called
greeting.txtin the workspace root. The file content is exactly one line:Hello, <name>!(note the comma, the space after it, and the trailing exclamation mark).- Reply with a single sentence that confirms the file was written and mentions the filename
greeting.txtexplicitly.
Two details make the skill straightforward to grade:
- The “exactly one line” + literal punctuation gives a script assertion something concrete to regex-match against. Models can paraphrase intent; they can’t paraphrase punctuation.
- The “must mention
greeting.txtexplicitly” rule gives the LLM-judge assertion a checkable claim — there’s a literal token to find in the assistant’s reply.
What evals/evals.json checks
Section titled “What evals/evals.json checks”Three cases:
{ "skill_name": "hello-world", "evals": [ { "id": "default-world", "prompt": "Create a greeting.", "expected_output": "greeting.txt exists in the workspace with 'Hello, world!' and the assistant reply names the file.", "assertions": [ { "type": "file-exists", "path": "greeting.txt" }, { "type": "regex-match", "pattern": "Hello, world!", "target": { "file": "greeting.txt" } }, "The response mentions the file `greeting.txt` by name" ] }, { "id": "named-ada", "prompt": "Create a greeting for Ada.", "expected_output": "greeting.txt contains 'Hello, Ada!' and the assistant reply confirms it greeted Ada.", "assertions": [ { "type": "file-exists", "path": "greeting.txt" }, { "type": "regex-match", "pattern": "Hello, Ada!", "target": { "file": "greeting.txt" } }, "The response confirms the greeting was written for Ada" ] }, { "id": "assistant-names-file", "prompt": "Please generate a greeting for Grace Hopper.", "expected_output": "greeting.txt exists; assistant text mentions 'greeting.txt' literally.", "assertions": [ { "type": "file-exists", "path": "greeting.txt" }, { "type": "regex-match", "pattern": "greeting\\.txt", "target": "assistant-text" }, { "type": "regex-match", "pattern": "Hello, Grace Hopper!", "target": { "file": "greeting.txt" } } ] } ]}Notable patterns:
- Two assertions per case from scripts, one from the judge. The script assertions cover the mechanical facts (file exists, regex matches the file body); the judge handles the prose claim.
target: "assistant-text"on the third case. The sameregex-matchmachinery that checks workspace files can check the assistant’s literal reply.Hello, Grace Hopper!— the third case proves multi-word names round-trip. If the skill mangles the name, the regex fails.- No
setup, nofiles. Every case starts from an empty workspace. Hello-world isn’t testing repo state; it’s testing one skill’s effect on an empty directory.
Run it
Section titled “Run it”arc-skill-eval run ./skills/hello-worldStdout shows one line per case (PASS/FAIL + id) plus a summary. Total wall-clock is dominated by the model call latency — typically under a minute on Anthropic, somewhat faster on OpenAI.
Read the artifacts
Section titled “Read the artifacts”After a clean run, you’ll have:
skills/hello-world/evals-runs/<runId>/├── eval-default-world/│ ├── assistant.md # "I wrote greeting.txt with the requested message."│ ├── outputs/greeting.txt # "Hello, world!"│ ├── timing.json│ ├── grading.json # 3 assertions, all passed│ ├── trace.json│ ├── tool-summary.json # tool_calls_by_name: { write: 1 }│ └── context-manifest.json # only attached skill: hello-world├── eval-named-ada/│ └── ...└── eval-assistant-names-file/ └── ...Three things to spot-check on a successful run:
grading.json.summary.pass_rate === 1.0for each case. Anything less and one of the assertions failed.tool-summary.json.tool_calls_by_nameshowswrite: 1and nothing else. If the skill ranbash, theSKILL.mdrule “Do not run shell commands” was violated.context-manifest.json.attached_skillshas exactly one entry —hello-world, roletarget. Anything else means an ambient skill leaked into the context.
Compare against the no-skill baseline
Section titled “Compare against the no-skill baseline”arc-skill-eval run ./skills/hello-world --compareCompare mode shows whether the skill instructions change the result. With the skill attached, modern Claude and GPT models usually pass all three cases. Without it, the model often writes a different filename or skips the exact punctuation contract. The benchmark.json delta records how pass rates differ between the two runs.
Use it as a template
Section titled “Use it as a template”If you’re authoring evals for your first skill, copy this directory verbatim and edit:
SKILL.md— replace the steps with your skill’s logic. Keep at least one mechanical detail the model can’t paraphrase (a literal filename, a specific word, an exact format).evals/evals.json— replace cases. Aim for 6–10. Layer your assertions: scripts first, judge for prose only.- (Optional)
evals/files/<fixture>/— add fixtures for cases that need pre-existing repo state. Hello-world doesn’t have any.
For more guidance, see Authoring evals.