Runtime and models
Skeval uses the Pi SDK by default. Pi supplies model providers, API keys, default models, and thinking levels unless you override them with CLI flags.
Optional CLI harnesses (--runtime codex, claude-code, cursor-agent, or copilot) run the same evals.json without Pi. They use credentials from environment variables or the harness’s login. See Multi-harness runtimes.
CLI harnesses do not support --sandbox just-bash (use --runtime pi-sdk for sandboxed bash). Harness stderr is redacted in traces; a non-zero CLI exit code fails the case even when partial assistant text was parsed.
Two model roles
Section titled “Two model roles”A run can use two different models:
- The runner model receives the prompt, skill instructions, tools, and fixture workspace.
- The judge model grades prose assertions. Deterministic assertions such as
file-exists,regex-match, andjson-validdo not call it.
Pin them independently:
arc-skill-eval run ./skills/arc-conventional-commits \ --model openai-codex/gpt-5.5:medium \ --judge-model mistral/ministral-8b-latestIf --judge-model is omitted, prose assertions use the same configured model path as the runner.
Provider/model syntax
Section titled “Provider/model syntax”Model flags use Pi’s provider/model form:
--model <provider/model[:thinking]>--judge-model <provider/model[:thinking]>Examples:
# GPT 5.5 with medium thinking.arc-skill-eval run ./skills/hello-world \ --model openai-codex/gpt-5.5:medium \ --judge-model openai-codex/gpt-5.5:medium
# Ollama Cloud model ID with a colon tag.arc-skill-eval run ./skills/hello-world \ --model ollama-cloud/gpt-oss:20b \ --judge-model ollama-cloud/gpt-oss:20bOllama model IDs commonly contain colon tags, for example gpt-oss:20b, qwen3.5:cloud, or qwen2.5-coder:1.5b. Skeval treats a final :suffix as a thinking level only when the suffix is a known thinking value such as off, low, medium, or high. Otherwise the colon stays part of the model ID.
Defaults without model flags
Section titled “Defaults without model flags”When no model flags are supplied, Skeval inherits Pi defaults from the configured Pi agent settings, normally:
~/.pi/agent/settings.jsonA minimal settings file looks like this:
{ "defaultProvider": "ollama-cloud", "defaultModel": "gpt-oss:20b", "defaultThinkingLevel": "off"}Set --model and --judge-model in CI or shared runs to use the same models across machines.
Ollama Cloud
Section titled “Ollama Cloud”Use Ollama Cloud for smoke tests that should not depend on a local model:
arc-skill-eval run ./skills/hello-world \ --model ollama-cloud/gpt-oss:20b \ --judge-model ollama-cloud/gpt-oss:20bIn one run, two of three hello-world cases passed. The failed case asked for a name instead of defaulting to Hello, world!; provider setup had succeeded.
Direct Ollama Cloud provider config
Section titled “Direct Ollama Cloud provider config”Set your key in the environment:
export OLLAMA_API_KEY=...Then add an ollama-cloud provider to Pi’s models.json:
{ "providers": { "ollama-cloud": { "baseUrl": "https://ollama.com/v1", "api": "openai-completions", "apiKey": "OLLAMA_API_KEY", "models": [ { "id": "gpt-oss:20b" }, { "id": "ministral-3:3b" }, { "id": "gemma3:4b" } ] } }}The apiKey value is the environment variable name. Do not commit literal API keys.
Eval-owned Pi runtime
Section titled “Eval-owned Pi runtime”Create an eval-owned runtime with init-runtime:
arc-skill-eval init-runtime ./.arc-skill-eval/pi-agent \ --provider ollama-cloud \ --model gpt-oss:20bThe command writes models.json and settings.json, refuses to overwrite existing files unless --force is supplied, and references OLLAMA_API_KEY instead of storing a literal secret.
Use --agent-dir when you want Skeval to load Pi settings, model registry, and auth from an eval-owned runtime directory instead of the normal personal Pi agent directory:
arc-skill-eval run ./skills/hello-world \ --agent-dir ./.arc-skill-eval/pi-agent \ --model ollama-cloud/gpt-oss:20b \ --judge-model ollama-cloud/gpt-oss:20bThe runtime can contain only the providers and settings required by the suite:
.arc-skill-eval/pi-agent/├── models.json└── settings.jsonWhen --agent-dir is supplied, both the skill runner and the default LLM judge use that directory for Pi config lookup. run preflights the directory before case execution and reports missing models.json, settings.json, provider/model entries, or required API-key environment variables once with an init-runtime remediation.
This setup provides:
- no dependence on personal
~/.pi/agentdefaults - reproducible team and CI runs
- no ambient skills, extensions, or prompt templates unless explicitly enabled
- provider setup that can live with the repo while secrets stay in environment variables
See the full design note in docs/agent-runtime-strategy.md.