Skip to content

Runtime and models

Skeval uses the Pi SDK by default. Pi supplies model providers, API keys, default models, and thinking levels unless you override them with CLI flags.

Optional CLI harnesses (--runtime codex, claude-code, cursor-agent, or copilot) run the same evals.json without Pi. They use credentials from environment variables or the harness’s login. See Multi-harness runtimes.

CLI harnesses do not support --sandbox just-bash (use --runtime pi-sdk for sandboxed bash). Harness stderr is redacted in traces; a non-zero CLI exit code fails the case even when partial assistant text was parsed.

A run can use two different models:

  • The runner model receives the prompt, skill instructions, tools, and fixture workspace.
  • The judge model grades prose assertions. Deterministic assertions such as file-exists, regex-match, and json-valid do not call it.

Pin them independently:

Terminal window
arc-skill-eval run ./skills/arc-conventional-commits \
--model openai-codex/gpt-5.5:medium \
--judge-model mistral/ministral-8b-latest

If --judge-model is omitted, prose assertions use the same configured model path as the runner.

Model flags use Pi’s provider/model form:

--model <provider/model[:thinking]>
--judge-model <provider/model[:thinking]>

Examples:

Terminal window
# GPT 5.5 with medium thinking.
arc-skill-eval run ./skills/hello-world \
--model openai-codex/gpt-5.5:medium \
--judge-model openai-codex/gpt-5.5:medium
# Ollama Cloud model ID with a colon tag.
arc-skill-eval run ./skills/hello-world \
--model ollama-cloud/gpt-oss:20b \
--judge-model ollama-cloud/gpt-oss:20b

Ollama model IDs commonly contain colon tags, for example gpt-oss:20b, qwen3.5:cloud, or qwen2.5-coder:1.5b. Skeval treats a final :suffix as a thinking level only when the suffix is a known thinking value such as off, low, medium, or high. Otherwise the colon stays part of the model ID.

When no model flags are supplied, Skeval inherits Pi defaults from the configured Pi agent settings, normally:

~/.pi/agent/settings.json

A minimal settings file looks like this:

{
"defaultProvider": "ollama-cloud",
"defaultModel": "gpt-oss:20b",
"defaultThinkingLevel": "off"
}

Set --model and --judge-model in CI or shared runs to use the same models across machines.

Use Ollama Cloud for smoke tests that should not depend on a local model:

Terminal window
arc-skill-eval run ./skills/hello-world \
--model ollama-cloud/gpt-oss:20b \
--judge-model ollama-cloud/gpt-oss:20b

In one run, two of three hello-world cases passed. The failed case asked for a name instead of defaulting to Hello, world!; provider setup had succeeded.

Set your key in the environment:

Terminal window
export OLLAMA_API_KEY=...

Then add an ollama-cloud provider to Pi’s models.json:

{
"providers": {
"ollama-cloud": {
"baseUrl": "https://ollama.com/v1",
"api": "openai-completions",
"apiKey": "OLLAMA_API_KEY",
"models": [
{ "id": "gpt-oss:20b" },
{ "id": "ministral-3:3b" },
{ "id": "gemma3:4b" }
]
}
}
}

The apiKey value is the environment variable name. Do not commit literal API keys.

Create an eval-owned runtime with init-runtime:

Terminal window
arc-skill-eval init-runtime ./.arc-skill-eval/pi-agent \
--provider ollama-cloud \
--model gpt-oss:20b

The command writes models.json and settings.json, refuses to overwrite existing files unless --force is supplied, and references OLLAMA_API_KEY instead of storing a literal secret.

Use --agent-dir when you want Skeval to load Pi settings, model registry, and auth from an eval-owned runtime directory instead of the normal personal Pi agent directory:

Terminal window
arc-skill-eval run ./skills/hello-world \
--agent-dir ./.arc-skill-eval/pi-agent \
--model ollama-cloud/gpt-oss:20b \
--judge-model ollama-cloud/gpt-oss:20b

The runtime can contain only the providers and settings required by the suite:

.arc-skill-eval/pi-agent/
├── models.json
└── settings.json

When --agent-dir is supplied, both the skill runner and the default LLM judge use that directory for Pi config lookup. run preflights the directory before case execution and reports missing models.json, settings.json, provider/model entries, or required API-key environment variables once with an init-runtime remediation.

This setup provides:

  • no dependence on personal ~/.pi/agent defaults
  • reproducible team and CI runs
  • no ambient skills, extensions, or prompt templates unless explicitly enabled
  • provider setup that can live with the repo while secrets stay in environment variables

See the full design note in docs/agent-runtime-strategy.md.