Documentation
Everything for running local, artifact-based agent evals with evalctl β grouped into guides, concepts, and project notes.
Start here
| If you want to⦠| Read |
|---|---|
| Install evalctl | install |
| Run your first suite | quickstart |
| The command surface | commands |
| Drive it from an agent | agent-guide |
| Decode an error code | errors |
| How it compares | comparison |
Buckets
Guides
Install, run your first suite, the command surface, and the agent workflow.
Concepts
How durable runs, replay, command scorers, and the spoolctl queue work.
Project
How evalctl compares, what changed, and the security posture.