From flakiness correction to multi-model comparison — the full toolkit, zero cloud required.
Compose freely — every evaluator works with N-run flakiness correction. Mix and match inside a single test case.
type: contains value: "deadline" case_sensitive: false
type: ends_with value: "."
type: max_length value: 100 unit: words
type: regex
value: '^\d{4}-\d{2}-\d{2}$'type: json_schema schema: type: object required: [name, id]
type: llm_judge criteria: "Is it polite?"
type: llm_judge
threshold: 7
rubric:
- dimension: tone
weight: 2type: custom evaluator: ./my-check.js
Run your full suite against multiple providers concurrently. Flakiness-corrected pass rates, per-check breakdowns, and estimated cost — all in one table.
Eight commands, all stable. From init to history — everything you need without the extras.
Features ship in order of demand. Star on GitHub to vote with your ⭐
No account. No cloud. No telemetry.
Node ≥ 18 and an API key is all you need.