Documentation
Everything runs from a clone today; package names are provisional until published. Start with the quickstart, then read how rubrics decide.
- QuickstartInstall from a clone, run a simulated evaluation, then a real one with your TypeSafe key.
- Evaluation casesThe case schema: input, output, messages, references, policy, tool events, metadata and expected labels.
- RubricsThe five check kinds, how deterministic rules compose with Jev questions, thresholds and statuses.
- CLIjeval init, run, report, compare, benchmark and labels.
- Reports and CI gatesPass-rate denominators, decided coverage, exit codes and how to loosen the policy deliberately.
- Jev integrationWhat is sent to TypeSafe, models, usage and cost, probabilities versus confidence.
- Benchmarking the judgeScore predicted outcomes against labels; tuning versus held-out splits; calibration.
- LimitationsPrompt injection, literal judging, synthetic labels, unverified integrations.