Course · 5 chapters
LLM Evaluation
Evaluation as a discipline — from your first eval to LLM-as-judge rigor, eval suites at scale, and CI gating for production AI
What you'll be able to do
- A 12-minute orientation to the LLM Evaluation skill path — the gateway chapter, then the three layers (judges, suites, gates) that turn eval-by-vibes into a discipline that ships.
- Stop checking outputs by vibes. Build a runnable eval — golden dataset, deterministic scorer, LLM judge — and read the result like an engineer.
- Design judges that survive CALM biases, calibrate against humans, and earn a place in your CI gate.
- Author, run, and visualize frontier-grade eval suites with UK AISI's open-source framework.
- Wire per-PR evals into GitHub Actions, pick thresholds that survive flakiness, and decide when a gate belongs on main.
What's inside
- 1LLM Evaluation: Start Here
A 12-minute orientation to the LLM Evaluation skill path — the gateway chapter, then the three layers (judges, suites, gates) that turn eval-by-vibes into a discipline that ships.
- 2Eval Foundations: Your First LLM Eval in 30 Minutes
Stop checking outputs by vibes. Build a runnable eval — golden dataset, deterministic scorer, LLM judge — and read the result like an engineer.
- 3LLM-as-Judge: Rubrics, Bias, and Reliability
Design judges that survive CALM biases, calibrate against humans, and earn a place in your CI gate.
- 4Inspect AI: Production Eval Suites at Scale
Author, run, and visualize frontier-grade eval suites with UK AISI's open-source framework.
- 5Eval Gating in CI: Blocking Bad Merges
Wire per-PR evals into GitHub Actions, pick thresholds that survive flakiness, and decide when a gate belongs on main.
Earn a certificate
Complete all chapters to receive your certificate of completion.