Curso · 5 capítulos

LLM Evaluation

Evaluation as a discipline — from your first eval to LLM-as-judge rigor, eval suites at scale, and CI gating for production AI

De pagoadvanced5 capítulosInglés + 6 idiomasCertificado al completar

Lo que sabrás hacer

  • A 12-minute orientation to the LLM Evaluation skill path — the gateway chapter, then the three layers (judges, suites, gates) that turn eval-by-vibes into a discipline that ships.
  • Stop checking outputs by vibes. Build a runnable eval — golden dataset, deterministic scorer, LLM judge — and read the result like an engineer.
  • Design judges that survive CALM biases, calibrate against humans, and earn a place in your CI gate.
  • Author, run, and visualize frontier-grade eval suites with UK AISI's open-source framework.
  • Wire per-PR evals into GitHub Actions, pick thresholds that survive flakiness, and decide when a gate belongs on main.

Qué incluye

  1. 1
    LLM Evaluation: Start Here

    A 12-minute orientation to the LLM Evaluation skill path — the gateway chapter, then the three layers (judges, suites, gates) that turn eval-by-vibes into a discipline that ships.

  2. 2
    Eval Foundations: Your First LLM Eval in 30 Minutes

    Stop checking outputs by vibes. Build a runnable eval — golden dataset, deterministic scorer, LLM judge — and read the result like an engineer.

  3. 3
    LLM-as-Judge: Rubrics, Bias, and Reliability

    Design judges that survive CALM biases, calibrate against humans, and earn a place in your CI gate.

  4. 4
    Inspect AI: Production Eval Suites at Scale

    Author, run, and visualize frontier-grade eval suites with UK AISI's open-source framework.

  5. 5
    Eval Gating in CI: Blocking Bad Merges

    Wire per-PR evals into GitHub Actions, pick thresholds that survive flakiness, and decide when a gate belongs on main.

Consigue un certificado

Completa todos los capítulos para recibir tu certificado de finalización.