Course
AI Evaluation
Measuring whether an AI system actually works -- and keeps working.
L3 · AdvancedFast-movingKnown~3 h
What you’ll learn
- Choose evaluation metrics that track the real objective
- Decide between offline and online evaluation and use LLM-as-judge carefully
- Evaluate RAG and agent systems on both outcome and process
Prerequisites
applied-llm-systems
Module 1. Foundations
What to measure, offline versus online, and LLM-as-judge.
Module 2. Evaluating Systems
Scoring RAG pipelines and multi-step agents.