14 section
Evaluation & Observability
How you know the system is good, and how you know it still is: judge design, retrieval scores, traces, token spend, and drift alerts on live traffic.
What is in here
Evaluation comes first, because an observability stack with no quality signal only tells you the errors were fast. That chapter covers LLM-as-judge calibration, RAGAS-style retrieval metrics and human annotation; Observability then wires the three pillars for non-deterministic systems and attributes cost per pipeline step. For the long version, the AI Evals Study Guide in Reference is roughly eight times the length.
01
10 min
LLM Evaluation
Building an eval that can actually fail: quality dimensions, calibrated LLM judges, human annotation, RAGAS-style retrieval scores, and monitoring that runs on live traffic.
evaluationragobservability
02
8 min
LLM Observability
Logs, metrics and traces reshaped for non-deterministic systems, where quality ranks alongside latency as a first-class signal and cost is attributed per pipeline step.
observabilitycostinfrastructure