Field Guides · Unbothered Publishing
Evals in a Nutshell
Measuring what your AI actually does
A vendor-neutral field guide to evaluating LLM applications and agents: datasets, graders, judges, statistics, RAG, tools, CI, production and red teams, and the habits that keep quality honest. Each chapter gets a labelled diagram that builds step by step, a one-minute film and the full chapter in the book.
100
chapters
10
parts
100
films
4
languages
Loading the hundred…