Field Guides · Unbothered Publishing

Evals in a Nutshell
Measuring what your AI actually does

A vendor-neutral field guide to evaluating LLM applications and agents: datasets, graders, judges, statistics, RAG, tools, CI, production and red teams, and the habits that keep quality honest. Each chapter gets a labelled diagram that builds step by step, a one-minute film and the full chapter in the book.

Read the book
100
chapters
10
parts
100
films
4
languages
Loading the hundred…
Roll the dice