LLM evaluation: in-house datasets and CI checks
An agent that works in a demo is not evaluated. This is how I evaluate the ones in AgendaGo, which is in production: in-house datasets, deterministic checks and model evals against the real stack.
What evaluating an LLM means
Evaluating is measuring the agent’s behaviour with fixed cases and criteria written in advance, not trying a few phrases by hand. Without that, every prompt or model change can break something and nobody finds out until a customer reports it.
How I did it in AgendaGo (in production)
- Dashboard assistant: a dataset of 100 phrases, with the expected tool for each one, and deterministic checks in CI.
- WhatsApp agent: a dataset of 52 scenarios and 15 conversations. Categories include greeting, price, booking, ambiguity, adversarial and prompt injection.
- Each scenario declares the expected route, the pass criteria and the forbidden behaviour.
- Model evals run against the real stack, with a spend cap, and a blind judge script exists.
Decisions and trade-offs
I separated what can be checked without a model from what cannot. Deterministic checks run in CI on every change; model evals run against the real stack and cost money, which is why they carry a spend cap.
Having each scenario declare what the agent must not do matters as much as what it must do: an adversarial or prompt-injection case passes only if the agent does not fall for it.
Evaluation is complemented by the per-call telemetry in the ai_usage_log table, which records model and prompt version, and by a test that requires every insert to declare both. I cover that on the observability page.
Related stack
- OpenAI
- Supabase
- CI