Open to remote or hybrid AI Engineer positions
Competency

LLM evaluation: in-house datasets and CI checks

An agent that works in a demo is not evaluated. This is how I evaluate the ones in AgendaGo, which is in production: in-house datasets, deterministic checks and model evals against the real stack.

Where I applied it

What evaluating an LLM means

Evaluating is measuring the agent’s behaviour with fixed cases and criteria written in advance, not trying a few phrases by hand. Without that, every prompt or model change can break something and nobody finds out until a customer reports it.

How I did it in AgendaGo (in production)

  • Dashboard assistant: a dataset of 100 phrases, with the expected tool for each one, and deterministic checks in CI.
  • WhatsApp agent: a dataset of 52 scenarios and 15 conversations. Categories include greeting, price, booking, ambiguity, adversarial and prompt injection.
  • Each scenario declares the expected route, the pass criteria and the forbidden behaviour.
  • Model evals run against the real stack, with a spend cap, and a blind judge script exists.

Decisions and trade-offs

I separated what can be checked without a model from what cannot. Deterministic checks run in CI on every change; model evals run against the real stack and cost money, which is why they carry a spend cap.

Having each scenario declare what the agent must not do matters as much as what it must do: an adversarial or prompt-injection case passes only if the agent does not fall for it.

Evaluation is complemented by the per-call telemetry in the ai_usage_log table, which records model and prompt version, and by a test that requires every insert to declare both. I cover that on the observability page.

Related stack

  • OpenAI
  • Supabase
  • CI