Open to remote or hybrid AI Engineer positions
Competency

LLM observability: what I record on every call and why

In AgendaGo, which is in production, every model call leaves a record of cost, latency and outcome. Langfuse is still part of a design, in a platform that is in construction.

Where I applied it

What observing an LLM system means

It means being able to answer, after the fact, what was called, with which model and prompt, how much it cost, how long it took and how it ended. Without that data you cannot explain a rising bill, a worsening latency or a bad answer.

AgendaGo: the ai_usage_log table (in production)

For every call the table stores: business, agent, user, model, prompt version, input, output and cache tokens, cost in dollars, latency, intent, tools used and result. The result can be ok, fallback, hand-off, error, budget exhausted, guardrail rejected or model rejected.

  • It stores no customer text.
  • Prompts are versioned in the repository, and the version is stored on every call.
  • A test requires every insert to declare model and prompt version.
  • The monthly per-business budget is enforced by a single database guard.

The decision: if it cannot be recorded, it is not served

Logging is fail-closed: if the usage insert fails, the turn is cut. It is an explicit trade-off. A logging failure is allowed to interrupt an answer, in exchange for never having unmeasured calls or a budget that cannot be controlled.

Langfuse and Nux: in construction

The AI agents platform for companies is in construction, in phase 0. Its design includes Langfuse for traces and costs, and evaluations over a question set. There is no code or data yet.

Nux, the Bonuxo assistant, has a usage log and per-plan cost caps, with the Langfuse integration in progress. Its LLM layer is built but switched off in production.

Related stack

  • PostgreSQL
  • Supabase
  • Langfuse