Open to remote or hybrid AI Engineer positions
Case study — In production

AgendaGo: a booking SaaS with three AI agents in production

I designed and run AgendaGo, a multi-business booking SaaS. This case explains how its three AI agents work, which barriers they have and how I evaluate and measure them.

AgendaGo — Multi-industry appointment-booking SaaS with three AI agents.
See the product

What it is and its status

AgendaGo is a multi-business booking SaaS: each business configures its schedule and its customers book. It is in production at agendago.com.ar. I designed it and I run it, and on top of the booking system I built three AI agents with tool calling.

This case covers the AI engineering side: how what a model can do is constrained, how it is evaluated and what is measured on every call.

The problem it solves

A business that lives on appointments has tasks that repeat every day. Each AgendaGo agent takes one of them:

  • Configuring the system: services, hours and rules. The dashboard assistant does it, at the owner’s request.
  • Answering and booking over WhatsApp. The WhatsApp agent does it, with the end customer.
  • Loading stock from supplier invoices and delivery notes. Reading with vision does it.

Architecture

The frontend is React with TypeScript, Vite, Tailwind and shadcn/ui. The backend lives in Supabase: PostgreSQL with RLS and Edge Functions. The external integrations are the WhatsApp Cloud API and MercadoPago.

AI runs in Edge Functions and calls OpenAI over direct HTTP, with no SDK, behind a provider abstraction layer. The response streams from the Edge to the frontend, so the owner sees the text as it is generated.

Dashboard assistant: it proposes, a person approves

The dashboard assistant has 21 tools: 11 read, 2 simulate or explain, and 8 propose. The propose tools never write. They create a proposal with the diff and its consequence, which expires after 15 minutes, and a human click applies it.

Applying the proposal does not depend on the model behaving: the database enforces it. It is a SECURITY DEFINER function, admin role only, with a row lock and an atomic step that applies the proposal and marks it as applied.

  • Three confirmation levels, derived from the risk of each action. Destructive ones need explicit confirmation and cannot go in a batch.
  • Per turn: at most 4 model tool rounds and 700 tokens, plus a per-turn spend cap.
  • A monthly budget per business, and runs can be cancelled.

WhatsApp agent: three barriers before booking

The WhatsApp agent has 9 tools: availability, booking, hours, address, services, coverage, requirements, re-ask and hand-off to a person.

Booking has three barriers, and none of them depends on what the model says:

  • The slot must come from the availability engine in the same turn.
  • The customer may not already have a booking that day.
  • A database trigger blocks overbooking as the final lock.

Everything sent to the customer goes through a single outbound path, which enforces the 24-hour window and WhatsApp templates. A deterministic router answers about 44% of messages without calling the model. The app passed Meta’s review on 31 August 2026.

Vision: supplier invoices and delivery notes

The model reads supplier invoices and delivery notes to update stock. It extracts supplier, number, date, total and line items, and the prompt tells it to use null and never invent a value.

  • Deterministic validation caps lines and lengths and enforces a date window.
  • Matching against the catalog is deterministic code, not the model.
  • What is read goes to a staging area: stock moves only when a person confirms.
  • A sha256 lock blocks uploading the same file twice. PDFs with a text layer use text extraction instead of vision.
  • Text inside a file is treated as data, not as instructions: that is the defence against prompt injection.

Evaluation and telemetry

For the dashboard assistant there is a 100-phrase dataset with the expected tool for each phrase, plus deterministic checks in CI. For WhatsApp there is a dataset of 52 scenarios and 15 conversations, with categories such as greeting, price, booking, ambiguity, adversarial and prompt injection. Each scenario declares the expected route, the criteria and the behaviour the agent must avoid.

Model evals run against the real stack with a spend cap, and a blind judge script exists.

The ai_usage_log table stores, per call: business, agent, user, model, prompt version, input, output and cache tokens, cost in dollars, latency, intent, tools used and result (ok, fallback, hand-off, error, budget exhausted, guardrail rejected or model rejected). It stores no customer text. If the insert fails, the turn is cut; a test requires every insert to declare model and prompt version, and prompts are versioned in the repository.

Reliability and security

  • A tool whitelist, filtered by role.
  • The business comes from the JWT, never from the model. Forbidden parameters (business ids, tokens, admin flags) are stripped.
  • Customer and file text is wrapped as data.
  • A detector for invented entities, plus the provenance of each claim.
  • Rejection text is built by code, not by the model.
  • An edge rate limit and a single database guard for the monthly budget.
  • Postgres RLS enabled across the schema, and SECURITY DEFINER RPCs with a fixed search_path.
  • The suite has 522 test files.

What is in production and what is not

Everything described on this page is in production: the three agents, the evals and the telemetry. What is still in construction does not belong to AgendaGo: it is Nux, the Bonuxo assistant, and the AI agents platform for companies, which is in phase 0, architecture.

Stack

  • React
  • TypeScript
  • Vite
  • Tailwind
  • shadcn/ui
  • Supabase (PostgreSQL, RLS, Edge Functions)
  • OpenAI
  • WhatsApp Cloud API
  • MercadoPago