Open to remote or hybrid AI Engineer positions
Competency

AI guardrails and safety: controls outside the prompt

A guardrail that depends on the model obeying is not a guardrail. This is how I put the controls in code and in the database in AgendaGo, and how I planned them for the platform, which is in construction.

Where I applied it

What guardrails are

They are the controls around the model, on the way in and on the way out: what it can receive, what it can call and what it is allowed to return. Asking the model in the prompt to behave helps, but it is not enough. What must not go wrong should be guaranteed by something the model does not control.

AgendaGo: the layers that are in production

  • A tool whitelist and a role filter.
  • The business comes from the JWT, never from the model. Forbidden parameters (business ids, tokens, admin flags) are stripped before execution.
  • Customer and file text is wrapped as data, not as instructions.
  • A detector for invented entities, and the provenance of each claim.
  • Rejection text is built by code, not by the model.
  • An edge rate limit, and a single database guard for the monthly per-business budget.
  • Postgres RLS enabled across the schema, and SECURITY DEFINER RPCs with a fixed search_path.

Where each control lives

The critical controls are in the database, not in the prompt. Applying a change proposed by the assistant requires a SECURITY DEFINER function for the admin role only; overbooking is blocked by a trigger; and stock moves only when a person confirms what the model read.

The cost is code and migrations for every rule. In exchange, none of those guarantees depends on the model paying attention.

The platform and Nux: in construction

The AI agents platform for companies is in construction, in phase 0, and is a design: it plans an input guardrail against prompt injection and an output guardrail that validates and cites, a read-only SQL Tool with an allowlist, validated SQL, a timeout and a row limit, and RLS per tenant_id.

In Nux, the Bonuxo assistant, writes need human confirmation, some of them strong confirmation, and respect the user’s permissions. Its LLM layer is built but switched off in production.

Related stack

  • PostgreSQL
  • Supabase Edge Functions
  • JWT