Open to remote or hybrid AI Engineer positions
Competency

Vision with multimodal models: reading invoices and delivery notes without inventing

In AgendaGo, which is in production, a multimodal model reads supplier invoices and delivery notes to update stock. What matters is what the model decides and what the code decides.

Where I applied it

What it is and its problem

A multimodal model takes images as well as text. Applied to documents, it can extract data from an invoice or a delivery note without a per-supplier template. The problem is that it can be wrong or invent a value with full confidence, and a misread value that moves stock is a business error.

How I did it in AgendaGo

The model extracts supplier, number, date, total and line items from each invoice or delivery note, and it is told to use null and never invent a value. The code does everything else:

  • Deterministic validation caps lines and lengths and enforces a date window.
  • Matching the lines against the catalog is deterministic code, not the model.
  • What is read goes to a staging area. Stock moves only when a person confirms.
  • A sha256 lock blocks uploading the same file twice.
  • PDFs that have a text layer use text extraction instead of vision.

Decisions and trade-offs

I left the model only the extraction, which is what code cannot do. The interpretation, the catalog matching and the decision to move stock stay in deterministic code and with a person. It costs one confirmation step per document; in exchange, a reading error does not reach stock without someone seeing it.

Text that arrives inside a file is treated as data, not as an instruction. It is the defence against prompt injection: a delivery note that says “ignore the above” is still a delivery note.

Related stack

  • OpenAI
  • Supabase
  • PostgreSQL