Vision with multimodal models: reading invoices and delivery notes without inventing
In AgendaGo, which is in production, a multimodal model reads supplier invoices and delivery notes to update stock. What matters is what the model decides and what the code decides.
What it is and its problem
A multimodal model takes images as well as text. Applied to documents, it can extract data from an invoice or a delivery note without a per-supplier template. The problem is that it can be wrong or invent a value with full confidence, and a misread value that moves stock is a business error.
How I did it in AgendaGo
The model extracts supplier, number, date, total and line items from each invoice or delivery note, and it is told to use null and never invent a value. The code does everything else:
- Deterministic validation caps lines and lengths and enforces a date window.
- Matching the lines against the catalog is deterministic code, not the model.
- What is read goes to a staging area. Stock moves only when a person confirms.
- A sha256 lock blocks uploading the same file twice.
- PDFs that have a text layer use text extraction instead of vision.
Decisions and trade-offs
I left the model only the extraction, which is what code cannot do. The interpretation, the catalog matching and the decision to move stock stay in deterministic code and with a person. It costs one confirmation step per document; in exchange, a reading error does not reach stock without someone seeing it.
Text that arrives inside a file is treated as data, not as an instruction. It is the defence against prompt injection: a delivery note that says “ignore the above” is still a delivery note.
Related stack
- OpenAI
- Supabase
- PostgreSQL