Open to remote or hybrid AI Engineer positions
Competency

RAG and semantic search: the pipeline I designed

The RAG in my agents platform is a design, not a system in use: it is in construction, in phase 0. This page describes each stage of the pipeline and what does not exist yet.

Where I applied it

What RAG is

RAG (retrieval-augmented generation) answers with information the model did not have: first the relevant chunks of the documents are searched for, then they are handed to the model together with the question. Semantic search uses embeddings, vectors that represent the meaning of a text, to find chunks similar to the question even if they share no words.

The planned pipeline

  • Input: unstructured documents such as PDF, Word, scans and images that need OCR, and email, TXT and MD.
  • Parser and OCR with PyMuPDF and Tesseract, then chunking and embeddings, in workers on Redis and Celery.
  • Storage in PostgreSQL with pgvector: documents are stored with a hash and versions, and chunks with their embeddings.
  • Retrieval: the Retrieval Tool fetches the top-k chunks, always within a tenant.
  • Answer: the RAG agent, routed by the LangGraph orchestrator, answers, and an output guardrail validates the answer and cites the source.

Status

Related stack

  • PostgreSQL + pgvector
  • PyMuPDF
  • Tesseract
  • LangChain
  • Celery