Skip to content
AI Lab
Contact

RAG

Ask a question over a document set and see the retrieved passages, scores and a cited answer.

Simulated
Retrieval-augmented generation
  • Embeddings
  • Vector search
  • Citations
Northwind knowledge baseReady

Ask a question about Northwind's internal documents

Pick a suggested question or write your own. Toggle documents on the left to change what the search can see.

  1. 1Embed the questionTurned into a vector with the same model used to index the documents.
  2. 2Retrieve passagesNearest chunks from the documents in scope, reranked and filtered by a threshold.
  3. 3Answer with citationsWritten only from those passages, with every claim linked to its source.

How it works in production

The demo above runs on a script in your browser. This is the architecture it stands in for.

Ingestion

Connectors

SharePoint, Drive, Confluence

Parser and chunker

Embeddings

Storage

pgvector

Chunks and metadata

Permissions

Access list per document

Query

Hybrid search

Vector and keyword

Reranker

LLM

Answer with citations

Feedback

Thumbs and comments

Unanswered questions

Retrieval pipeline

Documents are split into overlapping chunks that keep their headings, then embedded and stored with their source, page and access permissions. Search only returns chunks the person asking is allowed to read.

The model is instructed to answer only from the retrieved passages and to cite each claim. If nothing clears the similarity threshold, it says so instead of guessing, and the question is logged so a knowledge manager can fill the gap.

  • Hybrid retrieval (vector plus keyword) with a reranking pass
  • Citations link to the exact passage so answers can be verified
  • An evaluation set of real questions is re-run on every change to chunking or prompts

Build something like this

Tell me about the process you want to improve and the systems it touches. I'll come back with questions and a suggested approach.