RAG
Ask a question over a document set and see the retrieved passages, scores and a cited answer.
- Embeddings
- Vector search
- Citations
Ask a question about Northwind's internal documents
Pick a suggested question or write your own. Toggle documents on the left to change what the search can see.
- 1Embed the questionTurned into a vector with the same model used to index the documents.
- 2Retrieve passagesNearest chunks from the documents in scope, reranked and filtered by a threshold.
- 3Answer with citationsWritten only from those passages, with every claim linked to its source.
How it works in production
The demo above runs on a script in your browser. This is the architecture it stands in for.
Connectors
SharePoint, Drive, Confluence
Parser and chunker
Embeddings
pgvector
Chunks and metadata
Permissions
Access list per document
Hybrid search
Vector and keyword
Reranker
LLM
Answer with citations
Thumbs and comments
Unanswered questions
Documents are split into overlapping chunks that keep their headings, then embedded and stored with their source, page and access permissions. Search only returns chunks the person asking is allowed to read.
The model is instructed to answer only from the retrieved passages and to cite each claim. If nothing clears the similarity threshold, it says so instead of guessing, and the question is logged so a knowledge manager can fill the gap.
- Hybrid retrieval (vector plus keyword) with a reranking pass
- Citations link to the exact passage so answers can be verified
- An evaluation set of real questions is re-run on every change to chunking or prompts
Build something like this
Tell me about the process you want to improve and the systems it touches. I'll come back with questions and a suggested approach.