Document Intelligence
Classify a document, extract fields with confidence scores, and route low-confidence fields to review.
- Classification
- Extraction
- Human-in-the-loop
Sample document
- OCR
Read text and page layout
- Classifying
Work out what kind of document it is
- Extracting fields
Pull fields with a confidence score each
- Validating
Check the fields against business rules
Process the invoice to see what the model pulls out
It classifies the file, extracts each field with a confidence score and sends anything under 80% to you before it reaches the ERP.
Fields to extract
11 fields- VendorNot extracted
- Invoice numberNot extracted
- Invoice dateNot extracted
- Due dateNot extracted
- PO numberNot extracted
- Bill toNot extracted
- SubtotalNot extracted
- TaxNot extracted
- TotalNot extracted
- CurrencyNot extracted
- Payment termsNot extracted
How it works in production
The demo above runs on a script in your browser. This is the architecture it stands in for.
Email inbox
S3 upload
OCR
Classifier
Type and vendor
Field extraction
LLM with a schema
Business rules
Checked against ERP data
Confidence thresholds
Review queue
ERP REST API
Each document is classified first, because the type decides which schema to extract. Fields come back as structured output with a confidence per field, then business rules check them against data the ERP already holds.
Only uncertain fields reach a reviewer, highlighted on the source document. Their corrections are stored and used to tune prompts and thresholds, so the review queue shrinks over time rather than staying a manual step.
Build something like this
Tell me about the process you want to improve and the systems it touches. I'll come back with questions and a suggested approach.