Intelligent document extraction: from inbox to structured record
Thousands of documents across email, chat and portals, read, identified and filed without anyone re-keying them - and anything unsure flagged for a person.
By the numbers
One pattern, proven in two regulated industries
Financial services figures describe a 15,000-client tax practice. Healthcare figures describe a multi-site medical practice.
The problem
Skilled people spend the busiest weeks opening attachments
A regulated, seasonal business receives thousands of documents across email, chat and portals. Each one is opened, identified, filed and re-keyed by someone qualified to do far more valuable work - and it happens at exactly the time of year when there is least capacity to spare.
The documents are never tidy. They arrive locked behind passwords, photographed on phones, bundled together, and in dozens of near-identical formats. A single missed or misfiled certificate can hold up a whole client's submission.
The pattern
Ingest anywhere, normalise once
- 01
Ingest anywhere
Email attachments, chat uploads, a watched folder and a client portal all drop documents into the same place, filed against the right person and year.
- 02
Unlock
Password-protected PDFs are opened using details the business already holds, and the encrypted copy is replaced by a readable one.
- 03
Read
OCR turns pages and photographs into text, with tables kept intact, so the model sees what a person would.
- 04
Classify
The document is identified against a taxonomy the business owns, with tax year, a one-line note and a confidence score.
- 05
Extract
Type-specific fields are pulled out against a strict schema, so the output is a record, not a paragraph.
- 06
Write and link
One row per document, linked to the case it belongs to, with the checklist item ticked off automatically.
- 07
Route to a person
Anything unrecognised or uncertain is created as a "needs review" item and the owner is notified.
Proof point one
Financial services: a 15,000-client tax practice
Arrives by
One pipeline
Lands as
Income certificates get a second pass that returns 47 fields mapped to the tax authority's own source codes, and assessments return their outcome - refund, payable or nil. The extracted figures feed revenue data and client segmentation, for example flagging commission earners.
The approach has been through several generations: a custom form-recognition model for a single certificate type, then general OCR with a language model across all 34 types, then structured-output models for chat uploads. If the model cannot place a document, it is created as an "Unclassified / Needs Review" item and the consultant is notified in the app.
Proof point two
Healthcare: choosing the model on evidence
A multi-site pain-management practice needed structured data out of two legacy clinical systems holding roughly 65,900 records. Rather than assume, we graded candidate models against 420 hand-checked cells.
Source
Before it leaves
Extract
Checked
The bake-off - graded accuracy, 420 cells
Zero schema violations
Tied on accuracy, 2.3x the cost
Rejected for schema violations and hallucinated values.
Graded accuracy for the chosen mid-tier model, with zero schema violations.
Cheaper than the top-tier model, which tied it within a point on accuracy.
Total model cost to extract the full set of records.
Alongside this, scan reports are read into structured fields in production, with a person validating everything OCR extracts, and an early prototype reads historic scan reports into a surgical category with a per-row human-review flag, pending clinical validation.
Design decisions
What makes it hold up in production
Rules before the model
Precedence rules decide the ambiguous cases up front: a return that contains a certificate is the return, not the certificate. The model handles the reading, not the policy.
Taxonomy lives in data
The allowed document types are read from a table at run time, so adding a type is a row, not a prompt rewrite, and the model can only answer with a type that exists.
Minimum signals, not vibes
A document that matches too few identifying signals scores low by rule, which keeps confidence honest instead of flattering.
Explicit needs-review rows
A document the system cannot place is never dropped or guessed. It becomes a visible item, with a suggested type, flagged for a person.
Redact before extracting
In healthcare, identifiers are removed before any data leaves the environment, and validation by a person is mandatory for anything read by OCR.
Cost per document, tracked
Each flow logs OCR cost per page and model cost per call, so the economics of a document type are known rather than assumed.
The most common failure was never the model. It was the environment around it: a wrong folder path, a flow quietly switched off after a deployment, an import that changed a setting. So we monitor the pipeline as carefully as we tune it.
Where else it applies
The same pipeline, a different folder of paper
Supplier invoices
Read, matched to the purchase order and queued for accounts payable, with exceptions flagged.
Insurance claim bundles
Split a mixed bundle into its parts, classify each, and pull the claim fields.
KYC and onboarding packs
Identify each proof document, extract the details and tick the compliance checklist.
Leases and contracts
Abstract dates, parties and obligations into fields the business can search and report on.
Delivery notes and proof of delivery
Read the signed paperwork and close the job against the order automatically.
HR onboarding documents
File identity, qualification and bank documents against the new starter and chase what is missing.
Bring us a folder of your worst documents
We will show you what the pipeline makes of them.