Intelligent document extraction: from inbox to structured record

    Thousands of documents across email, chat and portals, read, identified and filed without anyone re-keying them - and anything unsure flagged for a person.

    Financial Services
    Healthcare
    OCR + LLM extraction

    By the numbers

    One pattern, proven in two regulated industries

    34
    Document types told apart by one classifier (financial services)
    4
    Intake channels feeding the same pipeline
    47
    Fields read from a single tax certificate
    99.3%
    Graded accuracy of the model chosen for clinical records (healthcare)
    ~$216
    To extract roughly 65,900 legacy records

    Financial services figures describe a 15,000-client tax practice. Healthcare figures describe a multi-site medical practice.

    The problem

    Skilled people spend the busiest weeks opening attachments

    A regulated, seasonal business receives thousands of documents across email, chat and portals. Each one is opened, identified, filed and re-keyed by someone qualified to do far more valuable work - and it happens at exactly the time of year when there is least capacity to spare.

    The documents are never tidy. They arrive locked behind passwords, photographed on phones, bundled together, and in dozens of near-identical formats. A single missed or misfiled certificate can hold up a whole client's submission.

    The pattern

    Ingest anywhere, normalise once

    1. 01

      Ingest anywhere

      Email attachments, chat uploads, a watched folder and a client portal all drop documents into the same place, filed against the right person and year.

    2. 02

      Unlock

      Password-protected PDFs are opened using details the business already holds, and the encrypted copy is replaced by a readable one.

    3. 03

      Read

      OCR turns pages and photographs into text, with tables kept intact, so the model sees what a person would.

    4. 04

      Classify

      The document is identified against a taxonomy the business owns, with tax year, a one-line note and a confidence score.

    5. 05

      Extract

      Type-specific fields are pulled out against a strict schema, so the output is a record, not a paragraph.

    6. 06

      Write and link

      One row per document, linked to the case it belongs to, with the checklist item ticked off automatically.

    7. 07

      Route to a person

      Anything unrecognised or uncertain is created as a "needs review" item and the owner is notified.

    Proof point one

    Financial services: a 15,000-client tax practice

    Arrives by

    Email attachments
    WhatsApp uploads
    Watched folder
    Client portal

    One pipeline

    Unlock protected PDFs
    OCR the pages
    Classify against 34 types
    Extract up to 47 fields

    Lands as

    One case per client per year
    Checklist item ticked
    Revenue and segment flags
    Unsure? Needs review

    Income certificates get a second pass that returns 47 fields mapped to the tax authority's own source codes, and assessments return their outcome - refund, payable or nil. The extracted figures feed revenue data and client segmentation, for example flagging commission earners.

    The approach has been through several generations: a custom form-recognition model for a single certificate type, then general OCR with a language model across all 34 types, then structured-output models for chat uploads. If the model cannot place a document, it is created as an "Unclassified / Needs Review" item and the consultant is notified in the app.

    Proof point two

    Healthcare: choosing the model on evidence

    A multi-site pain-management practice needed structured data out of two legacy clinical systems holding roughly 65,900 records. Rather than assume, we graded candidate models against 420 hand-checked cells.

    Source

    ~65,900 legacy records

    Before it leaves

    Identifiers redacted

    Extract

    Strict schema, graded model

    Checked

    A person validates OCR output

    The bake-off - graded accuracy, 420 cells

    Mid-tier model (chosen)99.3%

    Zero schema violations

    Top-tier model99%

    Tied on accuracy, 2.3x the cost

    Small Claude model - rejected Small GPT model - rejected

    Rejected for schema violations and hallucinated values.

    99.3%

    Graded accuracy for the chosen mid-tier model, with zero schema violations.

    2.3x

    Cheaper than the top-tier model, which tied it within a point on accuracy.

    ~$216

    Total model cost to extract the full set of records.

    Alongside this, scan reports are read into structured fields in production, with a person validating everything OCR extracts, and an early prototype reads historic scan reports into a surgical category with a per-row human-review flag, pending clinical validation.

    Design decisions

    What makes it hold up in production

    Rules before the model

    Precedence rules decide the ambiguous cases up front: a return that contains a certificate is the return, not the certificate. The model handles the reading, not the policy.

    Taxonomy lives in data

    The allowed document types are read from a table at run time, so adding a type is a row, not a prompt rewrite, and the model can only answer with a type that exists.

    Minimum signals, not vibes

    A document that matches too few identifying signals scores low by rule, which keeps confidence honest instead of flattering.

    Explicit needs-review rows

    A document the system cannot place is never dropped or guessed. It becomes a visible item, with a suggested type, flagged for a person.

    Redact before extracting

    In healthcare, identifiers are removed before any data leaves the environment, and validation by a person is mandatory for anything read by OCR.

    Cost per document, tracked

    Each flow logs OCR cost per page and model cost per call, so the economics of a document type are known rather than assumed.

    The most common failure was never the model. It was the environment around it: a wrong folder path, a flow quietly switched off after a deployment, an import that changed a setting. So we monitor the pipeline as carefully as we tune it.

    Where else it applies

    The same pipeline, a different folder of paper

    Supplier invoices

    Read, matched to the purchase order and queued for accounts payable, with exceptions flagged.

    Insurance claim bundles

    Split a mixed bundle into its parts, classify each, and pull the claim fields.

    KYC and onboarding packs

    Identify each proof document, extract the details and tick the compliance checklist.

    Leases and contracts

    Abstract dates, parties and obligations into fields the business can search and report on.

    Delivery notes and proof of delivery

    Read the signed paperwork and close the job against the order automatically.

    HR onboarding documents

    File identity, qualification and bank documents against the new starter and chase what is missing.

    Bring us a folder of your worst documents

    We will show you what the pipeline makes of them.