Automatic document processing, step by step, and the one that decides it
Documents arrive as they always have. We build the extraction in a business day. It costs nothing until it runs.
Automatic document processing turns documents into data without a person reading them. Five steps: the document arrives, it is identified, the fields are read, the result is checked, and it lands in a system. Reading the fields is the step that decides whether the other four were worth building.
Fabrx builds the API. Landing it in a given system is done by partners, so the only requirement is a system that accepts an API call.
What moves, and what it becomes
Invoices and credit notes
Line items in a table that changes shape per supplier, and a credit note that has to find the invoice it reverses.
A payable, with its lines
Accounts payable
Certificates and policies
Coverage sits in a limits table and the endorsements that qualify it are separate pages behind the certificate.
A coverage record per policy line
Vendor compliance
Identity documents
Short, dense, and unforgiving. One mistyped digit in a document number makes a contract legally awkward.
A verified party on a record
Onboarding and contracts
Contracts and agreements
The fields that matter are dates, parties, renewal terms and notice periods, and they are written in sentences instead of boxes.
A contract record with its dates
Legal and procurement
Delivery notes and receipts
Photographed at the point of delivery, creased, and often the only record that a short shipment happened.
A receipt against an order
Inventory
Forms and applications
Boxes that are reliable when ticked and ambiguous when they are not, next to free text that carries the actual answer.
A structured submission
Operations
What happens to one document
Your documents
Document arrives. Your documents
Email, a portal, a scanner or an upload. Nothing changes about how.
T+0
Fabrx
Identify what it is. Fabrx
A mixed inbox is sorted before anything is read out of it.
If a type the extraction was not built for
Hold for review. Fabrx. Needs a person.
A person sees the document and what stopped it.
needs a person
Then it rejoins the path above.
Read the fields. Fabrx
The fields that type of document carries, wherever they sit on the page.
Check what can be checked. Fabrx
Totals against lines, dates against each other, a field against its own format.
If a field left open, or a number that disagrees with itself
Hold for review. Fabrx. Needs a person.
A person sees the document and what stopped it.
needs a person
Then it rejoins the path above.
Send the payload. Fabrx
A structured record at the endpoint you nominate.
Where we hand over
The handoff is a JSON payload to an HTTPS endpoint you own, or a webhook we call on your side. We hold no credentials to your systems.
Your system
Lands in your system. Your system
Whatever happens next is the workflow you already run.
What we do, and what stays yours
Fabrx builds and runs
- Identifying what each document is, so a mixed inbox does not need sorting first
- Reading the document, whatever shape and layout it arrives in
- Pulling the fields that type of document carries
- Checking what the document can confirm on its own
- Holding anything it cannot settle, with the document attached
- Delivering a structured record to an endpoint you nominate
- Adding fields or document types when what you need changes
You keep
- What the data means, and the rule that acts on it
- The system it lands in, and who is allowed to see it
- The workflow it triggers, which does not move
- Any check that needs information the document does not carry
- Somebody to look at the output and tell us whether it is right
When a document is wrong
A pipeline with no failure path is a diagram of a good day. These are the ones that happen.
The scan is unreadable
The document stops before extraction and is never guessed at.
Who sees it. Whoever sent it, by reply, with the page that failed.
- Rescanned and reprocessed
- Keyed by hand, the way it is done today
It is a document type the extraction was not built for
It is identified as something else, set aside, and we are told. Nothing is sent to your endpoint.
Who sees it. Us first, then you, usually the same day.
- Added as its own document type, usually in one business day
- Confirmed as out of scope
A field we agreed on is genuinely not on the document
The field is returned empty and marked as absent. An empty field and a field we could not read are different states and they stay different.
Who sees it. Your reviewer, in the record.
- Chased with the sender
- Accepted as absent, which is your call
The document contradicts itself
Both numbers are returned and the document waits. We never pick the one that looks more likely.
Who sees it. Your reviewer, with both values shown.
- Corrected and released
- Sent back to the sender
What automatic document processing looks like, four ways
The same document, four approaches, and what each one actually hands back.
What it hands back
Plain OCR
Text, in reading order, with no idea which part is the invoice number.
Template capture
The fields you mapped, from the positions you mapped them to.
A platform you configure
The fields you configured, inside their system.
Fabrx
A structured record, at your endpoint, in the shape you asked for.
Who defines the fields
Plain OCR
Nobody. There are no fields, only text.
Template capture
You do, per document layout, and you maintain each one.
A platform you configure
You do, in their builder, and you maintain it.
Fabrx
You tell us on the call, and we build to it.
A document it has not seen before
Plain OCR
Read the same way as any other. Nothing identifies what it is.
Template capture
No template matches, so nothing useful comes back.
A platform you configure
Depends on the engine. Configuring for it is your job.
Fabrx
Identified as unfamiliar, set aside, and reported the same day.
What it costs to find out
Plain OCR
Low. The output still needs interpreting, so you learn little from it.
Template capture
A template build per layout, before you know it holds.
A platform you configure
A licence and an implementation, before you know it reads yours.
Fabrx
Nothing. The build is free and you pay per document once it runs.
Where the output lands
Plain OCR
A text file somebody still has to read.
Template capture
An export you move yourself.
A platform you configure
Their platform, then a connector into yours.
Fabrx
The system you already use, through an endpoint you own.
Someone has run this
They built solutions for any situation encountered along the way. For us, this type of partnership matters.
- 300+
- Field agents using it every working day
- 4,000+
- Documents processed every month
- 4 yrs
- In production, on the same workflow
Blitz runs identity documents into contracts. These figures describe that pipeline and that document type, on their volumes.
What a pilot measures
Measured on your documents
Field-level accuracy
Every field on every sample document, scored against your correction.
Sample size
How many of your real documents we ran. You choose which.
Spread of the sample
How many document types and senders your sample covers, so the figure means something beyond one kind of paperwork.
Exception rate
The share that stopped for a person, and what stopped them.
Build time
Calendar time from your samples reaching us to the pipeline running.
What we ask you for
How many minutes does one document take today, start to finish?
It is the only honest denominator for any claim about time saved.
How many do you process in a month?
It sets the pricing band, because the rate depends on volume.
How many different document types arrive in the same inbox?
Sorting them is a step of its own, and the answer changes what gets built.
What happens today when one comes through wrong?
It tells us where the exception should go, and who is waiting on it.
Which fields does somebody actually use afterwards?
Asking for every field on the page is the default. The ones nobody reads afterwards cost the same to build as the ones that matter.
What it costs
- 01
We build it
We build the extraction, usually within one business day of receiving your sample documents.
- 02
You check it
It costs you nothing until it runs on your real documents. If the accuracy is not good enough, nothing goes into production and nothing is charged.
- 03
Then you pay
A prepaid package priced per document, valid for a calendar year. The rate depends on the document type and your monthly volume.
- No subscription
- No per-seat charge
- No setup fee at any volume
Questions about automatic document processing
What is automatic document processing?
It is turning documents into data without a person reading them. A document arrives, software identifies what it is, reads the fields that type carries, checks what it can, and sends a structured record to a system. People stay in the loop for the ones that stop.
What are the steps in automatic document processing?
Capture, classification, extraction, validation and integration. Capture is how the document reaches the system. Classification decides what it is. Extraction reads the fields. Validation checks what the document can confirm on its own. Integration puts the result where the work happens.
What is the difference between OCR and intelligent document processing?
OCR turns a picture of text into text. It gives you words in reading order and no idea which of them is the invoice number. Intelligent document processing is the layer that decides what the document is and which values belong to which fields. OCR is a component of it, and on its own it leaves the interpreting to you.
Does it work on scanned documents and photographs?
Yes, for printed documents. A phone photograph of a printed page is ordinary input, including the creased and badly lit ones. Resolution low enough that a person would struggle is where it stops, and those are returned as unreadable instead of guessed at.
What happens to a document it has not seen before?
It is identified as something unfamiliar and set aside, and we are told. Nothing is sent to your endpoint on a guess. Adding it as its own type is usually a business day, or we tell you it is out of scope.
How much of it can actually run without a person?
The reading and the checking can. What cannot is any decision that needs information the document does not carry, or judgement about what to do when two things disagree. If something claims to need nobody, ask what it does with the documents it gets wrong.
Related
- invoice processing
The same five steps on the document type most teams automate first, with the line items that make it hard.
- certificate of insurance processing
What it looks like when the output feeds a compliance decision instead of a payment.
- how the build works
The call, the sample documents, and what happens in the business day between them.
Send us twenty documents
We build the extraction and show you what it got right, field by field, before you pay anything.