Automatic data extraction, and the two different jobs people mean by it
Files arrive as they always have. We build the extraction in a business day. It costs nothing until it runs.
Automatic data extraction covers two unrelated jobs. One pulls data off web pages, which breaks when the markup changes. The other pulls fields out of files, where there is no markup to break, and the difficulty is that nothing on the page says which number is the total. This page is about the second.
Fabrx builds the API. Landing it in a given system is done by partners, so the only requirement is a system that accepts an API call.
What moves, and what it becomes
A PDF that was generated, not scanned
The text is already in the file, which makes it the easy case, right up to the point where the layout is a table drawn with positioned text and no table.
Fields, read directly
The easy case
A scan of a printed page
Skew, speckle and a staple shadow across the column you need. The text has to be recovered before anything can be read out of it.
Fields, after recovery
The common case
A photograph taken on a phone
Perspective, a shadow across one half, and a finger in the corner. Usually taken by somebody standing up, in a hurry.
Fields, after correction
The field case
A spreadsheet attached to an email
Already structured, and structured the sender's way. The columns mean what their author meant, which is rarely what your system expects.
Fields, remapped
The deceptive one
The body of the email itself
No attachment at all. The order, the reference and the quantities are in the message, under a signature block and three levels of quoted reply.
Fields, from prose
The one people forget
A file inside a portal or an archive
Reachable only by logging in, or sitting in a zip with forty others and a naming convention somebody abandoned in March.
Fields, once it reaches us
The delivery problem
What happens to one document
Your documents
File arrives. Your documents
Email, a portal, a scanner or an upload. Getting it to us is your side.
T+0
Fabrx
Identify what it is. Fabrx
A mixed inbox is sorted before anything is read out of it.
If a file type the extraction was not built for
Hold for review. Fabrx. Needs a person.
A person sees the file and what stopped it.
needs a person
Then it rejoins the path above.
Read the fields. Fabrx
The fields that kind of file carries, wherever they sit in it.
Check what can be checked. Fabrx
What the file can confirm on its own, and nothing it cannot.
If a field left open, or a value that disagrees with itself
Hold for review. Fabrx. Needs a person.
A person sees the file and what stopped it.
needs a person
Then it rejoins the path above.
Send the payload. Fabrx
A structured record at the endpoint you nominate.
Where we hand over
The handoff is a JSON payload to an HTTPS endpoint you own, or a webhook we call on your side. We hold no credentials to your systems.
Your system
Lands in your system. Your system
Whatever happens next is the workflow you already run.
What we do, and what stays yours
Fabrx builds and runs
- Reading the file, whatever container it arrives in
- Identifying what it is, so a mixed inbox does not need sorting first
- Pulling the fields that kind of file carries
- Checking what the file can confirm on its own
- Holding anything it cannot settle, with the file attached
- Delivering a structured record to an endpoint you nominate
- Adding fields or file types when what you need changes
You keep
- Getting the file to us, including anything behind a login
- What the data means, and the rule that acts on it
- The system it lands in, and who is allowed to see it
- Any check that needs information the file does not carry
- Somebody to look at the output and tell us whether it is right
When a document is wrong
A pipeline with no failure path is a diagram of a good day. These are the ones that happen.
The scan is unreadable
The file stops before extraction and is never guessed at.
Who sees it. Whoever sent it, by reply, with the page that failed.
- Rescanned and reprocessed
- Keyed by hand, the way it is done today
The file is locked or damaged
A password-protected PDF or a truncated download cannot be opened, so nothing is attempted on it.
Who sees it. Whoever sent it, with which of the two it was.
- Resent unlocked
- Password supplied once, for that sender
It is a file type the extraction was not built for
It is identified as something else, set aside, and we are told. Nothing is sent to your endpoint.
Who sees it. Us first, then you, usually the same day.
- Added as its own type, usually in one business day
- Confirmed as out of scope
A field we agreed on is genuinely not in the file
The field is returned empty and marked as absent. An empty field and a field we could not read are different states and they stay different.
Who sees it. Your reviewer, in the record.
- Chased with the sender
- Accepted as absent, which is your call
Automatic data extraction, four ways to get it
The first column is here to be ruled out. It is the right tool for a different job.
What it is actually for
Write a scraper
Pages built for people to read, reached by a URL.
Plain OCR
Turning a picture of text into text.
A platform you configure
Files, once you have told it what to look for.
Fabrx
Files, with the reading built for yours.
What it hands back
Write a scraper
Whatever you selected, as long as the markup holds.
Plain OCR
Text in reading order, with no idea which value is which.
A platform you configure
The fields you configured, inside their system.
Fabrx
A structured record, at your endpoint, in the shape you asked for.
What breaks it
Write a scraper
A markup change, a rate limit, or a login.
Plain OCR
Nothing, because it was never interpreting anything.
A platform you configure
A layout the configuration did not anticipate.
Fabrx
A file type nobody mentioned, which comes back as an exception.
Who maintains it
Write a scraper
You do, every time the site changes.
Plain OCR
Nobody, because there is nothing configured to maintain.
A platform you configure
You do, in their builder.
Fabrx
We do. Adding a field is a request, not a procurement cycle.
What it costs to find out
Write a scraper
An afternoon, and the maintenance starts immediately.
Plain OCR
Low, and the work of interpreting the text is still ahead of you.
A platform you configure
A licence and an implementation, before you know it reads yours.
Fabrx
Nothing. The build is free and you pay per document once it runs.
Someone has run this
They built solutions for any situation encountered along the way. For us, this type of partnership matters.
- 300+
- Field agents using it every working day
- 4,000+
- Documents processed every month
- 4 yrs
- In production, on the same workflow
Blitz runs identity documents into contracts, photographed in the field on phones. These figures describe that pipeline and that kind of file.
What a pilot measures
Measured on your documents
Field-level accuracy
Every field on every sample file, scored against your correction.
Sample size
How many of your real files we ran. You choose which.
Spread of the sample
How many containers your sample covers, scans and photographs included, so the figure is not measured only on the clean ones.
Exception rate
The share that stopped for a person, and what stopped them.
Build time
Calendar time from your samples reaching us to the pipeline running.
What we ask you for
How many minutes does one file take today, start to finish?
It is the only honest denominator for any claim about time saved.
How many do you process in a month?
It sets the pricing band, because the rate depends on volume.
How do they reach you, and in how many different ways?
The container is often harder than the document. A portal login or a zip of forty changes what gets built before anything is read.
What proportion are scans or photographs?
It is the difference between an easy build and a real one.
Which fields does somebody actually use afterwards?
Asking for every field on the page is the default. The ones nobody reads cost the same to build as the ones that matter.
What it costs
- 01
We build it
We build the extraction, usually within one business day of receiving your sample documents.
- 02
You check it
It costs you nothing until it runs on your real documents. If the accuracy is not good enough, nothing goes into production and nothing is charged.
- 03
Then you pay
A prepaid package priced per document, valid for a calendar year. The rate depends on the document type and your monthly volume.
- No subscription
- No per-seat charge
- No setup fee at any volume
Questions about automatic data extraction
What is automatic data extraction?
Pulling specific values out of a source and handing them over as structured data, without a person reading the source first. The source can be a web page, a database, or a file. Which one it is decides almost everything about how it is done.
What is the difference between data extraction and web scraping?
Scraping is extraction from pages built for people to read. It works by following the markup, so it breaks when the markup changes, and it has to live with logins, rate limits and what the site permits. Extracting from a file has none of those problems and one that scraping does not have: there is no markup at all, so nothing in the file says which number is the total.
Can data be extracted automatically from a PDF?
Yes. A PDF that was generated rather than scanned already carries its text, which makes it the easier case. The harder part is that a PDF has layout and no structure, so a table can be a set of positioned words that only look like columns.
What about a scan or a photograph of a document?
Both are ordinary input for printed pages, including the creased, skewed and badly lit ones. Where it stops is resolution low enough that a person would struggle too, and those come back as unreadable instead of guessed at.
What happens to a file it cannot read?
It stops and is reported. A locked or damaged file is returned with which of the two it was. An unfamiliar type is set aside and we are told, usually the same day. Nothing is sent to your endpoint on a guess.
Does it need a template for each layout?
No. The extraction reads fields by what they are, so it does not depend on a sender keeping one layout. Template based tools are tied to where each field sits, which is why a redesign breaks them.
Related
- automatic document processing
The same work described by its five steps, from capture through to the system it lands in.
- invoice processing
What it looks like on the document type most teams automate first.
- how the build works
The call, the sample files, and what happens in the business day between them.
Send us twenty files
We build the extraction and show you what it got right, field by field, before you pay anything.