Data extraction companies, and the three industries that answer to the name
Three industries share the name. Only one reads the files you receive, and that is the one we are in.
Data extraction companies fall into three groups that do not substitute for each other. Web data companies read public web pages at scale. Data pipeline companies move records between systems that already hold them. Document extraction companies read files you receive, where the data has no structure at all.
Fabrx builds the API. Landing it in a given system is done by partners, so the only requirement is a system that accepts an API call.
What moves, and what it becomes
Web data companies
Apify, Oxylabs, Bright Data, ScraperAPI and others read public pages at scale, through the defences sites put up to stop them. If your data is on somebody else's website, this is the right group and nothing below will help.
Pages, as records
The largest group
Data pipeline companies
Hevo, Fivetran and their kind move data that is already structured, between systems that already hold it. Nothing is being read here, only transported and reshaped.
Records, moved
Often in the same lists
Document extraction platforms
Software you configure yourself. You describe the fields in their builder, test it against your documents, and maintain it as those documents change.
Fields, once configured
Document extraction
Outsourced document services
People reading your documents, often with software behind them. You brief the fields once and the provider interprets the edge cases.
Fields, by arrangement
Document extraction
Document extraction builds
The extraction is built for your documents by whoever sells it, and runs as an API behind the systems you already use. Fabrx is in this group.
Fields, at your endpoint
Document extraction
Directories and ranked lists
These are not vendors. They rank the three groups together, which is how a shortlist ends up with a scraping API and a document parser on the same page.
A shortlist worth checking
How shortlists get built
What happens to one document
Your documents
File arrives. Your documents
Email, a portal, a scanner or an upload. Getting it to us is your side.
T+0
Fabrx
Identify what it is. Fabrx
A mixed inbox is sorted before anything is read out of it.
If a file type the extraction was not built for
Hold for review. Fabrx. Needs a person.
A person sees the file and what stopped it.
needs a person
Then it rejoins the path above.
Read the fields. Fabrx
The fields that kind of file carries, wherever they sit in it.
Check what can be checked. Fabrx
What the file can confirm on its own, and nothing it cannot.
If a field left open, or a value that disagrees with itself
Hold for review. Fabrx. Needs a person.
A person sees the file and what stopped it.
needs a person
Then it rejoins the path above.
Send the payload. Fabrx
A structured record at the endpoint you nominate.
Where we hand over
The handoff is a JSON payload to an HTTPS endpoint you own, or a webhook we call on your side. We hold no credentials to your systems.
Your system
Lands in your system. Your system
Whatever happens next is the workflow you already run.
What we do, and what stays yours
Fabrx builds and runs
- Reading the file, whatever container it arrives in
- Identifying what it is, so a mixed inbox does not need sorting first
- Pulling the fields that kind of file carries
- Checking what the file can confirm on its own
- Holding anything it cannot settle, with the file attached
- Delivering a structured record to an endpoint you nominate
You keep
- Anything on a public website, which is work for a web data company
- Anything already in a database, which is work for a pipeline company
- Getting the file to us, including anything behind a login
- What the data means, and the system it lands in
- Somebody to look at the output and tell us whether it is right
When a document is wrong
A pipeline with no failure path is a diagram of a good day. These are the ones that happen.
The scan is unreadable
The file stops before extraction and is never guessed at.
Who sees it. Whoever sent it, by reply, with the page that failed.
- Rescanned and reprocessed
- Keyed by hand, the way it is done today
The file is locked or damaged
A password-protected PDF or a truncated download cannot be opened, so nothing is attempted on it.
Who sees it. Whoever sent it, with which of the two it was.
- Resent unlocked
- Password supplied once, for that sender
It is a file type the extraction was not built for
It is identified as something else, set aside, and we are told. Nothing is sent to your endpoint.
Who sees it. Us first, then you, usually the same day.
- Added as its own type, usually in one business day
- Confirmed as out of scope
The data turns out to be on a website
We say so on the call and name the kind of company that does it. There is no version of this where we take the work and subcontract it.
Who sees it. You, before anything is built.
- Sent to a web data company
- Scoped down to the files you do receive
Data extraction companies, and what each kind actually does
The three ways to buy document extraction, against keeping it on a person.
What you are actually buying
By hand
Nothing. The work stays where it is.
A platform you configure
Software, and the job of configuring it.
An outsourced service
Capacity, and somebody else making the judgement calls.
Fabrx
An extraction built for your files, running as an API.
Who defines the fields
By hand
Whoever is keying, differently each time.
A platform you configure
You do, in their builder, and you maintain it.
An outsourced service
You brief it once and the provider interprets it.
Fabrx
You tell us on the call, and we build to it.
What it costs to find out whether it works
By hand
No new spend. The cost is hours your team is already paid for.
A platform you configure
A licence and an implementation, before you know it reads yours.
An outsourced service
Onboarding, then a rate per document or per seat.
Fabrx
Nothing. The build is free and you pay per document once it runs.
Who maintains it
By hand
Nobody, because there is nothing built to maintain.
A platform you configure
You do, in their builder, whenever your documents change.
An outsourced service
The provider, at the provider's pace.
Fabrx
We do. Adding a field is a request, not a procurement cycle.
Where the output lands
By hand
Retyped into the system you already use.
A platform you configure
Their platform, then a connector into yours.
An outsourced service
The provider's portal, or a file they send you.
Fabrx
A record at an endpoint you own, inside the system you already use.
Someone has run this
They built solutions for any situation encountered along the way. For us, this type of partnership matters.
- 300+
- Field agents using it every working day
- 4,000+
- Documents processed every month
- 4 yrs
- In production, on the same workflow
Blitz runs identity documents into contracts, photographed in the field on phones. These figures describe that pipeline and that kind of file.
What a pilot measures
Measured on your documents
Field-level accuracy
Every field on every sample file, scored against your correction.
Sample size
How many of your real files we ran. You choose which.
Spread of the sample
How many senders and containers your sample covers, so the figure is not measured only on the clean ones.
Exception rate
The share that stopped for a person, and what stopped them.
Build time
Calendar time from your samples reaching us to the pipeline running.
What we ask you for
Where does the data live right now?
On a website, in a database, or in files you receive. The answer decides which of the three industries you are shopping in, and it is worth settling before anyone quotes you.
How do the files reach you, and in how many different ways?
The container is often harder than the document. A portal login or a zip of forty changes what gets built before anything is read.
How many do you process in a month?
It sets the pricing band, because the rate depends on volume.
How many minutes does one take today, start to finish?
It is the only honest denominator for any claim about time saved.
Which fields does somebody actually use afterwards?
Asking for every field on the page is the default. The ones nobody reads cost the same to build as the ones that matter.
What it costs
- 01
We build it
We build the extraction, usually within one business day of receiving your sample documents.
- 02
You check it
It costs you nothing until it runs on your real documents. If the accuracy is not good enough, nothing goes into production and nothing is charged.
- 03
Then you pay
A prepaid package priced per document, valid for a calendar year. The rate depends on the document type and your monthly volume.
- No subscription
- No per-seat charge
- No setup fee at any volume
Questions about data extraction companies
What do data extraction companies actually do?
Three different things, depending which kind you have found. Web data companies read public pages at scale. Pipeline companies move structured records between systems. Document extraction companies read files you receive and return the fields. The ranked lists on this subject mix all three.
What is the difference between a data extraction company and a web scraping company?
A scraping company reads pages built for people, which means following the markup, working around rate limits and logins, and rebuilding when a site changes. Reading a file has none of those problems and one they do not have: a file carries no markup, so nothing in it says which number is the total. The two solve different problems and neither can do the other job.
How much do data extraction services cost?
It depends which group you are buying from, and they charge in different shapes. Web data companies usually price by request or by volume of pages. Platforms charge a licence plus the implementation. Outsourced services charge per document or per seat. Fabrx builds the extraction for nothing and charges per document once it runs.
Should you buy a platform, outsource it, or have it built?
A platform suits a team with somebody who will own the configuration and keep owning it. Outsourcing suits volume that moves and judgement calls you are content to delegate. A build suits a team that wants the fields in the system they already use and does not want a new interface for anyone to log into.
How long does it take to get working?
A platform takes as long as the configuration and the testing take, which is yours to schedule. Outsourcing takes an onboarding. Fabrx usually builds the extraction within one business day of receiving sample documents, and it costs nothing until it runs on your real ones.
What should you ask a vendor before signing?
Where your data currently lives, and whether they do that kind. What a quoted accuracy figure was measured on, since a number from header fields is a much easier measurement than one that includes line items. What happens to a document they cannot read. And who maintains it when your documents change.
Related
- automatic data extraction
How the document side actually works, once you have established that your data is in files.
- automatic document processing
The same work described by its five steps, capture through to integration.
- how the build works
The call, the sample files, and what happens in the business day between them.
Send us twenty files
We build the extraction and show you what it got right, field by field, before you pay anything.