Skip to content
Fundamentals

Data extraction companies, and the three industries that answer to the name

Three industries share the name. Only one reads the files you receive, and that is the one we are in.

See the pipeline

Data extraction companies fall into three groups that do not substitute for each other. Web data companies read public web pages at scale. Data pipeline companies move records between systems that already hold them. Document extraction companies read files you receive, where the data has no structure at all.

Fabrx builds the API. Landing it in a given system is done by partners, so the only requirement is a system that accepts an API call.

What moves, and what it becomes

  • Web data companies

    Apify, Oxylabs, Bright Data, ScraperAPI and others read public pages at scale, through the defences sites put up to stop them. If your data is on somebody else's website, this is the right group and nothing below will help.

    Pages, as records

    The largest group

  • Data pipeline companies

    Hevo, Fivetran and their kind move data that is already structured, between systems that already hold it. Nothing is being read here, only transported and reshaped.

    Records, moved

    Often in the same lists

  • Document extraction platforms

    Software you configure yourself. You describe the fields in their builder, test it against your documents, and maintain it as those documents change.

    Fields, once configured

    Document extraction

  • Outsourced document services

    People reading your documents, often with software behind them. You brief the fields once and the provider interprets the edge cases.

    Fields, by arrangement

    Document extraction

  • Document extraction builds

    The extraction is built for your documents by whoever sells it, and runs as an API behind the systems you already use. Fabrx is in this group.

    Fields, at your endpoint

    Document extraction

  • Directories and ranked lists

    These are not vendors. They rank the three groups together, which is how a shortlist ends up with a scraping API and a document parser on the same page.

    A shortlist worth checking

    How shortlists get built

What happens to one document

  1. Your documents

    1. File arrives. Your documents

      Email, a portal, a scanner or an upload. Getting it to us is your side.

      T+0

  2. Fabrx

    1. Identify what it is. Fabrx

      A mixed inbox is sorted before anything is read out of it.

    If a file type the extraction was not built for

    1. Hold for review. Fabrx. Needs a person.

      A person sees the file and what stopped it.

      needs a person

    Then it rejoins the path above.

    1. Read the fields. Fabrx

      The fields that kind of file carries, wherever they sit in it.

    1. Check what can be checked. Fabrx

      What the file can confirm on its own, and nothing it cannot.

    If a field left open, or a value that disagrees with itself

    1. Hold for review. Fabrx. Needs a person.

      A person sees the file and what stopped it.

      needs a person

    Then it rejoins the path above.

    1. Send the payload. Fabrx

      A structured record at the endpoint you nominate.

    Where we hand over

    The handoff is a JSON payload to an HTTPS endpoint you own, or a webhook we call on your side. We hold no credentials to your systems.

  3. Your system

    1. Lands in your system. Your system

      Whatever happens next is the workflow you already run.

The document extraction path, which is the only one of the three we can describe from the inside. A file arrives, Fabrx identifies it, reads the fields, checks what can be checked, and a structured record reaches your endpoint.
Where we stop

What we do, and what stays yours

Fabrx builds and runs

  • Reading the file, whatever container it arrives in
  • Identifying what it is, so a mixed inbox does not need sorting first
  • Pulling the fields that kind of file carries
  • Checking what the file can confirm on its own
  • Holding anything it cannot settle, with the file attached
  • Delivering a structured record to an endpoint you nominate

You keep

  • Anything on a public website, which is work for a web data company
  • Anything already in a database, which is work for a pipeline company
  • Getting the file to us, including anything behind a login
  • What the data means, and the system it lands in
  • Somebody to look at the output and tell us whether it is right

When a document is wrong

A pipeline with no failure path is a diagram of a good day. These are the ones that happen.

The scan is unreadable

The file stops before extraction and is never guessed at.

Who sees it. Whoever sent it, by reply, with the page that failed.

  • Rescanned and reprocessed
  • Keyed by hand, the way it is done today
The file is locked or damaged

A password-protected PDF or a truncated download cannot be opened, so nothing is attempted on it.

Who sees it. Whoever sent it, with which of the two it was.

  • Resent unlocked
  • Password supplied once, for that sender
It is a file type the extraction was not built for

It is identified as something else, set aside, and we are told. Nothing is sent to your endpoint.

Who sees it. Us first, then you, usually the same day.

  • Added as its own type, usually in one business day
  • Confirmed as out of scope
The data turns out to be on a website

We say so on the call and name the kind of company that does it. There is no version of this where we take the work and subcontract it.

Who sees it. You, before anything is built.

  • Sent to a web data company
  • Scoped down to the files you do receive

Data extraction companies, and what each kind actually does

The three ways to buy document extraction, against keeping it on a person.

  • What you are actually buying

    By hand

    Nothing. The work stays where it is.

    A platform you configure

    Software, and the job of configuring it.

    An outsourced service

    Capacity, and somebody else making the judgement calls.

    Fabrx

    An extraction built for your files, running as an API.

  • Who defines the fields

    By hand

    Whoever is keying, differently each time.

    A platform you configure

    You do, in their builder, and you maintain it.

    An outsourced service

    You brief it once and the provider interprets it.

    Fabrx

    You tell us on the call, and we build to it.

  • What it costs to find out whether it works

    By hand

    No new spend. The cost is hours your team is already paid for.

    A platform you configure

    A licence and an implementation, before you know it reads yours.

    An outsourced service

    Onboarding, then a rate per document or per seat.

    Fabrx

    Nothing. The build is free and you pay per document once it runs.

  • Who maintains it

    By hand

    Nobody, because there is nothing built to maintain.

    A platform you configure

    You do, in their builder, whenever your documents change.

    An outsourced service

    The provider, at the provider's pace.

    Fabrx

    We do. Adding a field is a request, not a procurement cycle.

  • Where the output lands

    By hand

    Retyped into the system you already use.

    A platform you configure

    Their platform, then a connector into yours.

    An outsourced service

    The provider's portal, or a file they send you.

    Fabrx

    A record at an endpoint you own, inside the system you already use.

Someone has run this

They built solutions for any situation encountered along the way. For us, this type of partnership matters.

Operations lead, Operations, BlitzOctober 2026
300+
Field agents using it every working day
4,000+
Documents processed every month
4 yrs
In production, on the same workflow

Blitz runs identity documents into contracts, photographed in the field on phones. These figures describe that pipeline and that kind of file.

Before you pay

What a pilot measures

Measured on your documents

  • Field-level accuracy

    Every field on every sample file, scored against your correction.

  • Sample size

    How many of your real files we ran. You choose which.

  • Spread of the sample

    How many senders and containers your sample covers, so the figure is not measured only on the clean ones.

  • Exception rate

    The share that stopped for a person, and what stopped them.

  • Build time

    Calendar time from your samples reaching us to the pipeline running.

What we ask you for

  • Where does the data live right now?

    On a website, in a database, or in files you receive. The answer decides which of the three industries you are shopping in, and it is worth settling before anyone quotes you.

  • How do the files reach you, and in how many different ways?

    The container is often harder than the document. A portal login or a zip of forty changes what gets built before anything is read.

  • How many do you process in a month?

    It sets the pricing band, because the rate depends on volume.

  • How many minutes does one take today, start to finish?

    It is the only honest denominator for any claim about time saved.

  • Which fields does somebody actually use afterwards?

    Asking for every field on the page is the default. The ones nobody reads cost the same to build as the ones that matter.

What it costs

  1. 01

    We build it

    We build the extraction, usually within one business day of receiving your sample documents.

  2. 02

    You check it

    It costs you nothing until it runs on your real documents. If the accuracy is not good enough, nothing goes into production and nothing is charged.

  3. 03

    Then you pay

    A prepaid package priced per document, valid for a calendar year. The rate depends on the document type and your monthly volume.

  • No subscription
  • No per-seat charge
  • No setup fee at any volume

Questions about data extraction companies

What do data extraction companies actually do?

Three different things, depending which kind you have found. Web data companies read public pages at scale. Pipeline companies move structured records between systems. Document extraction companies read files you receive and return the fields. The ranked lists on this subject mix all three.

What is the difference between a data extraction company and a web scraping company?

A scraping company reads pages built for people, which means following the markup, working around rate limits and logins, and rebuilding when a site changes. Reading a file has none of those problems and one they do not have: a file carries no markup, so nothing in it says which number is the total. The two solve different problems and neither can do the other job.

How much do data extraction services cost?

It depends which group you are buying from, and they charge in different shapes. Web data companies usually price by request or by volume of pages. Platforms charge a licence plus the implementation. Outsourced services charge per document or per seat. Fabrx builds the extraction for nothing and charges per document once it runs.

Should you buy a platform, outsource it, or have it built?

A platform suits a team with somebody who will own the configuration and keep owning it. Outsourcing suits volume that moves and judgement calls you are content to delegate. A build suits a team that wants the fields in the system they already use and does not want a new interface for anyone to log into.

How long does it take to get working?

A platform takes as long as the configuration and the testing take, which is yours to schedule. Outsourcing takes an onboarding. Fabrx usually builds the extraction within one business day of receiving sample documents, and it costs nothing until it runs on your real ones.

What should you ask a vendor before signing?

Where your data currently lives, and whether they do that kind. What a quoted accuracy figure was measured on, since a number from header fields is a much easier measurement than one that includes line items. What happens to a document they cannot read. And who maintains it when your documents change.

Related

Send us twenty files

We build the extraction and show you what it got right, field by field, before you pay anything.