Skip to content
Fundamentals

Automatic data extraction, and the two different jobs people mean by it

Files arrive as they always have. We build the extraction in a business day. It costs nothing until it runs.

See the pipeline

Automatic data extraction covers two unrelated jobs. One pulls data off web pages, which breaks when the markup changes. The other pulls fields out of files, where there is no markup to break, and the difficulty is that nothing on the page says which number is the total. This page is about the second.

Fabrx builds the API. Landing it in a given system is done by partners, so the only requirement is a system that accepts an API call.

What moves, and what it becomes

  • A PDF that was generated, not scanned

    The text is already in the file, which makes it the easy case, right up to the point where the layout is a table drawn with positioned text and no table.

    Fields, read directly

    The easy case

  • A scan of a printed page

    Skew, speckle and a staple shadow across the column you need. The text has to be recovered before anything can be read out of it.

    Fields, after recovery

    The common case

  • A photograph taken on a phone

    Perspective, a shadow across one half, and a finger in the corner. Usually taken by somebody standing up, in a hurry.

    Fields, after correction

    The field case

  • A spreadsheet attached to an email

    Already structured, and structured the sender's way. The columns mean what their author meant, which is rarely what your system expects.

    Fields, remapped

    The deceptive one

  • The body of the email itself

    No attachment at all. The order, the reference and the quantities are in the message, under a signature block and three levels of quoted reply.

    Fields, from prose

    The one people forget

  • A file inside a portal or an archive

    Reachable only by logging in, or sitting in a zip with forty others and a naming convention somebody abandoned in March.

    Fields, once it reaches us

    The delivery problem

What happens to one document

  1. Your documents

    1. File arrives. Your documents

      Email, a portal, a scanner or an upload. Getting it to us is your side.

      T+0

  2. Fabrx

    1. Identify what it is. Fabrx

      A mixed inbox is sorted before anything is read out of it.

    If a file type the extraction was not built for

    1. Hold for review. Fabrx. Needs a person.

      A person sees the file and what stopped it.

      needs a person

    Then it rejoins the path above.

    1. Read the fields. Fabrx

      The fields that kind of file carries, wherever they sit in it.

    1. Check what can be checked. Fabrx

      What the file can confirm on its own, and nothing it cannot.

    If a field left open, or a value that disagrees with itself

    1. Hold for review. Fabrx. Needs a person.

      A person sees the file and what stopped it.

      needs a person

    Then it rejoins the path above.

    1. Send the payload. Fabrx

      A structured record at the endpoint you nominate.

    Where we hand over

    The handoff is a JSON payload to an HTTPS endpoint you own, or a webhook we call on your side. We hold no credentials to your systems.

  3. Your system

    1. Lands in your system. Your system

      Whatever happens next is the workflow you already run.

A file arrives, Fabrx identifies what it is, reads the fields, checks what can be checked, and a structured record reaches your endpoint. Anything it cannot settle waits for a person.
Where we stop

What we do, and what stays yours

Fabrx builds and runs

  • Reading the file, whatever container it arrives in
  • Identifying what it is, so a mixed inbox does not need sorting first
  • Pulling the fields that kind of file carries
  • Checking what the file can confirm on its own
  • Holding anything it cannot settle, with the file attached
  • Delivering a structured record to an endpoint you nominate
  • Adding fields or file types when what you need changes

You keep

  • Getting the file to us, including anything behind a login
  • What the data means, and the rule that acts on it
  • The system it lands in, and who is allowed to see it
  • Any check that needs information the file does not carry
  • Somebody to look at the output and tell us whether it is right

When a document is wrong

A pipeline with no failure path is a diagram of a good day. These are the ones that happen.

The scan is unreadable

The file stops before extraction and is never guessed at.

Who sees it. Whoever sent it, by reply, with the page that failed.

  • Rescanned and reprocessed
  • Keyed by hand, the way it is done today
The file is locked or damaged

A password-protected PDF or a truncated download cannot be opened, so nothing is attempted on it.

Who sees it. Whoever sent it, with which of the two it was.

  • Resent unlocked
  • Password supplied once, for that sender
It is a file type the extraction was not built for

It is identified as something else, set aside, and we are told. Nothing is sent to your endpoint.

Who sees it. Us first, then you, usually the same day.

  • Added as its own type, usually in one business day
  • Confirmed as out of scope
A field we agreed on is genuinely not in the file

The field is returned empty and marked as absent. An empty field and a field we could not read are different states and they stay different.

Who sees it. Your reviewer, in the record.

  • Chased with the sender
  • Accepted as absent, which is your call

Automatic data extraction, four ways to get it

The first column is here to be ruled out. It is the right tool for a different job.

  • What it is actually for

    Write a scraper

    Pages built for people to read, reached by a URL.

    Plain OCR

    Turning a picture of text into text.

    A platform you configure

    Files, once you have told it what to look for.

    Fabrx

    Files, with the reading built for yours.

  • What it hands back

    Write a scraper

    Whatever you selected, as long as the markup holds.

    Plain OCR

    Text in reading order, with no idea which value is which.

    A platform you configure

    The fields you configured, inside their system.

    Fabrx

    A structured record, at your endpoint, in the shape you asked for.

  • What breaks it

    Write a scraper

    A markup change, a rate limit, or a login.

    Plain OCR

    Nothing, because it was never interpreting anything.

    A platform you configure

    A layout the configuration did not anticipate.

    Fabrx

    A file type nobody mentioned, which comes back as an exception.

  • Who maintains it

    Write a scraper

    You do, every time the site changes.

    Plain OCR

    Nobody, because there is nothing configured to maintain.

    A platform you configure

    You do, in their builder.

    Fabrx

    We do. Adding a field is a request, not a procurement cycle.

  • What it costs to find out

    Write a scraper

    An afternoon, and the maintenance starts immediately.

    Plain OCR

    Low, and the work of interpreting the text is still ahead of you.

    A platform you configure

    A licence and an implementation, before you know it reads yours.

    Fabrx

    Nothing. The build is free and you pay per document once it runs.

Someone has run this

They built solutions for any situation encountered along the way. For us, this type of partnership matters.

Operations lead, Operations, BlitzOctober 2026
300+
Field agents using it every working day
4,000+
Documents processed every month
4 yrs
In production, on the same workflow

Blitz runs identity documents into contracts, photographed in the field on phones. These figures describe that pipeline and that kind of file.

Before you pay

What a pilot measures

Measured on your documents

  • Field-level accuracy

    Every field on every sample file, scored against your correction.

  • Sample size

    How many of your real files we ran. You choose which.

  • Spread of the sample

    How many containers your sample covers, scans and photographs included, so the figure is not measured only on the clean ones.

  • Exception rate

    The share that stopped for a person, and what stopped them.

  • Build time

    Calendar time from your samples reaching us to the pipeline running.

What we ask you for

  • How many minutes does one file take today, start to finish?

    It is the only honest denominator for any claim about time saved.

  • How many do you process in a month?

    It sets the pricing band, because the rate depends on volume.

  • How do they reach you, and in how many different ways?

    The container is often harder than the document. A portal login or a zip of forty changes what gets built before anything is read.

  • What proportion are scans or photographs?

    It is the difference between an easy build and a real one.

  • Which fields does somebody actually use afterwards?

    Asking for every field on the page is the default. The ones nobody reads cost the same to build as the ones that matter.

What it costs

  1. 01

    We build it

    We build the extraction, usually within one business day of receiving your sample documents.

  2. 02

    You check it

    It costs you nothing until it runs on your real documents. If the accuracy is not good enough, nothing goes into production and nothing is charged.

  3. 03

    Then you pay

    A prepaid package priced per document, valid for a calendar year. The rate depends on the document type and your monthly volume.

  • No subscription
  • No per-seat charge
  • No setup fee at any volume

Questions about automatic data extraction

What is automatic data extraction?

Pulling specific values out of a source and handing them over as structured data, without a person reading the source first. The source can be a web page, a database, or a file. Which one it is decides almost everything about how it is done.

What is the difference between data extraction and web scraping?

Scraping is extraction from pages built for people to read. It works by following the markup, so it breaks when the markup changes, and it has to live with logins, rate limits and what the site permits. Extracting from a file has none of those problems and one that scraping does not have: there is no markup at all, so nothing in the file says which number is the total.

Can data be extracted automatically from a PDF?

Yes. A PDF that was generated rather than scanned already carries its text, which makes it the easier case. The harder part is that a PDF has layout and no structure, so a table can be a set of positioned words that only look like columns.

What about a scan or a photograph of a document?

Both are ordinary input for printed pages, including the creased, skewed and badly lit ones. Where it stops is resolution low enough that a person would struggle too, and those come back as unreadable instead of guessed at.

What happens to a file it cannot read?

It stops and is reported. A locked or damaged file is returned with which of the two it was. An unfamiliar type is set aside and we are told, usually the same day. Nothing is sent to your endpoint on a guess.

Does it need a template for each layout?

No. The extraction reads fields by what they are, so it does not depend on a sender keeping one layout. Template based tools are tied to where each field sits, which is why a redesign breaks them.

Related

Send us twenty files

We build the extraction and show you what it got right, field by field, before you pay anything.