Skip to content

AI

Document intelligence

Extraction and classification of information from documents, with human verification where it matters.

What is delivered

  • Extraction with a confidence threshold and review below it

  • Handling of documents in Portuguese, English and Spanish

  • Accuracy measured per field, not just per document

What determines whether a document is easy or hard?

Format consistency, far more than document type. Invoices from thirty different suppliers with thirty different layouts are harder than long contracts with a stable structure, and intuition usually says the opposite.

Then, scan quality, which is the variable that most frequently dominates cost. A flow with variable scanning costs more to process than twice the volume well prepared, and fixing the source is frequently the best investment available before any automation.

And whether there is a source of truth to validate against. Extracting a number from an invoice is easy; knowing whether it is right requires matching it against the order, and it is the matching that produces value. Extraction without validation produces new data nobody trusts.

How is acceptable accuracy defined?

Per field rather than as a single number, because the consequence of an error varies enormously between fields. Ninety-nine point five per cent means very different things applied to a notes field and to a tax identification number.

On critical fields we use double checking, which is expensive and reserved for cases where an error has material consequence. On the rest, statistical sampling at an agreed confidence level, which rises automatically when the error rate leaves the expected range.

And we define what happens when the system lacks sufficient confidence: the case goes to human review rather than being filled with a best guess. A system that always fills in has higher apparent accuracy and produces silent errors, which are the expensive ones.

How does the human layer work?

It handles what automation returns and what fails validation, with the document and the context in view. Interface design matters more than it seems: a review where the person has to hunt for the field in the document is several times slower than one where the field is highlighted.

Each human correction feeds automation improvement, which only works if corrections are recorded in a structured way. A team correcting directly in the destination system without recording what it corrected is resolving cases and wasting the most valuable information it produces.

And the proportion between automation and the human layer is reported and discussed rather than hidden. A supplier presenting an automation rate without saying how many cases the human layer handled is describing half the operation.

What data framework applies?

Business documents almost always contain personal data: names, contacts, identification numbers and sometimes more. Processing requires a legal basis, and extraction into an analytical system is frequently a purpose distinct from the one the document was received for.

This service runs onshore, from Brazil or from the global network depending on the model chosen. Outside the European Economic Area, standard contractual clauses and an impact assessment apply, and the measure that most reduces the problem is excluding unnecessary fields at ingestion.

For clinical documentation or documents processed on behalf of public bodies, we propose only the onshore model. It is the same answer we give across every service and it does not change because the volume is large or the deadline short.

How does it integrate with destination systems?

Through a programmatic interface or secure file transfer in an agreed format. The alternative, someone copying results from one system to another, introduces a layer of error that cancels part of the gain and is surprisingly common in operations that are otherwise well built.

Each delivered batch comes with its report: volume handled, exception rate, accuracy measured on the sample, and a list of returned cases with reasons. A batch delivered without that report requires you to trust, and trust is not a control.

Access negotiation happens in the design phase. Access to third-party systems has its own authorisation processes taking weeks that nobody can accelerate on launch day, and it is the most common cause of delayed launches in this service.

When is this service not worth it?

When volume is low and the format is stable. A few thousand documents a year with a consistent layout are better solved by direct integration with whoever issues them, and proposing extraction there is selling a layer that did not need to exist.

When the document is the easy part and the decision is the hard one. Extracting ten fields from a contract is trivial; deciding what they mean for your business is not, and an extraction service does not solve the second problem even while solving the first perfectly.

And when the source can be fixed. If half the documents arrive badly scanned because someone upstream uses inadequate equipment, fixing that costs less and improves more than any downstream automation, and we say so even when it means a smaller project.

How does the service improve over time?

With human corrections feeding the automation, which only works if they are recorded in a structured way rather than corrected directly in the destination. An operation that corrects without recording resolves cases and wastes the most valuable information it produces daily.

And with periodic review of what keeps going to the human layer. A category that was two per cent of exceptions and became fifteen indicates something changed at source, and catching that early avoids months of manually handling what a new rule would resolve.

Where document intelligence can be run from

Not every delivery model suits every service. The table shows only those that make sense for this work, with the data residency position of each.

ModelWhereWhen it makes sensePersonal data
Onshore PortugalLisbon, Porto, Braga, Coimbra, Aveiro, Faro, Funchal and Ponta DelgadaWhen data cannot leave the EEA, or when the end customer is PortugueseStay inside the EEA. No transfer.
BrazilSao PauloWhen scale and cost are the priority, or the market served is BrazilianNo adequacy decision. Requires standard contractual clauses and a transfer impact assessment.
Global networkUzbekistan, the Philippines, Poland, the Dominican Republic, Mexico, Colombia, Turkiye and AfricaWhen you need continuous cover, specific languages or the lowest costPoland is inside the EEA. The others require standard contractual clauses.

The data column describes the applicable framework and is not legal advice. The detail is in international data transfers.

Frequently asked questions

Which formats do you handle?

Scanned, native PDF, images and structured feeds. Format consistency affects cost more than document type does.

What accuracy is agreed?

Per field rather than as a single number, with double checking on fields where an error has material consequence.

What happens when the system lacks confidence?

The case goes to human review. Always filling in with a best guess produces silent errors, which are the expensive ones.

Do you validate against other sources?

We do, and it is the matching that produces value. Extraction without validation produces new data nobody trusts.

Can the data leave the EEA?

Depending on the model, with standard contractual clauses and an impact assessment. For clinical documentation we propose onshore only.

Do you report the automation rate?

We do, with the number of cases handled by the human layer alongside. One without the other describes half the operation.

How do you integrate with our system?

Through a programmatic interface or secure file transfer. We avoid manual copying, which introduces error without adding value.

When is it not worth it?

Low volume with a stable format, or when the source can be fixed. We say so even when it means a smaller project.

Let us look at the numbers for your case

Tell us which processes you want to outsource, in which languages and at what volume. We come back with a euro estimate and an operating design, with no commitment.

We reply within 6 hours on working days. If you would rather write: info@corpshore.solutions