BPO
Data processing
Extraction, classification, validation and enrichment of data at volume, with measured quality control.
What is delivered
Agreed validation rules, applied consistently
Double-keying on the critical fields you define
Per-batch quality reporting, with error causes
Where does automation reach and where does it not?
It reaches the normal case, which in a mature document process is most of the volume. Extracting fields from structured documents, classifying by type, validating against rules and external sources: all of it is automatable with high accuracy and with cheap, visible error.
It does not reach the badly scanned document, the format nobody anticipated, the discrepancy between two sources that both look reliable, and the case where the rule does not clearly apply. That fraction is small in volume and consumes most of the time, and it is work for people with context.
The design that works does not choose between the two: it automates what is automatable, routes the rest to people with the context to hand, and uses the human-resolved cases to improve the automation. What must not happen is automation closing cases that are not resolved, and we measure that explicitly.
How is accuracy measured and assured?
With double entry on critical fields and statistical sampling on the rest. Double entry is expensive and reserved for fields where an error has material consequence; sampling covers the rest at an agreed confidence level and adjusts automatically when the error rate leaves the range.
Agreed accuracy is defined per field rather than as a single number. Ninety-nine point five per cent means very different things applied to a notes field and to a tax identification number, and a contract with a single percentage is hiding that difference.
Errors are classified by cause, and the largest category is almost always ambiguous instruction rather than carelessness. An operation correcting people when it should be correcting the instruction repeats the same error indefinitely with different people.
What volumes does this model support?
From a few thousand to several million documents a year, with the team's composition changing with scale. Below a certain volume, automation does not pay for itself and the work is predominantly human; above it, the proportion inverts and the team mostly handles exceptions.
Seasonal peaks are handled with capacity shared between clients, which is why this model works for operations whose volume varies three or four times between trough and peak. Sizing internally for the peak is expensive; sizing for the trough misses the peak.
The practical limit is rarely processing capacity and almost always input quality. A document flow with inconsistent formats and variable scanning costs more to handle than twice the volume well prepared, and fixing the source is frequently the best investment available.
What data framework applies?
This service usually runs from Brazil, the Lusophone corridor or the global network, because it is asynchronous work where the time difference does not penalise. None of those destinations holds an adequacy decision, so all require standard contractual clauses and a transfer impact assessment.
The measure that most reduces the problem is pseudonymisation before transfer, with the information needed to re-identify staying at origin. For many document processes this is workable and turns an uncomfortable assessment into a defensible one.
Where documents contain health data or are processed on behalf of public bodies, we propose the onshore model and say why. It is the same answer we give across every service and it does not change because the volume is large.
How is the work delivered and integrated?
Preferably by direct integration with your system, through an API or secure file transfer in an agreed format. The alternative, someone copying results from one system to another, introduces a layer of error that cancels part of the gain and is surprisingly common.
Access negotiation happens in the design phase rather than in transition. Access to third-party systems has its own authorisation processes that take weeks and that nobody can accelerate on launch day.
Each delivered batch comes with its own quality report: volume handled, exception rate, accuracy measured on the sample, and a list of returned cases with reasons. A batch delivered without that report requires you to trust, and trust is not a control.
How do you prepare the source before automating?
By measuring the variability of what arrives. A flow with ten stable formats is handled completely differently from one with a hundred irregular variations, and the cost difference between them is larger than the volume difference.
Frequently the best-return intervention is upstream rather than with us. A supplier switching to an agreed format, or a scanner replaced, reduces cost more than any downstream extraction improvement, and we say so even when it means a smaller project.
Where the source cannot be fixed, we design for the variability rather than ignoring it: classification by format before extraction, and rules per class rather than one general rule that fails on half the cases with nobody knowing which.
What happens to the data at the end of the contract?
It is returned in the agreed format and deleted from our systems within the period written into the processing agreement, with written confirmation that deletion happened. It is not a courtesy: it is a processor obligation and should be quantified before signing.
It is worth asking any supplier this in the first meeting, including what happens to backups and test environments. It is the question that separates those who thought about the end of the contract from those who only thought about the start, and the answer takes seconds to give when it exists.
Where data processing can be run from
Not every delivery model suits every service. The table shows only those that make sense for this work, with the data residency position of each.
| Model | Where | When it makes sense | Personal data |
|---|---|---|---|
| Brazil | Sao Paulo | When scale and cost are the priority, or the market served is Brazilian | No adequacy decision. Requires standard contractual clauses and a transfer impact assessment. |
| Lusophone corridor | Coordination in Lisbon, operations in Luanda, Maputo and Praia | When you operate in Angola or Mozambique and need European governance | The Lisbon tier stays in the EEA. Local operations require standard contractual clauses. |
| Global network | Uzbekistan, the Philippines, Poland, the Dominican Republic, Mexico, Colombia, Turkiye and Africa | When you need continuous cover, specific languages or the lowest cost | Poland is inside the EEA. The others require standard contractual clauses. |
The data column describes the applicable framework and is not legal advice. The detail is in international data transfers.
Related services
- Customer supportContact centre operations in Portuguese, English and more than thirty languages, across voice, email, chat and social.
- Technical supportLevel 1 and 2 support for software, telecoms and hardware products, with structured escalation.
- Back officeAdministrative processing, document management, data entry and internal operations support.
- Finance and accountingAccounts payable and receivable, reconciliations, invoicing and close support, alongside your own accountants.
Frequently asked questions
Which document formats do you handle?
Scanned, native PDF, images and structured feeds. Scan quality affects cost more than format does.
What accuracy is agreed?
Per field rather than as a single number, because the consequence of an error varies greatly between fields.
Can data be pseudonymised before it leaves?
It can, and for most document processes it is what we recommend.
How do you handle seasonal peaks?
With capacity shared between clients, which is what makes this model viable for volumes varying three or four times.
Who decides on ambiguous cases?
The team, within what is documented; the rest is returned with the reason rather than guessed.
Do you integrate with our system?
By API or secure file transfer. We avoid manual copying between systems, which introduces error without adding value.
How long does launch take?
Six to ten weeks for a well-documented process, more if the documentation has to be built first.
Does it suit health data?
Only in the onshore model, inside the EEA. For that category we propose no destination outside the European Economic Area.
Let us look at the numbers for your case
Tell us which processes you want to outsource, in which languages and at what volume. We come back with a euro estimate and an operating design, with no commitment.
We reply within 6 hours on working days. If you would rather write: info@corpshore.solutions