AI
Training data
Collection, annotation and review of model training data, including European Portuguese and the African variants.
What is delivered
Annotation guidelines calibrated before volume, not during it
Inter-annotator agreement measured and reported
Coverage of PT-PT, PT-BR and the Angolan and Mozambican variants
Why is European Portuguese scarce in data?
Because of internet demographics. Brazil has around twenty times Portugal's population and a proportionally larger digital presence, so any collection weighted by volume ends up overwhelmingly Brazilian, and the language tag rarely distinguishes the variant.
The practical effect shows up in service: gerund constructions where European Portuguese uses the prepositional infinitive, Brazilian vocabulary in Portuguese contexts, and a familiar register where the customer expects formal courtesy. These are errors no factual accuracy test catches.
The African variants are even less represented. Angola, Mozambique and Cape Verde follow the European written standard with their own local vocabulary, and administrative terminology differs enough that using the wrong term reads as not knowing the country.
How is an evaluation set built?
From real, already-decided cases, chosen to cover the real distribution rather than the convenient one. Between five hundred and two thousand well-chosen examples are enough to measure with useful confidence whether a system behaves differently across the two variants.
The part done badly is labelling. If two experienced annotators disagree on the correct answer in fifteen per cent of cases, no system will do better than that, and that number is valuable information: it measures how much of the problem is genuinely ambiguous.
The set is built before any tuning work and kept separate. A set used for tuning no longer serves to evaluate, and the temptation to reuse it is strongest precisely when results start looking good.
How is the variant identified and verified?
With unambiguous markers, applied to a manually labelled sample. Gerund versus prepositional infinitive construction is the most reliable: está a fazer against está fazendo appears in any text of length and has no overlap between variants.
Then, high-frequency lexical pairs that work well for service text, and pronoun placement, which differs systematically and is hard to imitate without native knowledge. Three independent, agreeing markers give a classification that can be trusted.
What does not work is spelling. The orthographic agreement brought spelling close enough that classifying by it produces false positives in quantity, and a classifier built on spelling appears to work in testing and fails in production.
What legal questions arise in collection?
The legal basis for processing personal data appearing in the material. Call transcripts, support messages and email almost always contain personal data, and using them to train a system is a purpose distinct from the one they were collected for.
And copyright over the collected text. Public content is not free content, and mass collection from the internet has its own framework in the European Union, with text and data mining exceptions subject to rights reservation by the holder.
The approach avoiding most of these problems is to start with the client's own data, pseudonymised before annotation and with a documented legal basis for the new purpose. It is less material than the whole internet and describes exactly the domain the system will operate in.
How is annotation quality assured?
With guidelines written before starting and with examples from both sides of each boundary. The difference between well-briefed and badly briefed annotators is larger than the difference between good and average annotators, and costs far less to fix.
With regular calibration: sessions where several annotators mark the same cases and discuss disagreements. They reveal ambiguities in the guidelines that no solitary review uncovers, and should happen at the start and periodically, because interpretation drifts.
And in conditions that sustain attention. Annotation is repetitive and quality falls measurably after a few continuous hours. Operations measuring only volume per hour get volume per hour, and the degradation only shows when someone audits a late sample.
Where is this work done?
European Portuguese annotation is done by native speakers from Portugal, which in practice means Portugal. It is not a preference but the requirement, and an annotator who is not native to the variant does not reliably detect what they are assessing.
For the African variants, the Lusophone corridor, with annotators from Angola, Mozambique and Cape Verde. None of those countries holds an adequacy decision, so annotated material moves under standard contractual clauses and an impact assessment, or is synthetic and the question does not arise.
Where the material contains personal data and pseudonymisation is insufficient, we propose the European model and say why. Annotation is work where a person reads the content in full, which makes the residency question more sensitive than in automated processing.
Who owns what is produced?
You do, including the annotations, the guidelines and the evaluation set. It is written into the contract before starting, because it is the point most frequently left ambiguous in data work and the most expensive to resolve once material exists in quantity.
The annotation guidelines are the part whose value is most underestimated. They describe how to decide on hard cases, were built against real cases and calibrated between annotators, and rebuilding them from scratch costs more than the initial annotation. They are the asset rather than the by-product.
Where training data can be run from
Not every delivery model suits every service. The table shows only those that make sense for this work, with the data residency position of each.
| Model | Where | When it makes sense | Personal data |
|---|---|---|---|
| Lusophone corridor | Coordination in Lisbon, operations in Luanda, Maputo and Praia | When you operate in Angola or Mozambique and need European governance | The Lisbon tier stays in the EEA. Local operations require standard contractual clauses. |
| Global network | Uzbekistan, the Philippines, Poland, the Dominican Republic, Mexico, Colombia, Turkiye and Africa | When you need continuous cover, specific languages or the lowest cost | Poland is inside the EEA. The others require standard contractual clauses. |
The data column describes the applicable framework and is not legal advice. The detail is in international data transfers.
This service carries a variant decision
Are you serving customers in Portugal, in Brazil or in both? The answer changes the operating design, the scripts, who reviews quality and how results are reported.
Related services
- AI managed servicesOngoing operation of AI-assisted workflows, with people in the loop where the risk requires it.
- Agents and automationAutomation of repetitive processes and assisted agents, with supervision and defined limits.
- Evaluation and safetyModel output evaluation, adversarial testing and safety review, by native speakers.
- Document intelligenceExtraction and classification of information from documents, with human verification where it matters.
Frequently asked questions
Are annotators native to the variant?
They are. An annotator who is not native does not reliably detect what they are assessing.
How much data is needed?
For evaluation, five hundred to two thousand well-chosen examples. For correction, far more, which is why evaluation comes first.
How do you identify a text's variant?
By structural markers such as verb construction and pronoun placement. Spelling has not worked since the orthographic agreement.
Do you use our data or collect it?
We prefer yours, pseudonymised and with a documented legal basis for the new purpose. It describes exactly the domain.
Do you cover the African variants?
We do, through the Lusophone corridor, with annotators from Angola, Mozambique and Cape Verde.
How do you ensure consistency between annotators?
With guidelines written before starting and periodic calibration sessions, because interpretation drifts over time.
Is the evaluation set reused?
No. A set used for tuning no longer serves to evaluate, and the temptation to reuse it grows as results improve.
Are there copyright questions?
There are, in collection from the internet. Public content is not free content, and starting with your own data avoids most of it.
Let us look at the numbers for your case
Tell us which processes you want to outsource, in which languages and at what volume. We come back with a euro estimate and an operating design, with no commitment.
We reply within 6 hours on working days. If you would rather write: info@corpshore.solutions