Skip to content

Artificial intelligence

Training data in European Portuguese: why it is scarce and what to do

The overwhelming majority of Portuguese in public corpora is Brazilian. Systems trained on that data answer a Portuguese customer in the wrong register.

Corpshore Portugal editorial team

Written by the team that builds these operations. No individual byline: this is internally reviewed work, not personal opinion.

Published

Training data in European Portuguese: why it is scarce and what to do

Where does the imbalance come from?

From internet demographics. Brazil has around twenty times Portugal's population and a proportionally larger digital presence. Any collection weighted by volume ends up overwhelmingly Brazilian, and the language tag rarely distinguishes the variant.

The practical effect shows up in service: gerund constructions where European Portuguese uses the infinitive, Brazilian vocabulary in Portuguese contexts, and a familiar register where the customer expects formal courtesy.

What do you do in practice?

Collect and annotate European-variant data with native speakers from Portugal, tagging the variant explicitly rather than assuming Portuguese is one thing. That means building the evaluation set first, because without it there is no way to know whether the correction worked.

Evaluate per variant rather than in aggregate. A system at ninety per cent overall accuracy can be at seventy in European Portuguese, and the average hides it. The same principle applies to the Angolan, Mozambican and Cape Verdean variants.

How do you measure a dataset's variant?

With unambiguous markers, applied to a manually labelled sample. Gerund versus prepositional infinitive construction is the most reliable: está a fazer against está fazendo appears in any text of length and has no overlap between variants.

Then come high-frequency lexical pairs, which work well for service text: utilizador and usuário, ecrã and tela, telemóvel and celular, morada and endereço. And pronoun placement, which differs systematically and is hard to imitate without native knowledge.

What does not work is spelling. The orthographic agreement brought the two variants' spelling close enough that classifying by spelling produces false positives in quantity. A classifier built on spelling appears to work in testing and fails in production, which is the worst possible failure mode.

How much data is genuinely needed?

For evaluation, less than people expect: between five hundred and two thousand well-chosen, well-labelled examples are enough to measure with useful confidence whether a system behaves differently across the two variants. That is weeks of work rather than months, and it is what should be done first.

For correction, far more, and it depends on what is being corrected. Adjusting register and vocabulary in text generation needs orders of magnitude more than evaluation. That is why sequence matters: collecting training data before knowing where the system fails collects the wrong thing in quantity.

There is also a middle path that is frequently sufficient: retrieval over a knowledge base written in the right variant, rather than model tuning. If the system answers from text someone wrote in European Portuguese, much of the register problem disappears with no training data at all.

Who should annotate, and under what conditions?

Native speakers of the variant, not Portuguese speakers in general, and with written guidelines produced before starting. The difference between well-briefed and badly briefed annotators is larger than the difference between good and average annotators, and it costs far less to fix.

With regular calibration. Sessions where several annotators mark the same cases and discuss disagreements reveal ambiguities in the guidelines that no solitary review uncovers, and should happen at the start and then periodically, because interpretation drifts over time.

And in working conditions that sustain attention. Annotation is repetitive and quality falls measurably after a few continuous hours. Operations measuring only volume per hour get volume per hour, and the degradation only shows when someone audits a late sample.

What legal questions arise in collection?

Two main ones. The first is the legal basis for processing personal data appearing in the material: call transcripts, support messages and email almost always contain personal data, and using them to train a system is a purpose distinct from the original purpose of collecting them.

The second is copyright over the collected text. Public content is not free content, and mass collection from the internet has its own framework in the European Union, with text and data mining exceptions subject to rights reservation by the holder.

The approach that avoids most of these problems is to start with your own data, pseudonymised before annotation and with a documented legal basis for the new purpose. It is less material than the whole internet and it is material describing exactly the domain the system will operate in.

How do you know the correction worked?

By comparing European-variant performance before and after, on the separate evaluation set that was never used for tuning. The number that matters is not the overall one but the variant that was worse, because that is what the whole exercise was for.

It is worth also measuring whether the other variant got worse. Tuning a system to produce European Portuguese can degrade its Brazilian Portuguese output, and an operation serving both markets can solve one problem by creating another without anyone noticing for months.

And confirm with people. Automatic text-similarity metrics do not capture register, which is precisely what is at stake. A blind assessment by native speakers, where they do not know which output is which, is the test that counts and is considerably cheaper than it sounds.

Which sources exist and which are usable?

Public Portuguese corpora exist and are mostly Brazilian, with the variant rarely tagged. They work as a starting point and require prior classification, and that classification has to be manually validated on a sample before being trusted at scale.

Portuguese institutional sources, such as official publications, legislation and administrative documentation, are unambiguously European variant and of high linguistic quality. The problem is register: it is formal administrative language, and training service responses on it produces replies that sound like official correspondence.

The operation's own data is the most valuable and the most restricted. Call transcripts, support conversations and email describe exactly the domain and register that matter, and contain personal data requiring a legal basis for the new purpose and pseudonymisation before annotation.

Synthetic data produced by native speakers from real cases is the compromise that works most often: it keeps the register and the domain, carries no personal data, and can be produced in volume aimed at the gaps the evaluation identified.

How do you measure register, and not just accuracy?

With blind human assessment, the only instrument that reliably captures register. Native speakers of the variant compare outputs without knowing which came from where and score appropriateness, and differences between systems appear clearly even with relatively small samples.

With automatic markers as a complement rather than a substitute. Gerund construction frequency, pronoun placement and the incidence of the other variant's vocabulary are automatically measurable and detect gross degradation, which is useful as a continuous alarm between human evaluations.

And with real complaints as ground truth. If customers in one market start commenting on the tone of replies, that is higher-quality information than any internal metric, and an operation with no channel for it to reach the team tuning the system is wasting it.

The mistake to avoid is measuring only what is easy. Factual accuracy is easily measured and register appropriateness is not, and an operation optimising what it can measure ends with a factually correct system that sounds foreign, which is precisely the problem it started with.

Frequently asked questions

Why so little European data?
Because collection weights by volume and Brazil has around twenty times Portugal's population.
Do models distinguish the variants?
Not always, and rarely without being explicitly asked. It has to be evaluated.
Where to start?
With the European-variant evaluation set, before collecting training data.
Who should annotate?
Native speakers of the variant, not Portuguese speakers in general.
How is the variant tagged?
Explicitly, at collection time. Afterwards it is practically impossible to separate reliably.
What does the global average hide?
A much weaker performance on the minority variant, which is often the market that pays.
How much data is needed?
Less than expected for evaluation, far more for correction. Start with evaluation.
What about the African variants?
They are even less represented, and the same tagging and evaluation discipline applies.

Let us look at the numbers for your case

Tell us which processes you want to outsource, in which languages and at what volume. We come back with a euro estimate and an operating design, with no commitment.

We reply within 6 hours on working days. If you would rather write: info@corpshore.solutions