Skip to content

AI

AI in Portuguese

AI work that treats Portuguese as two or more operations, because that is what it is.

What is delivered

  • Datasets separated by variant, with provenance metadata

  • Evaluation that catches the wrong variant, not merely the wrong language

  • Coverage of the Angolan, Mozambican and Cape Verdean variants

Why does a Portuguese-language system fail in Portugal?

Because the overwhelming majority of Portuguese available in public corpora is Brazilian, and the language tag rarely distinguishes the variant. A system trained on that material produces correct Portuguese in a variant that is not the Portuguese customer's, and no factual accuracy test reveals it.

What the customer notices is verb construction, service vocabulary and forms of address. Está a verificar against está verificando, utilizador against usuário, ecrã against tela, and a familiar register where measured courtesy was expected. That is ten to fifteen occurrences in a typical purchase journey.

And they notice it in a context where they are already dissatisfied, because whoever contacts support usually has a problem. The wrong register in a moment of frustration reads as not being taken seriously, which is a considerably harder complaint to resolve than the original one.

Where does the correction start?

With the European-variant evaluation set, before any tuning work. Without it there is no way to say whether a change improved anything, and the reverse order produces months of work whose effect nobody can demonstrate.

Between five hundred and two thousand well-chosen examples labelled by native speakers are enough to measure with useful confidence. It is weeks of work rather than months, and it is always the first thing we do, even when the client arrives with a solution already chosen.

Frequently the measurement changes the plan. A system expected to need rewriting turns out to be adequate with a knowledge base written in the right variant, and the intervention becomes a content one rather than a model one, at a fraction of the cost.

Which correction paths exist?

The cheapest and frequently sufficient path is retrieval over content written natively in the right variant. If the system answers from text a native speaker wrote, much of the register problem disappears without any work on the model.

The middle path is explicit instruction on variant and register, with examples, verified against the evaluation set. It works reasonably for vocabulary and less well for verb construction, which is structural and resurfaces when the answer is long.

The most expensive is tuning with variant data, which requires orders of magnitude more material than evaluation and is only justified when the first two were not enough. Proposing this path first is common and is selling the most expensive work before knowing whether it is needed.

How do you verify the correction worked?

By comparing European-variant performance before and after, on the separate set never used for tuning. The number that matters is not the overall one but the variant that was worse, because that is what the whole exercise was for.

Also checking whether the other variant got worse. Tuning a system to produce European Portuguese can degrade its Brazilian Portuguese output, and an operation serving both markets can solve one problem by creating another without anyone noticing for months.

And confirming with people, in blind assessment. Automatic text-similarity metrics do not capture register, which is precisely what is at stake, and an assessment by native speakers who do not know which output is which is the test that counts.

How do the African variants fit in?

Angola, Mozambique, Cape Verde, Guinea-Bissau and São Tomé and Príncipe follow the European written standard with their own local vocabulary and substantial influence from national languages. A European Portuguese system serves them reasonably better than a Brazilian one, and does not serve them well.

The usual correction is a local glossary added over the European base rather than a separate system. It is proportionally cheap and resolves most of the friction, particularly in administrative terminology, where using the wrong term reads as not knowing the country.

In Cape Verde there is also Creole, which is a language in its own right rather than a variant, and treating it as a Portuguese variant is a category error that shows immediately. Where Creole is needed, it is separate work and we say so.

What result can be expected?

A measurable improvement in the variant that was worse, and that is what we promise: the measurement before, the measurement after, and the difference. We do not promise a specific number before measuring, because doing so is promising without information.

Frequently the commercial effect appears before the technical one. Complaints about tone fall, requests to speak to someone else fall, and first-contact resolution in the Portuguese market rises, and these three indicators already exist in your operation and do not need to be created.

And there is a result that is not about the system: a way of knowing comes into existence. An organisation with a per-variant evaluation set can answer the question of how the system behaves in its market, and that capability survives any change of technology.

What stays with you after this work?

The European-variant evaluation set, which is the project's most durable asset. It survives any change of technology and lets you answer the question of how the system behaves in your market whenever it is asked, without redoing the work from scratch.

And a written description of the system's known limitations in the variant you care about. It is the part missing from almost all documentation and it is what answers the question someone asks when a customer complains about tone, which is when the answer is needed urgently.

Where ai in portuguese can be run from

Not every delivery model suits every service. The table shows only those that make sense for this work, with the data residency position of each.

ModelWhereWhen it makes sensePersonal data
Lusophone corridorCoordination in Lisbon, operations in Luanda, Maputo and PraiaWhen you operate in Angola or Mozambique and need European governanceThe Lisbon tier stays in the EEA. Local operations require standard contractual clauses.
Global networkUzbekistan, the Philippines, Poland, the Dominican Republic, Mexico, Colombia, Turkiye and AfricaWhen you need continuous cover, specific languages or the lowest costPoland is inside the EEA. The others require standard contractual clauses.

The data column describes the applicable framework and is not legal advice. The detail is in international data transfers.

This service carries a variant decision

Are you serving customers in Portugal, in Brazil or in both? The answer changes the operating design, the scripts, who reviews quality and how results are reported.

See PT-PT and PT-BR

Frequently asked questions

Why does our system sound Brazilian?

Because most Portuguese in public corpora is Brazilian and the language tag rarely distinguishes the variant.

Where do we start?

With the European-variant evaluation set. Without it there is no way to say whether a change improved anything.

Do we need to train a new model?

Frequently not. Retrieval over content written natively in the right variant solves much of the problem.

How long does measurement take?

Weeks rather than months, with five hundred to two thousand examples labelled by native speakers.

Can the other variant get worse?

It can, and we measure both. An operation serving both markets can solve one problem by creating another.

Do you cover the African variants?

We do, usually with a local glossary over the European base rather than a separate system.

And Cape Verdean Creole?

It is a language in its own right rather than a Portuguese variant. Where it is needed, it is separate work and we say so.

What result do you promise?

The measurement before, the measurement after, and the difference. We do not promise a specific number before measuring.

Let us look at the numbers for your case

Tell us which processes you want to outsource, in which languages and at what volume. We come back with a euro estimate and an operating design, with no commitment.

We reply within 6 hours on working days. If you would rather write: info@corpshore.solutions