Skip to content

An AI research lab

Model evaluation across four Portuguese variants

Native evaluators of European, Brazilian, Angolan and Mozambican Portuguese assessed model output by variant rather than merely by language.

Model evaluation across four Portuguese variants

Case facts

Delivery model
Lusophone corridor
Variant
multilingue
Team
19 evaluators, Lusophone corridor
Time to launch
5 weeks of calibration

The problem

The model answered correctly in Portuguese but with a Brazilian register to Portuguese users, and existing evaluation sets did not catch it because they treated Portuguese as a single language.

What we built

Per-variant evaluation rubrics with boundary examples, and native evaluators for each. Results are reported by variant and by risk category, with reproducible examples.

Compliance

Part of the operation sits outside the EEA, with standard contractual clauses. Client data is used only for that client's work, by contractual clause.

Results

variants assessed separately
4variants assessed separately
inter-annotator agreement
0,82inter-annotator agreement
native evaluators
19native evaluators
calibration before volume
5 sem.calibration before volume

We treated Portuguese as one language. It is at least four operations.

Head of evaluation, An AI research lab

Let us look at the numbers for your case

Tell us which processes you want to outsource, in which languages and at what volume. We come back with a euro estimate and an operating design, with no commitment.

We reply within 6 hours on working days. If you would rather write: info@corpshore.solutions