AI
Evaluation and safety
Model output evaluation, adversarial testing and safety review, by native speakers.
What is delivered
Written evaluation rubrics, with boundary examples
Evaluators native to the variant assessed, not merely to the language
Results by risk category, with reproducible examples
Why is overall accuracy not enough?
Because it hides exactly what matters. A system at ninety per cent overall accuracy can be at seventy in one case category or one language variant, and the average does not reveal it. The question a regulator asks is about the group where the system performs worst, and it is the one the average prevents you answering.
We measure per case category, per variant and per user segment where relevant. That requires an evaluation set built to cover those splits, which in turn requires deciding in advance which splits matter, and that is a business decision rather than a technical one.
And we measure against human agreement, which is the ceiling. If two experienced reviewers disagree on fifteen per cent of cases, no system will exceed eighty-five, and announcing ninety is a sign the evaluation is measuring something else.
How do you test adversarially?
By actively seeking the inputs that make the system fail, rather than confirming it works on cases someone chose. It is human work and it is deliberately hostile to the system, because the real user will be too, even without intending to.
The categories that fail most frequently are predictable: ambiguous inputs, inputs in a variant or register different from training, very long or very short inputs, and inputs combining two requests in one sentence. An adversarial test not covering those four is not adversarial.
We also test what happens when the system is pushed outside its scope. A system designed to answer about orders that answers about anything else when someone insists is a reputational risk and, depending on what it says, a legal one.
How do you assess register and not just accuracy?
With blind human assessment, the only instrument that reliably captures register. Native speakers of the variant compare outputs without knowing which came from where and score appropriateness, and differences between systems appear clearly even on relatively small samples.
With automatic markers as a complement rather than a substitute: gerund construction frequency, pronoun placement and the incidence of the other variant's vocabulary are automatically measurable and detect gross degradation between human evaluations.
The mistake to avoid is measuring only what is easy. Factual accuracy is easily measured and register appropriateness is not, and an operation optimising what it can measure ends with a factually correct system that sounds foreign to its own market.
What does continuous evaluation add?
It detects degradation, which happens without anything changing in the system. Input data changes, users change behaviour, and a system evaluated once at launch is being assessed against a world that no longer exists.
We do it with a sample periodically reviewed by people and with known control cases seeded into the queue. Control cases measure two things at once: whether the system is still getting things right and whether human review is still happening.
And with a written threshold beyond which automation for a case type is suspended. Having the threshold defined beforehand avoids the argument about defining it during the incident, which is when the pressure not to suspend is greatest.
What documentation comes out of this evaluation?
A description of the evaluation set, the results per category and per variant, and the system's known limitations. That last part is missing from almost all automated system documentation and it is what answers the question someone asks when something goes wrong.
A change log with measured effect: when the system changed, why, and what happened to the numbers afterwards. A system that changes without a record makes any historical analysis invalid, and the first time that matters is always during an investigation.
And the material supporting the Artificial Intelligence Regulation's obligations according to risk level: risk management, data quality, technical documentation and evidence of effective human oversight. We produce the material; management responsibility stays with the client, because it is not delegable.
Where is this work done?
European Portuguese evaluation is done by native speakers from Portugal, because a reviewer who is not native to the variant does not reliably detect what they are assessing. It is the same requirement that applies to annotation and for the same reason.
For the African variants, the Lusophone corridor. For evaluation over synthetic sets, the global network, because no personal data is involved and the residency question does not arise.
Where evaluation involves reading real outputs containing personal data, the same rules apply as to any other processing, and human review makes the question more sensitive than automated processing, because a person reads the content in full.
What uncomfortable things does an honest evaluation produce?
Usually a number worse than the organisation expected, and that is the value. An evaluation confirming what was already believed added no information, and the resistance we most frequently meet is not technical: it is that someone has already communicated a number the measurement does not support.
We deliver the number with the context needed to interpret it, including the ceiling imposed by human agreement, which frequently explains much of the gap. The objective is not showing the system is bad; it is knowing where it fails in order to decide what to do with that information.
Where evaluation and safety can be run from
Not every delivery model suits every service. The table shows only those that make sense for this work, with the data residency position of each.
| Model | Where | When it makes sense | Personal data |
|---|---|---|---|
| Lusophone corridor | Coordination in Lisbon, operations in Luanda, Maputo and Praia | When you operate in Angola or Mozambique and need European governance | The Lisbon tier stays in the EEA. Local operations require standard contractual clauses. |
| Global network | Uzbekistan, the Philippines, Poland, the Dominican Republic, Mexico, Colombia, Turkiye and Africa | When you need continuous cover, specific languages or the lowest cost | Poland is inside the EEA. The others require standard contractual clauses. |
The data column describes the applicable framework and is not legal advice. The detail is in international data transfers.
This service carries a variant decision
Are you serving customers in Portugal, in Brazil or in both? The answer changes the operating design, the scripts, who reviews quality and how results are reported.
Related services
- AI managed servicesOngoing operation of AI-assisted workflows, with people in the loop where the risk requires it.
- Training dataCollection, annotation and review of model training data, including European Portuguese and the African variants.
- Agents and automationAutomation of repetitive processes and assisted agents, with supervision and defined limits.
- Document intelligenceExtraction and classification of information from documents, with human verification where it matters.
Frequently asked questions
Why is overall accuracy not enough?
Because it hides the category or variant where the system fails, which is precisely the question a regulator asks.
What is the accuracy ceiling?
Agreement between human reviewers. Announcing above it is a sign the evaluation is measuring something else.
How do you test adversarially?
By actively seeking inputs that fail, including a different variant, ambiguous inputs and combined requests in one sentence.
How do you assess register?
With blind human assessment by native speakers, with automatic markers only as a complement.
Is evaluation continuous?
It is, because degradation happens without anything changing in the system: data and user behaviour change.
What is a control case?
A known case seeded into the queue, measuring whether the system still gets it right and whether human review is still happening.
Who decides to suspend automation?
The threshold is written beforehand, which avoids the argument during the incident, when the pressure not to suspend is greatest.
Do you produce the regulatory documentation?
We produce the material. Management responsibility stays with the client, because it is not delegable.
Let us look at the numbers for your case
Tell us which processes you want to outsource, in which languages and at what volume. We come back with a euro estimate and an operating design, with no commitment.
We reply within 6 hours on working days. If you would rather write: info@corpshore.solutions