Skip to content

Artificial intelligence

Artificial intelligence as a managed service for the mid-market

Most mid-sized companies do not need a data science team. They need two or three automated processes, with someone accountable when they fail.

Corpshore Portugal editorial team

Written by the team that builds these operations. No individual byline: this is internally reviewed work, not personal opinion.

Published

Artificial intelligence as a managed service for the mid-market

Why do in-house projects stall?

Because the prototype is the easy part. What fails afterwards is the operation: who corrects it when the model is wrong, who reassesses when the data shifts, who answers the customer who got the wrong reply, and who documents all of it for the regulator. None of those questions is settled by a demo.

In a managed service, those responsibilities are assigned by contract before the first case passes through the system. That is the main difference, and it is worth more than any accuracy gain.

Where do you start without taking risk?

With triage and classification work, where an error is cheap and visible: routing messages by subject, extracting fields from documents, proposing replies that a person approves. Human review always stays on cases affecting customer rights.

The European Union Artificial Intelligence Regulation applies in phases and classifies systems by risk. For most of these cases the risk is limited, but transparency obligations apply: the user has to know they are interacting with an automated system.

How do you tell whether it is working?

By total process time and reopening rate, not by percentage of cases automated. Automating eighty per cent and reopening half of them is worse than automating forty and closing nearly all.

It is also worth measuring the time supervision consumes. A system that automates half the volume and requires a review layer the size of the previous team has released no capacity at all, it has only changed the kind of work, and that only shows if someone counts the hours on both sides.

Which processes are good candidates?

Processes with high volume, stable rules and an outcome verifiable within hours. Routing messages by subject, extracting fields from invoices and delivery notes, classifying requests by type, and proposing answers to frequent questions with human approval before they go out.

The criterion that matters most is the cost of an error. If a wrong classification means a message reaches the wrong team and someone forwards it in two minutes, the error is cheap and visible and the process is a good candidate. If it means a customer receives a wrong decision about an entitlement, it is not.

The second criterion is having historical examples. A process that has run for two years with a record of what was decided has evaluation material. A new process does not, and starting by automating something never done manually is solving two unknown problems at once.

How do you build the evaluation set?

From real cases already decided, chosen to cover the real distribution rather than the convenient one. Two to five hundred cases is enough for most mid-market processes, provided they include the borderline cases in the proportion they actually occur.

The part done badly is labelling. If two experienced people disagree on the correct decision in fifteen per cent of cases, no system will do better than that, and the number is valuable information: it measures how much of the process is genuinely ambiguous. It is worth measuring before measuring the system.

The evaluation set should be built before any automation work and kept separate. A set used to tune the system no longer serves to assess it, and the temptation to reuse it is strongest precisely when results start looking good.

What does the AI Regulation require in practice?

Risk classification, first. Most mid-market cases fall under limited risk, where the central obligation is transparency: the user has to know they are interacting with an automated system, and artificially generated content has to be identifiable as such.

Some cases rise to high risk, particularly where the system participates in decisions about employment, access to essential services, credit or the assessment of people. There, substantial obligations are added on risk management, data quality, technical documentation, event logging and effective human oversight.

Effective human oversight is the phrase most often met only on paper. A person approving three hundred proposals an hour is not supervising, they are rubber-stamping. If the human rejection rate is near zero, either the system is perfect or the oversight is not real, and the first possibility is rare.

What does it cost and when do you see a return?

For two or three mid-market processes, setup typically takes between eight and sixteen weeks: four to six describing the process and building the evaluation set, four to six building and tuning, and the rest running assisted before supervision is reduced.

The return rarely comes from headcount reduction and almost always from cycle time and released capacity. A process that took two days and now takes two hours changes what the business can promise, and that is worth more than the direct cost difference in most cases.

The recurring cost people forget is maintenance. Data changes, processes change, and a system that is not reassessed degrades silently. Budgeting zero for maintenance is the usual way to have an excellent project in year one and a problem in year three.

What signals say you should stop?

Reopening rate rising while automation rate rises. It is the clearest sign that the system is closing cases that are not resolved, and it is invisible on any dashboard that measures only automated volume. Automating forty per cent and closing nearly all of it beats automating eighty and reopening half.

Complaints clustering on one case type. If grievances group into a pattern, the system is systematically wrong on that pattern, and the correct response is to suspend automation for that case type while investigating, not to adjust a global threshold and hope.

And human reviewers who have stopped disagreeing. When the human correction rate falls sharply with nothing having changed in the system, it is usually not the system that improved: it is the oversight that became routine. It is worth seeding known control cases to measure whether review is still happening.

How is human oversight designed?

With an explicit rule on which cases require approval before going out, written by case type rather than by system confidence threshold. Confidence thresholds are useful and are not a defensible criterion on their own: a system can be very confident and very wrong, and the difference is not observable from outside.

With enough time for the review to be real. If the operating design implies each reviewer approving one proposal every ten seconds, the oversight is nominal. Sizing the review layer by how long an honest review takes is what separates compliance from theatre.

With real power to reject and escalate. A reviewer who can correct but cannot stop automation for a case type is not supervising the system, they are cleaning up after it. The authority to suspend has to exist and has to be usable without personal cost.

And with measurement of the oversight itself. Known control cases seeded into the review queue measure whether review is actually happening. It is the only way to detect drift towards rubber-stamping, which always happens and never announces itself.

What documentation has to exist?

A description of the system and its purpose, in language someone outside the technical team can assess. If the documentation is only comprehensible to whoever built the system, it serves neither management oversight nor a regulator, which are the two audiences that matter.

The provenance of the data used, with a legal basis for each use, and a description of the evaluation set and the results obtained per case category. Aggregate results are not enough: the question a regulator asks is about the group where the system performs worst.

The event log, with defined retention, allowing reconstruction of why a specific case was decided as it was. Without it, answering an individual challenge is impossible, and an individual challenge is the most likely scenario of all.

And a change log: when the system was changed, by whom, why, and what the measured effect was. A system that changes without a record makes any historical analysis invalid, and the first time that matters is always during an investigation.

Frequently asked questions

Do I need data scientists?
For most mid-market cases, no. You need defined processes and assigned accountability.
Does the AI Regulation apply to me?
It does, with obligations proportionate to risk. Transparency is the most common.
Should there always be human review?
On cases affecting customer rights, yes, and that should be written down.
What is the right metric?
Total process time and reopening rate, not percentage automated.
Who answers when the system is wrong?
In a managed service, the supplier, with accountability assigned by contract before opening.
Which cases do you start with?
Triage and classification, where an error is cheap and visible and correction is immediate.
Does the user have to know it is automated?
They do, and that is the AI Regulation's most common transparency obligation.
When is the system reassessed?
When the data or the process changes, and at a periodic review set in the contract.

Let us look at the numbers for your case

Tell us which processes you want to outsource, in which languages and at what volume. We come back with a euro estimate and an operating design, with no commitment.

We reply within 6 hours on working days. If you would rather write: info@corpshore.solutions