Artificial intelligence learns from examples. This simple statement carries an important consequence: the examples we choose influence what a model can recognize.

In healthcare, differences in population, equipment, protocols, and access are part of the landscape. A solution developed far from this context may reach its limits when it enters daily practice.

Regionalization is more than translation

Adapting a system to Brazil goes beyond language. It means observing population diversity, epidemiological characteristics, and the many healthcare structures across the country.

It also means understanding how professionals work, what information is available, and which constraints must be respected.

What changes between where the model learned and where it runs

In medical imaging, the difference between two services is not abstract: it is recorded in the file itself.

The equipment manufacturer and model change. Magnetic field strength changes. The CT reconstruction kernel changes. Slice thickness, repetition time, contrast phase, whether a particular sequence is acquired at all: these vary between protocols, and each variation alters the appearance of the image the model receives.

The population changes too. Age, body habitus, the prevalence of the conditions being investigated, the point at which people reach the exam — in a service that receives late referrals, disease appears at a different stage than it does in routine screening.

A model learns the distribution it was shown, not disease in the abstract. When that distribution changes, performance can fall without anything in the system reporting a failure. It is the most uncomfortable kind of error: the model keeps answering with its usual confidence, and what changed never appears on screen.

The same accuracy produces different results

There is a common confusion between two questions that look alike and are not.

The first is about the model: among the exams where the condition is present, how many does it identify? Among the normal ones, how many does it leave alone? Those are sensitivity and specificity, measured in the validation study.

The second is about the decision: this positive result, here in front of me, how likely is it to be true? That answer does not depend on the model alone. It also depends on how common the condition is in the population passing through that service.

The rarer the condition, the larger the share of false alarms among positive results — even with the model working exactly as measured. A system validated at a referral center, where cases arrive already selected, can produce a flood of unconfirmed positives when applied to a general population. Nothing broke. The question changed.

Research begins with listening

Questions come before the model. Which problem deserves attention? Who will the technology support? How will results be evaluated? What happens under uncertainty?

These answers do not belong to a single discipline. They require collaboration among healthcare professionals, researchers, data specialists, and affected communities.

Responsible AI does not arrive ready-made for a context. It is built in dialogue with it.

Who labeled the examples, and by what criteria

Every supervised model learns to reproduce a label. Before asking what it got right, it is worth asking what was called right.

Labels extracted automatically from report text carry the vocabulary and the hesitations of whoever wrote them: “cannot be excluded,” “probable,” “consistent with.” Turning that into a yes-or-no column requires decisions that are rarely explained afterwards.

Labels produced by dedicated re-reading are more consistent, but they depend on how many people read each exam and on what was done when they disagreed. And they do disagree: variation between observers is a known fact in radiology, not a defect of the team.

What the model learns, then, is not the truth. It is the judgment of a specific group of people, with their conventions and their blind spots, frozen into a table. Knowing who that group was is part of knowing where the model applies.

Working in the study and working in the service

There is a wide gap between a model that performs well on a held-out portion of the same dataset and a model that performs well at another service, with other equipment and another population.

The first measurement says the model learned something beyond memorizing. The second says whether that something survives away from home. They are different questions, and the second is the one that matters to whoever is going to deploy it.

It is also worth separating what was measured looking backwards, on archived exams, from what was measured following exams as they happened. Retrospective evaluation is cheaper and faster, and it is a good first filter. But it does not show what happens when the model’s output arrives in time to influence a decision — which is precisely the moment the tool starts having consequences.

After deployment, the work continues

A model is not a finished construction. It is a component that ages along with the service it lives in.

Equipment gets replaced. Protocols get revised. The demand profile shifts after a new insurance agreement or a screening campaign. Each of these changes nudges the distribution the model learned, and the accumulated effect can be silent.

Watching for it requires a few simple, continuous things: observing whether the proportion of positive results changes without a clinical explanation; periodically comparing the model’s answers against what was later confirmed; recording when a professional disagrees, in a way that makes disagreeing easy and leaves a trace; and keeping documented which population, which equipment, and which protocols that model was validated on, so the question “does it apply here?” has somewhere to be answered.

Support, observe, and learn

In healthcare, responsible technology must make its role and limits clear. This includes supervision, continuous monitoring, and room for professionals to question results.

Regionalization is an ongoing learning process. The goal is not a universal promise, but tools that are more attentive to the reality where they will be used — and to the people who will live with their effects.