Italy’s EU-funded recovery plan pushed every region toward electronic health records and clinical decision support in emergency departments, but early reporting shows sharply uneven results. Comparing Veneto and Puglia points to an uncomfortable conclusion: what determines success is not the predictive model but the quality of the data feeding it and how nurses actually use it at the triage desk.
Italy’s EU-funded recovery plan pushed every region toward electronic health records and clinical decision support in emergency departments, but early reporting shows sharply uneven results. Comparing Veneto and Puglia points to an uncomfortable conclusion: what determines success is not the predictive model but the quality of the data feeding it and how nurses actually use it at the triage desk.
At the triage desk of an Italian emergency department, the nurse receiving a patient has a few minutes to assign a colour code. For some years now, in a growing number of hospitals, something else appears on the screen alongside vital signs and the assessment scale: a score computed by an artificial intelligence model — a machine learning system trained on the hospital’s record of past admissions — estimating the risk of clinical deterioration or the probability of admission. It decides nothing. It suggests. And in that gap between suggestion and decision lies the most interesting — and least discussed — part of Italy’s publicly funded push into digital health.
Some context for readers outside Italy. The country’s National Recovery and Resilience Plan, financed through the EU’s post-pandemic recovery instrument, includes a dedicated health mission with two linked goals: making the Fascicolo Sanitario Elettronico (the regional Electronic Health Record) genuinely usable everywhere, and equipping facilities with telemedicine and decision-support tools. Crucially, Italian healthcare is administered at regional level: twenty regions, twenty procurement processes, twenty different starting points. The 2026 deadline is approaching, first progress reports are circulating, and the picture is not a simple north-south divide. It is more fragmented than that.
The record is not an archive, it is a data infrastructure
The first misconception concerns what an electronic health record actually is. In public communication it is presented as a folder: lab reports, prescriptions, discharge letters, accessible to citizens through the national digital identity system. From the point of view of a triage algorithm, however, a folder full of PDFs is close to useless.
A model estimating the risk of a patient walking into an emergency department needs structured information: coded chronic conditions, active medications, previous visits, machine-readable test results. If the clinical history exists only as a scanned document or as unnormalised free text, the decision-support system works essentially blind, relying only on what the nurse types in that moment.
This is where regional differences stop being cosmetic. Veneto arrived at this moment with a historical advantage: a regional data-collection infrastructure built over years, a stable patient registry, and a tradition of integration between local health authorities that long predates the recovery plan. That is not a credit to the plan — it is the reason the plan had something to stand on there. Puglia, by contrast, had to run two construction sites at once: building interoperability and installing the tools that presuppose it, against the same deadline.
No predictive model compensates for data that isn’t there. Software cannot distinguish between “patient with no chronic conditions” and “patient whose chronic conditions were never coded.”
What triage support software actually does
It is worth being precise, because public debate swings between “a robot deciding who gets treated” and “another useless piece of software.” The technical reality is more modest and more interesting. What the label “decision support” covers is several distinct families of machine learning models, learning from different data and differing widely in how dependable they are.
- Deterioration scores: models combining vital signs, age and history to estimate the probability that a patient’s condition will worsen in the following hours. They target the emergency department’s worst-case scenario — the underestimated patient who crashes in the waiting room.
- Demand forecasting: expected arrivals by time slot, used for staff rostering and bed management. Paradoxically the most reliable component, because it works on aggregate historical data rather than on individual patients.
- Free-text extraction: systems that read the nurse’s symptom description and derive standardised codes. Useful, but sensitive to abbreviations, local jargon and typing errors — and Italian clinical shorthand varies noticeably between regions.
- Inconsistency alerts: warnings about drug interactions or recorded allergies. Here the link to the health record is direct and the benefit immediate — when the record is populated.
None of these functions assigns a colour code autonomously. Responsibility stays with the clinician, which is also consistent with the European regulatory framework placing AI systems in healthcare in a high-risk category, with obligations around human oversight and traceability.
The problem no tender funds: actual use
Anyone who has worked with clinical alerting systems knows the phenomenon of alert fatigue. When software flags too much, and especially when it flags badly, staff learn to close the window without reading it. The system stays switched on, logs record its operation, the progress report classifies it as active. In practice, it does not exist.
This is the blind spot of recovery-plan reporting, which by design measures verifiable outputs: number of facilities equipped, percentage of documents uploaded to the record, training modules delivered. Legitimate indicators, but they capture none of what matters clinically — whether time-to-treatment for at-risk patients actually fell, and whether assigned codes became more accurate.
The difference between an implementation that works and one that exists on paper turns on things that are mundane and expensive: who stays on the ward in the first weeks explaining how to read the score, who collects nurses’ feedback when the model is systematically wrong about a certain patient type, who has the mandate to recalibrate the algorithm on the local population. These items rarely appear in a procurement specification.
Why a different population needs a different model
There is a technical point that explains why the same product performs differently in two regions. A risk model trained on a given case mix reflects the distribution of that case mix: disease prevalence, age structure, patterns of emergency-department use, availability of primary care.
Veneto and Puglia differ along all these axes. The extent to which people use emergency departments for problems that primary care could handle radically changes the composition of low-acuity presentations. A model calibrated where community medicine absorbs most low-intensity demand will, once moved to a context where that filter is weaker, produce a flood of uninformative flags. Not because it is a poor model — because it was tuned on another reality.
The operational consequence is that local recalibration is not a post-installation extra, it is part of the system. And it requires skills — a clinical lead who understands data, a data scientist who understands emergency medicine — that local health authorities struggle to keep in-house and that capital funding covers poorly.
The finding that reverses the priorities
Here is the least intuitive part of the story. The most tangible benefit reported by people who have worked with these systems almost never concerns the critical patient. Experienced triage nurses are very good at spotting those, and an algorithm adds little.
The real gain sits at the opposite end: apparently stable patients with complex histories, where a properly populated record surfaces in seconds a recent admission, an anticoagulant therapy, repeated visits in preceding weeks. Information the nurse would otherwise extract only by questioning a confused patient or a relative who happens to be present.
Which inverts the investment hierarchy. The predictive model is the visible, sellable part; the systematic coding of clinical information upstream — in wards and in general practitioners’ offices — is the invisible, laborious part. The first can be installed through a tender. The second is built year after year, and it does not end in 2026.