Skip to content

NotablePeer-reviewed

Clinical usability of an explainable AI decision support tool and evaluation of multimodal models in NSCLC

ByArsela Prelaj, V. Miskovic, Matteo Sacco, Alberto Ferrarin, Cristina Licciardello, Leonardo Provenzano and 2 others

Fondazione IRCCS Istituto Nazionale dei Tumori · Politecnico di Milano · University of Chicago · European Institute of Oncology and 4 more

What it is

The I3LUNG study enrolled 2,396 patients to test whether AI can guide immunotherapy selection in non-small cell lung cancer, integrating clinical and blood data, CT images, digital pathology, and genomics into early-fusion and intermediate-fusion models. Machine-learning and deep-learning models using clinical-and-blood data alone reached an area under the curve of up to 0.77 in the test set and significantly surpassed PD-L1, ECOG performance status, neutrophil-to-lymphocyte ratio, LDH, and the Lung Immune Prognostic Index. In external validation the AUC dropped to a range of 0.55 to 0.72, which the authors attribute to population differences, and in a usability study both expert and nonexpert physicians improved their predictions using the explainable AI tool.

Why it matters

Immunotherapy selection in NSCLC still rests on imperfect PD-L1 and clinical scores, so a tool that outperforms those markers and, importantly, measurably improves physicians' own predictions addresses a real decision-making gap. The authors report that multimodal integration did not clearly beat the clinical-and-blood-only model, a sober finding that keeps the practical claim tied to the simpler, more deployable data.

Underlined numbers link to their source. Every metric and quoted figure is listed under Sources and data below.

Filed underRadiomics and Machine Learning in Medical Imaging, Cancer Immunotherapy and Biomarkers, Lung Cancer Diagnosis and Treatment

Scored 2026-09-28How scoring works

The Frontier Score, in full

How this entry's score is built, term by term. Open any pillar to see each sub-metric's raw value, how it maps to 0 to 100, and the record that holds it.
Evidence 91Impact 42Novelty 61
The three pillars as one shape: a result strong on every axis fills the triangle.
How the Frontier Score is calculated for this breakthrough
TermScoreWeightPoints
EvidenceHow established and verifiable the result is.91× 0.2522.7
ImpactHow much the result matters, judged field-relative and by current momentum.42× 0.2510.5
NoveltyHow genuinely new the result is, and how fast-moving.61× 0.5030.7
Weighted pillars63.8
Frontier recency1 month since publication× 0.96
IntegrityNot retracted× 1
Frontier Score61
Confidence: Early signal. Early signal: little citation history yet, so impact is an estimate and may shift as the work is taken up.Early signal: little citation history yet, so impact is an estimate and may shift as the work is taken up.

Every sub-metric, traced to its source

Evidence91
Evidence sub-metrics
Sub-metricRaw value0 to 100WeightPointsSource
Peer review & venueWhether the work has cleared peer review, and how selective its venue is. Peer review is necessary but not sufficient: before citations accrue, a fresh result in the most selective venues (Nature, Science, Cell, NEJM, the Lancet, PNAS, Physical Review Letters/X) is a stronger evidence signal than one in a legitimate but very high-volume mega-journal, which in turn outranks a preprint that has not been reviewed at all.Maps to 0 to 100: Peer-reviewed article graded by venue prestige: flagship 100, elite-family 90, high-volume mega-journal 78. Review 90, book 70, dataset 60, preprint 45, other 55.Peer-reviewed900.3027.0OpenAlex type / primary_location / source
CorroborationIndependent corroboration that the result is real and being taken up, read as the BREADTH of agreement rather than the size of the citation pile. Counts how many of the five independent citation indices (OpenAlex, Crossref, Semantic Scholar, OpenCitations, Europe PMC) report the work at all, plus orthogonal, non-citation lines of corroboration: entry into the encyclopedia, registered clinical trials, released datasets or software, public funding on record, and technical-community discussion. Citation MAGNITUDE is deliberately judged under Impact, not here, so Evidence and Impact measure genuinely different things instead of both rewarding the same citation pile twice.Maps to 0 to 100: Starts at 28 and rises with each of the five independent citation indices that agree the work is cited (+9 each) and each orthogonal non-citation corroboration line (+6 each: Wikipedia, clinical trials, open datasets or software, public funding, community discussion). Capped at 100.5 of 5 citation indices agree, 1 independent non-citation line790.3023.7OpenAlex, Crossref, Semantic Scholar, OpenCitations, Europe PMC, DataCite, NIH RePORTER, Wikipedia, Hacker News count of agreeing indices + orthogonal reach
VerifiabilityHow openly the result can be read, reused, and reproduced. Rewards open access, a permissive reuse license (CC-BY/CC0), a green repository copy anyone can archive, openly minable full text, and released datasets or software a reader can actually run. Cross-checked across OpenAlex, Unpaywall, Europe PMC, and DataCite.Maps to 0 to 100: Closed 45; open access 80+; a permissive CC license 100; a repository copy, open full text, or a released dataset/software artifact each raise the floor.Open access (hybrid), CC-BY, repository copy1000.2020.0OpenAlex, Unpaywall, Europe PMC, DataCite is_oa / license / has_repository_copy / open datasets
IntegrityRetraction and post-publication correction status. A retracted result is not evidence of anything and collapses the whole Frontier Score; a correction or expression of concern is a softer flag. Read across OpenAlex and Crossref (both update-to and updated-by notices).Maps to 0 to 100: Clean 100; a correction or expression of concern 75; retracted 0. Flagged if OpenAlex OR Crossref records the notice.No retraction or concern on record1000.2020.0OpenAlex, Crossref is_retracted / update-to / updated-by
Impact42
Impact sub-metrics
Sub-metricRaw value0 to 100WeightPointsSource
Field-normalized impactImpact relative to the world average for the same field, year, and type, so a small field and a large field are judged fairly. Triangulates three independent field-normalized measures: OpenAlex's Field-Weighted Citation Impact, its citation percentile within the exact field-and-year cohort, and the NIH iCite Relative Citation Ratio. 1.0x is average; the percentile is the share of same-field, same-year work it out-cites.Maps to 0 to 100: Each ratio (1.0 = field average) maps 1.0 to 50, 3.0 to 75, 9.0 to 90; the field+year percentile maps directly (top 1% -> ~99). The available measures are averaged. Falls back to a saturating citation count when none is available yet.5.1x FWCI, top 4% of field-year vs field900.4540.5OpenAlex, NIH iCite fwci / citation_normalized_percentile / relative_citation_ratio
Influential citationsCitations that genuinely build on the work rather than mention it in passing (Semantic Scholar's influential-citation measure). A sharper signal of real impact than a raw count.Maps to 0 to 100: Saturating: 15 influential citations maps to 50. Falls back to a fraction of consensus citations when unavailable.0 influential00.200.0Semantic Scholar influentialCitationCount
Citation velocityCitations received in the trailing twelve months, judged as a rate rather than a lifetime total. Momentum is what a frontier index cares about: a result being taken up fast right now is landing, whether its lifetime pile is large or still small. Where an expected field citation rate is known, the momentum is also read relative to it, so a fast-moving result in a quiet field is not overlooked.Maps to 0 to 100: Saturating: 30 citations in the last twelve months maps to 50. When NIH iCite provides the field's expected rate, this is blended evenly with the ratio of observed to expected momentum.1 in last 12 months (0.0x field rate)40.351.3OpenAlex, NIH iCite counts_by_year / field_citation_rate
Novelty61
Novelty sub-metrics
Sub-metricRaw value0 to 100WeightPointsSource
Conceptual noveltyHow genuinely new the contribution is, measured at publication from the work's OWN content, not its date: how atypical its combination of research areas is versus all prior work, and whether it builds on a broad, deep base rather than only the newest papers in a fast-moving area. This is what separates a real advance from a recent-but-commodity result, and it needs no forward citations, so it is stable for brand-new work. An honestly noisy proxy (even state-of-the-art novelty measures agree only moderately with expert judgment), so it carries bounded weight and never drives the ranking alone.Maps to 0 to 100: Blends two publication-time signals, each 0-100: an atypical-combination score (how unusual the pairing of the work's research topics is versus the prior literature, by pointwise mutual information) and a foundational-reach score (median age and depth of the work's references). The stronger signal weighs 0.62, the weaker 0.38. When neither is measurable (a work too thin in topics or references), the pillar renormalizes over its recency signals instead.Radiomics and Machine Learning in Medical Imaging x Cancer Immunotherapy and Biomarkers (5,247 prior works made this pairing); builds on refs a median 2.8y deep350.4515.9OpenAlex topics + referenced_works (publication-time)
RecencyHow recently the work appeared. The frontier is now, so newer scores higher.Maps to 0 to 100: Exponential decay with a 12-month half-life from the publication date.This month950.3028.5OpenAlex publication_date
Attention accelerationWhether attention is rising: citations in the latest full year versus the year before. Accelerating interest signals an active, opening frontier.Maps to 0 to 100: Latest/prior-year ratio: 1x maps to 50, 2x to 75, 4x to 100. Unknown is neutral.Not enough history500.105.0OpenAlex counts_by_year
Edge of reviewPreprints and brand-new work sit ahead of the peer-review process. That earns novelty here (while it is discounted under Evidence), because it is where the frontier forms first.Maps to 0 to 100: Preprint 95, article <6 months 80, article <18 months 65, else 45.Peer-reviewed record800.1512.0OpenAlex type / publication_date
ProvenanceAll sources

Sources and data

Every number behind this entry, grouped by the source that holds it: up to five citation indices cross-checked, field-normalized impact, then real-world reach (Wikipedia, community discussion, clinical trials, released data and public funding) where it exists. Each value links to its record.

OpenAlex

Primary record
Total citations
1
Field-weighted citation impact
5.1x field average
Citation percentile
Top 4% of its field-year
Citations, recent 12 months
1
Publication date
2026-09-01

Crossref

Citations
1
Works it builds on
38
Funders on record
77

Semantic Scholar

Citations, all versions
1
Influential citations
0

OpenCitations

Citations
0

Europe PMC

Citations
1

NIH iCite

Potential to translate
5%

Unpaywall

Reuse license
CC-BY

Figures quoted in the write-up

FigureValueSource
Patients enrolled (International real-world multimodal AI study (I3LUNG))2,396Nature Medicine
Best test-set AUC (Clinical-and-blood-only ML and DL models in the test set)up to 0.77Nature Medicine
AI & ComputingThe whole field

Back to the Frontier Index, or browse every breakthrough by date.

The Frontier Brief

The week's frontier, scored.

Every Monday, the highest-scoring new breakthroughs in the index, each with its score and a one-line read on why it matters. What you get

One email a week. The top new breakthroughs, scored. Unsubscribe anytime.