Skip to content

MajorPeer-reviewed

Accelerating scientific discovery with Co-Scientist

A multi-agent AI system on Gemini that generates and validates scientific hypotheses

ByJuraj Gottweis, Wei‐Hung Weng, Alexander Daryin, Tao Tu, Petar Sirkovic, Anatoly Myaskovsky and 2 others

Google (Switzerland) · Google (United States) · Stanford Medicine · Stanford University and 2 more

What it is

Co-Scientist is a multi-agent AI system built on Gemini in which agents continuously generate, critique, and refine research hypotheses, with quality improving as test-time compute is scaled. Its two stated contributions are (1) a multi-agent architecture with an asynchronous task-execution framework for flexible compute scaling and (2) a tournament evolution process for self-improving hypothesis generation. The authors validate the system across three biomedical applications (drug repurposing, novel-target discovery, and explaining antimicrobial-resistance mechanisms), where it identified drug-repurposing candidates and synergistic combination therapies for acute myeloid leukaemia that were confirmed in vitro.

Why it matters

Automated hypothesis generation has mostly produced plausible but unverified ideas; here the outputs were carried through to in vitro experiments in acute myeloid leukaemia, tying an AI system's proposals to wet-lab confirmation. The reported continued benefit of test-time compute scaling means hypothesis quality improved with more deliberation rather than plateauing.

Underlined numbers link to their source. Every metric and quoted figure is listed under Sources and data below.

Filed undermulti-agent systems, hypothesis generation, gemini, drug repurposing, test-time compute

Scored 2026-09-30How scoring works

The Frontier Score, in full

How this entry's score is built, term by term. Open any pillar to see each sub-metric's raw value, how it maps to 0 to 100, and the record that holds it.
Evidence 94Impact 87Novelty 72
The three pillars as one shape: a result strong on every axis fills the triangle.
How the Frontier Score is calculated for this breakthrough
TermScoreWeightPoints
EvidenceHow established and verifiable the result is.94× 0.2523.4
ImpactHow much the result matters, judged field-relative and by current momentum.87× 0.2521.7
NoveltyHow genuinely new the result is, and how fast-moving.72× 0.5036.2
Weighted pillars81.2
Frontier recency4 months since publication× 0.84
IntegrityNot retracted× 1
Frontier Score68
Confidence: Firming up. Gaining evidence: about 70 citations so far, its trajectory is still forming. Corroborated beyond citations (cited on Wikipedia).Gaining evidence: about 70 citations so far, its trajectory is still forming. Corroborated beyond citations (cited on Wikipedia).

Every sub-metric, traced to its source

Evidence94
Evidence sub-metrics
Sub-metricRaw value0 to 100WeightPointsSource
Peer review & venueWhether the work has cleared peer review, and how selective its venue is. Peer review is necessary but not sufficient: before citations accrue, a fresh result in one of the five general flagships (Nature, Science, Cell, NEJM, the Lancet) is a stronger evidence signal than one in a selective specialist or society journal, which in turn outranks one in a legitimate but very high-volume mega-journal, and any of them outranks a preprint that has not been reviewed at all.Maps to 0 to 100: Peer-reviewed article graded by venue tier: flagship journal (Nature, Science, Cell, NEJM, the Lancet) 100; selective journal (the specialist journals of those families, JAMA, PNAS, Physical Review Letters and X) 90; high-volume mega-journal (Nature Communications, Science Advances) 78. Review 90, book 70, dataset 60, preprint 45, other 55.Peer-reviewed, flagship journal1000.3030.0OpenAlex type / primary_location / source
CorroborationIndependent corroboration that the result is real and being taken up, read as the BREADTH of agreement rather than the size of the citation pile. Counts how many of the five independent citation indices (OpenAlex, Crossref, Semantic Scholar, OpenCitations, Europe PMC) report the work at all, plus orthogonal, non-citation lines of corroboration: entry into the encyclopedia, registered clinical trials, released datasets or software, public funding on record, and technical-community discussion. Citation MAGNITUDE is deliberately judged under Impact, not here, so Evidence and Impact measure genuinely different things instead of both rewarding the same citation pile twice.Maps to 0 to 100: Starts at 28 and rises with each of the five independent citation indices that agree the work is cited (+9 each) and each orthogonal non-citation corroboration line (+6 each: Wikipedia, clinical trials, open datasets or software, public funding, community discussion). Capped at 100.5 of 5 citation indices agree, 2 independent non-citation lines850.3025.5OpenAlex, Crossref, Semantic Scholar, OpenCitations, Europe PMC, DataCite, NIH RePORTER, Wikipedia, Hacker News count of agreeing indices + orthogonal reach
VerifiabilityHow openly the result can be read, reused, and reproduced. Rewards open access, a permissive reuse license (CC-BY/CC0), a green repository copy anyone can archive, openly minable full text, and released datasets or software a reader can actually run. Cross-checked across OpenAlex, Unpaywall, Europe PMC, and DataCite.Maps to 0 to 100: Closed 45; open access 80+; a permissive CC license 100; a repository copy, open full text, or a released dataset/software artifact each raise the floor.Open access (hybrid), repository copy900.2018.0OpenAlex, Unpaywall, Europe PMC, DataCite is_oa / license / has_repository_copy / open datasets
IntegrityRetraction and post-publication correction status. A retracted result is not evidence of anything and collapses the whole Frontier Score; a correction or expression of concern is a softer flag. Read across OpenAlex and Crossref (both update-to and updated-by notices).Maps to 0 to 100: Clean 100; a correction or expression of concern 75; retracted 0. Flagged if OpenAlex OR Crossref records the notice.No retraction or concern on record1000.2020.0OpenAlex, Crossref is_retracted / update-to / updated-by
Impact87
Impact sub-metrics
Sub-metricRaw value0 to 100WeightPointsSource
Field-normalized impactImpact relative to the world average for the same field, year, and type, so a small field and a large field are judged fairly. Triangulates three independent field-normalized measures: OpenAlex's Field-Weighted Citation Impact, its citation percentile within the exact field-and-year cohort, and the NIH iCite Relative Citation Ratio. 1.0x is average; the percentile is the share of same-field, same-year work it out-cites.Maps to 0 to 100: Each ratio (1.0 = field average) maps 1.0 to 50, 3.0 to 75, 9.0 to 90; the field+year percentile maps directly (top 1% -> ~99). The available measures are averaged. Falls back to a saturating citation count when none is available yet.397.8x FWCI, 16.3x RCR, top 0.0% of field-year vs field980.4544.1OpenAlex, NIH iCite fwci / citation_normalized_percentile / relative_citation_ratio
Influential citationsCitations that genuinely build on the work rather than mention it in passing (Semantic Scholar's influential-citation measure). A sharper signal of real impact than a raw count.Maps to 0 to 100: Saturating: 15 influential citations maps to 50. Falls back to a fraction of consensus citations when unavailable.33 influential690.2013.8Semantic Scholar influentialCitationCount
Citation velocityCitations received in the trailing twelve months, judged as a rate rather than a lifetime total. Momentum is what a frontier index cares about: a result being taken up fast right now is landing, whether its lifetime pile is large or still small. Where an expected field citation rate is known, the momentum is also read relative to it, so a fast-moving result in a quiet field is not overlooked.Maps to 0 to 100: Saturating: 30 citations in the last twelve months maps to 50. When NIH iCite provides the field's expected rate, this is blended evenly with the ratio of observed to expected momentum.89 in last 12 months (8.4x field rate)820.3528.7OpenAlex, NIH iCite counts_by_year / field_citation_rate
Novelty72
Novelty sub-metrics
Sub-metricRaw value0 to 100WeightPointsSource
Conceptual noveltyHow genuinely new the contribution is, measured at publication from the work's OWN content, not its date: how atypical its combination of research areas is versus all prior work, and whether it builds on a broad, deep base rather than only the newest papers in a fast-moving area. This is what separates a real advance from a recent-but-commodity result, and it needs no forward citations, so it is stable for brand-new work. An honestly noisy proxy (even state-of-the-art novelty measures agree only moderately with expert judgment), so it carries bounded weight and never drives the ranking alone.Maps to 0 to 100: Blends two publication-time signals, each 0-100: an atypical-combination score (how unusual the pairing of the work's research topics is versus the prior literature, by pointwise mutual information) and a foundational-reach score (median age and depth of the work's references). The stronger signal weighs 0.62, the weaker 0.38. When neither is measurable (a work too thin in topics or references), the pillar renormalizes over its recency signals instead.Genomics and Rare Diseases x Cell Image Analysis Techniques (114 prior works made this pairing)630.4528.1OpenAlex topics + referenced_works (publication-time)
RecencyHow recently the work appeared. The frontier is now, so newer scores higher.Maps to 0 to 100: Exponential decay with a 12-month half-life from the publication date.4 months old780.3023.3OpenAlex publication_date
Attention accelerationWhether attention is rising: citations in the latest full year versus the year before. Accelerating interest signals an active, opening frontier.Maps to 0 to 100: Latest/prior-year ratio: 1x maps to 50, 2x to 75, 4x to 100. Unknown is neutral.3.0x rising900.109.0OpenAlex counts_by_year
Edge of reviewPreprints and brand-new work sit ahead of the peer-review process. That earns novelty here (while it is discounted under Evidence), because it is where the frontier forms first.Maps to 0 to 100: Preprint 95, article <6 months 80, article <18 months 65, else 45.Peer-reviewed record800.1512.0OpenAlex type / publication_date
ProvenanceAll sources

Sources and data

Every number behind this entry, grouped by the source that holds it: up to five citation indices cross-checked, field-normalized impact, then real-world reach (Wikipedia, community discussion, clinical trials, released data and public funding) where it exists. Each value links to its record.

OpenAlex

Primary record
Total citations
96
Field-weighted citation impact
397.8x field average
Citation percentile
Top 0.1% of its field-year
Citations, recent 12 months
89
Publication date
2026-05-19

Crossref

Citations
70
Works it builds on
37

Semantic Scholar

Citations, all versions
494
Influential citations
33

OpenCitations

Citations
0

Europe PMC

Citations
36

NIH iCite

Relative Citation Ratio
16.26x field average
Potential to translate
75%

Unpaywall

Reuse license
OTHER-OA

Wikipedia

Wikipedia articles citing it
1

Figures quoted in the write-up

FigureValueSource
first key contribution: multi-agent architecture with asynchronous task execution for flexible compute scaling(1)Nature
second key contribution: tournament evolution process for self-improving hypothesis generation(2)Nature
biomedical application areas used for validationthreeNature
AI & ComputingThe whole field

Back to the Frontier Index, or browse every breakthrough by date.

The Frontier Brief

The week's frontier, scored.

Every Monday, the highest-scoring new breakthroughs in the index, each with its score and a one-line read on why it matters. What you get

One email a week. The top new breakthroughs, scored. Unsubscribe anytime.