Skip to content

54 min read

The Latest Scientific Breakthroughs of 2026, Scored and Ranked

An evidence-scored ranking of the latest and biggest scientific breakthroughs of 2026, drawn from Frontier's public index of 101 results across eight fields of science.

By Yuma Heymans. Scores quoted are live Frontier Scores; see how they are computed.

A year of science, weighed against the evidence instead of the mood of the room.

The Frontier Index closed its 2026 view with 101 published breakthroughs, each carrying one transparent number, and the highest of them reaches only 70 out of 100. That gap between the best result of the year and the top of the scale is not a disappointment. It is the entire point. A perfect 100 would demand a discovery that is at once ironclad, field-changing, and wholly without precedent, and nothing in the past year clears all three bars together. The score you see is therefore an argument, not a trophy, and every part of it traces back to a public record you can pull yourself.

Most "breakthrough of the year" lists are, by their own publishers' descriptions, works of editorial taste. A staff debates, a committee deliberates, a shortlist emerges, and you are handed a curated opinion to trust. That method has genuine value and a hard limit: it is fast and readable, and it is impossible to check. Frontier was built to invert it. Every entry is scored against public evidence signals with a published, reproducible method, so a ranking can be examined line by line and disagreed with on the merits rather than accepted on authority. - Science's own Breakthrough of the Year process

This report is Frontier's own account of what its 2026 data shows. We built the index, so we explain it in the open: what an evidence score actually measures, how the three pillars are assembled from independent sources, which breakthroughs rose to the top and why, how each field of science is faring, and what the numbers reveal once you stop reading them one at a time. It is a state-of-the-frontier piece, honest about its limits, concrete about its methods, and skeptical of its own conclusions wherever the evidence is still thin.

Contents

  1. The shape of the 2026 frontier
  2. How Frontier scores a breakthrough
  3. The breakthroughs that defined the year
  4. Field by field: where each science stands
  5. What the scores actually reveal
  6. The outlook: 2026 into 2027
  7. How to read the index yourself

A note on how to read what follows. Each numbered section stands on its own, so you can begin with the field that interests you and double back. Every figure quoted here is drawn from the live index rather than written by hand, and the underlying table is downloadable in full, so nothing below has to be taken on faith. Where the data is genuinely uncertain we say so plainly rather than rounding the doubt away. The most useful way to read a report like this is with the open dataset open in another tab, checking the claims against the records as you go.

1. The shape of the 2026 frontier

Before any single result, it is worth asking what a breakthrough score is even trying to capture. A discovery matters for reasons that pull in different directions. Some findings are unimpeachably solid yet incremental, others are dazzling yet fragile, and a rare few are solid, consequential, and new all at once. An opinion list collapses those dimensions into one act of editorial judgment and hands you the verdict. An evidence score does the opposite: it holds the dimensions apart, attaches each to public signals anyone can retrieve, and only then combines them, so the final number carries its own reasoning. You are not asked to believe a result is important. You are shown the parts and can redo the arithmetic.

That distinction is not academic. The most recent editorial lists are, by their authors' own admission, best guesses rather than measurements. One prominent technology list calls its selections "educated guesses" and notes candidly that it does not always get them right. - MIT Technology Review's 10 Breakthrough Technologies 2026 There is nothing wrong with an honest guess from expert editors. The problem is that a guess cannot be audited, cannot be re-run when new evidence lands, and gives a reader no way to locate exactly where their judgment and the editors' part ways. A score built from public signals can do all three, which is the difference Frontier exists to make.

Seen through that lens, the 2026 index has a distinct shape. Across the 101 breakthroughs, Frontier Scores run from a low of 21 to a high of 70, with a mean of 45.7 and a median of 45. The near-identical mean and median tell you the distribution is roughly symmetric, without a long tail of outliers dragging the average around. There is no cluster of nineties and no lone entry sprinting away from the pack. Instead there is a densely packed middle and a modest spread, which is precisely what a scale designed to reserve its upper reaches should produce. The frontier of 2026 is broad and busy, not dominated by one towering event, and that breadth is unevenly distributed across the sciences.

Biomedicine is the center of gravity. With 27 entries, Biomedicine and Health holds more than a quarter of the whole index, as many as the next two fields combined, a reflection of how much of 2026's most heavily corroborated work came out of genomics, gene editing, and controlled clinical trials. Neuroscience follows with 14, carried by a run of connectomics and single-cell studies, and AI and Computing sits close behind at 13. The planetary and physical sciences (space, materials, climate, quantum, and physics) each contribute smaller cohorts. This is not a claim that biology is intrinsically more important than physics. It is a measurement of where highly corroborated, high-signal results are currently concentrated, which is a narrower and more honest statement than a ranking of whole disciplines.

Concentration by field, though, says nothing about how high any of these results actually scored. A field can be crowded with solid but middling work, or thin and topped by a standout. To see that, you have to read the curve of scores itself, from the single highest-ranked entry down to the hundred-and-first.

The line descends gently and never touches the top of the scale. The highest-scored entry of the year, an AI co-scientist that proposes and evaluates its own research hypotheses, reaches 70, and it is instructive that even the single best result stops a full thirty points short of the ceiling. - the Co-Scientist study in Nature From there the slope is smooth: the fifteenth-ranked entry sits at 56, the median entry near 45, and the tail settles at 21. There is no cliff and no artificial floor, only a steady, unforced decline. That shape is itself a result. It says the year produced a deep bench of strong science and no runaway singular event, and the reason the curve never reaches 100 is worth stating plainly.

A perfect score would require a result that is at once ironclad, field-changing, and wholly unprecedented, and rigorously scored on every axis. Real science almost never arrives in that form. A landmark clinical trial is ironclad and consequential but seldom conceptually new. A brilliant method is new but needs years of replication before its evidence is beyond question. The ceiling is deliberately left unreached so the scale keeps its meaning as the field advances, rather than saturating the first strong year it meets. When a result eventually does land in the eighties, that number will say something precise, because the top of the scale was never spent cheaply on a merely excellent year.

All of this is inspectable rather than asserted, which is the property that separates a measured index from a confident opinion. The full ranking is live and sortable on the Frontier Index, and every figure behind it, each score, pillar, and source, is published as an open dataset under a permissive license, downloadable as CSV and JSON. The remainder of this report explains how those numbers are built and what they reveal when read together, but the invitation underneath it stays simple: do not take the ranking on trust, take it apart.

2. How Frontier scores a breakthrough

The hard problem behind any index like this is not opinion, it is commensurability. A gene-editing trial, a trapped-ion quantum processor, and a genome foundation model are not obviously measurable on the same axis, and the easy move is to wave a hand and rank them by feel. Frontier's answer is to refuse that shortcut and route every candidate through the same fixed pipeline, so that a neuroscience result and an astronomy result face identical questions, drawn from identical kinds of public evidence, before either receives a number. The score is not a verdict announced at the end. It is the accumulated output of a repeatable process, and that process is the same for all 101 entries, which is what makes two very different results comparable at all.

That process rests on three pillars, each answering a separate question about a result. Evidence asks whether the finding is real and independently corroborated. Impact asks whether it moves its field and the wider world. Novelty asks whether it combines ideas in a genuinely new way. These are deliberately held apart, because a result can be strong on one and weak on another, and collapsing them too early is exactly the error an evidence score exists to prevent. The pillars are computed from up to 13 independent public data sources, refreshed weekly, so scores shift as the evidence around a result accumulates rather than freezing at the moment of publication. The path from raw literature to a ranked entry runs in five stages.

Each stage is mechanical and inspectable rather than a matter of taste. Discovery pulls candidate results from a broad, open catalog of the literature and its companion feeds, so inclusion is driven by what is being published and cited, not by what an editor happened to notice. Enrichment gathers the cross-source signals attached to each work: citation behavior, publication venue, replication traces, and the structure of its references. Scoring turns those signals into the three pillar values. Ranking combines the pillars into a single 0 to 100 Frontier Score. Publishing exposes the entry with every underlying figure linked to its primary record, the step most opinion lists skip entirely. The inputs are not proprietary. They are the ordinary infrastructure of modern scientometrics, and they are named openly at our sources page.

  • OpenAlex, the fully open index of the global research literature, built by a nonprofit as the successor to Microsoft Academic Graph.
  • Crossref and DOI records, anchoring each result to a citable, permanent identifier.
  • NIH iCite, supplying field-normalized citation measures for the biomedical literature.
  • Reference and citation graphs, used to judge how atypical a work's combination of prior ideas really is.

Naming the sources is what makes a score reproducible rather than merely reported. The open catalog at the base of the pipeline is itself a public, documented system that anyone can query. - the OpenAlex paper Anyone who disagrees with where a breakthrough landed can pull the same records, recompute the same pillars, and pinpoint the exact place their judgment departs from Frontier's. That is a fundamentally different posture from a list you are asked to accept. The full source register and the current methodology version are published in the open, so the machinery is available for audit rather than described in the abstract. What the three pillars actually measure is worth seeing laid out side by side.

Each pillar leans on established method rather than a formula invented for the occasion. Evidence is the most straightforward in spirit: it rewards independent corroboration, replication, and the standing of the venue, and it is the reason a result has to be well-sourced before it is admitted at all. Because that bar is applied at the door, the Evidence pillar tends to run uniformly high, and its real job is to keep unverified claims out of the index rather than to separate the strong entries from the merely good ones.

Impact is where the measurement gets more careful, because raw citation counts punish slower-citing fields and flatter faster ones. Frontier leans on field-weighted approaches that normalize a result against the world average for its own discipline, year, and type, so a much-cited genomics paper and a much-cited astronomy paper can be compared on one scale. - the field-weighted citation impact definition

A second, article-level lens complements the first by benchmarking a paper against its own co-citation neighborhood, the specific cluster of work it sits among, rather than against a broad disciplinary average. That approach, developed and validated in the biomedical literature, is what lets the Impact pillar register influence even for results whose field is small or unusually defined. - the Relative Citation Ratio in PLOS Biology

Novelty is the subtlest of the three. It builds on a well-known finding that the highest-impact science tends to pair conventional ideas with atypical ones, which makes genuine newness measurable through the structure of a work's references rather than left to a reviewer's intuition. - the atypical-combinations study in Science A result that fuses two rarely-joined literatures reads as more novel than one that extends a single crowded line of work.

There is a pattern hiding in these three pillars, and it is the single most revealing thing the 2026 data has to say. Because Frontier only admits work that is already well-sourced and corroborated, Evidence scores sit near the top for almost every entry, which means evidence is rarely what separates a high score from a low one. The real spread lives in Impact and Novelty. That asymmetry quietly reshapes how the whole index should be read, and Section 5 takes it apart in full. For now it is enough to notice the setup: when the science is nearly always sound, the interesting question stops being whether a result is true and becomes whether it is also new and consequential.

3. The breakthroughs that defined the year

The index holds 101 entries, but a small group sits clearly above the rest. The ten profiled here run from the year's single highest score, a 70, down to a tight cluster at 60, and together they are what the method ranks as the most consequential science of 2026. Almost none is a household name. Read them in order and a pattern surfaces that no opinion list would show you: the leaders are not the results with the loudest press, they are the ones that pair ironclad evidence with genuine novelty and real reach. Each profile below gives the Frontier Score, the pillar breakdown that produced it, what the work actually did in plain terms, and why it earned its place near the top.

One note on how to read a score before we begin. Every entry is graded on three independent questions, evidence (is it real), impact (does it matter), and novelty (is it new), each on a 0 to 100 scale, and the three fold into the headline number. Because Frontier only admits well-corroborated work, evidence scores cluster near the ceiling for nearly everything here, so the ranking is settled mostly by impact and novelty. Keep that in mind as you go: when one of these entries outranks another, the gap almost always lives in how new or how consequential the result is, not in whether it is true. The full grading rules are public at Frontier's methodology page.

The year's highest score, a 70, belongs to an AI co-scientist, a multi-agent system built on Gemini that does something narrow language tools do not: it proposes original scientific hypotheses, argues against them, and ranks the survivors. Its pillars read evidence 91, impact 85, and novelty 74, and it is the high impact figure, rare in this index, that lifts it clear of the field. The system runs several specialized agents in a kind of tournament, generating ideas, critiquing them, and evolving the strongest, with answers improving as it is given more time to think. What makes the result credible rather than a demo is the wet-lab validation behind it: proposed drug-repurposing candidates showed real activity against a leukemia cell line, its liver-fibrosis suggestions had measurable anti-fibrotic effects in human organoids, and on antibiotic resistance it independently reconstructed a mechanism that a separate lab had found but not yet published.

The announcement framed the work in a stylized illustration of the molecular scales it operates on.

Why it matters is a shift in role, not just capability. For years the honest description of AI in science was that it summarizes and retrieves. This is a system that generates and prioritizes, compressing the slow early stage where a researcher stares at a problem and guesses what to try. The novelty score of 74 reflects that this is a genuinely new use of the tools, and the impact of 85 reflects that partner institutions are already running it. The sober caveat, which the score honors by stopping at 70 rather than approaching the ceiling, is that every proposed hypothesis still has to be tested at the bench by people; the machine narrows the search, it does not close it.

Second, at 68, is a complete diploid human genome benchmark, and it is the clearest case in the top ten of evidence carrying a score. Its pillars are evidence 96, impact 61, and novelty 63: not a flashy result, but an unusually solid one. The plain version is that this is a reference answer key for reading a person's DNA, and for the first time it covers both inherited copies of the genome, including the stretches that older references simply left blank. The benchmark reaches near-perfect accuracy across 99.4% of the sequence, adds 701.4 Mb of new autosomal DNA plus 216.8 Mb of both sex chromosomes, and fills in the roughly 15.3% of the genome that prior standards omitted. It matters because clinical sequencing is only as trustworthy as the yardstick it is measured against, and until now that yardstick had holes in exactly the repetitive, hard-to-read regions where real disease variants hide. Led by teams at the NHGRI and NIST, it is quiet, foundational infrastructure, the sort of work that never trends and quietly enables everything downstream.

Tied at 67 are two neuroscience results, and the first is a study of a cortical output channel for perceptual categorization, scoring evidence 91, impact 52, and novelty 64. Researchers trained mice to sort sounds into categories by how fast the sound fluttered, then watched three different populations of cortical neurons as the animals learned. Only one type, the deep-layer output neurons that carry signals out of the cortex to the rest of the brain, developed clean categorical responses; the other populations did not reorganize. The finding people will remember is that this learned code was context-gated: it appeared while the mouse was doing the task and vanished when the same neurons heard the same sounds passively, on the same day. In plain terms, the study locates where a small perceptual decision is written in the brain, and shows it is routed through a specific export channel rather than smeared across the whole circuit. The modest impact score reflects that this is basic-science depth rather than an immediate application, which is exactly what an honest scale should record.

The second entry at 67, and the higher-novelty of the pair, is the first brain-and-cord connectome of an adult fruit fly, with pillars of evidence 97, impact 63, and novelty 71. A connectome is a full wiring diagram, every neuron and every connection, and until now the only complete ones came from tiny animals with a few thousand synapses. This map joins a fly's brain and its ventral nerve cord, the insect equivalent of a spinal cord, into a single reconstruction of roughly 160,000 neurons and on the order of 100 million synapses. Reading the wiring, the team found that much of motor control is local: neurons that drive the body are mostly steered by sensory input from the same body part in tight feedback loops, with the brain acting more as a modulator than a central puppeteer. It matters because a synapse-resolution map of a complete nervous system, brain and body together, is the substrate on which theories of how brains produce behavior can finally be checked against the actual circuit. This is a landmark for the whole field of neuroscience.

Also at 67, and carrying the highest novelty in the top ten at 75, is a therapy: a single base-editing infusion that cuts LDL cholesterol up to 62%. Its pillars are evidence 97, impact 60, and novelty 75. The medicine, VERVE-102, is given as one intravenous infusion that carries a base editor to the liver in a targeted lipid nanoparticle and rewrites a single DNA letter to switch off the PCSK9 gene, the body's brake-release on cholesterol clearance. In an early trial of 35 participants across several ascending dose cohorts, PCSK9 protein fell by up to 88% and LDL cholesterol dropped by as much as 62%, an absolute cut of 78 mg per deciliter at the top dose, with reductions durable up to 18 months and no dose-limiting toxicity. The significance is a change in kind: high cholesterol is normally managed with pills taken every day for life, and this is a one-time edit aimed at making the change permanent. The high novelty score reflects how new in-body gene editing at this scale is, and the honest limit is that this is a small, early-phase study; the score's restraint says durability and safety across large populations are still to be proven. This is one of the most consequential results in biomedicine this year.

The sixth entry, at 65, is the most instructive score in the whole section, because its number is built almost entirely from evidence and novelty with barely any impact. The study, electronic and solvent reorganization in proton-coupled electron transfer captured by ultrafast X-rays, scores evidence 94, novelty 68, and an unusually low impact of 34. In plain language, the team used ultrafast X-ray pulses to film, step by step, one of chemistry's most important elementary moves: a proton-coupled electron transfer, where an electron and a proton shift together. Working with a ruthenium model molecule in water, they watched the electron redistribute after a flash of light, then a proton attach to the molecule at around 460 picoseconds, while the surrounding shell of water molecules rearranged in concert. It matters because this exact coupled motion sits at the heart of photosynthesis, fuel cells, and catalysis, and seeing it resolved in real time is a genuine first. The low impact score is not a criticism; it is the method being honest that this is deep, foundational chemistry whose payoff is years out, not a result that moves an industry today.

Those six entries are what Chart F breaks apart. Placing their pillars side by side makes the section's core lesson visible at a glance: the evidence bars are tall and nearly level across all six, so they are not what separates a 70 from a 65. The daylight between these leaders opens up in the impact and novelty bars instead, which is why the AI co-scientist, with the strongest impact of the group, tops the list, and why the ultrafast-X-ray study, brilliant but early, sits at the bottom of this cluster despite evidence as solid as any.

Read the chart and the ordering of the whole index starts to make intuitive sense. A result cannot buy its way to the top on solid evidence alone, because almost everything admitted has that. It has to also be new, or matter, or ideally both, and the rare entries that manage all three at once are the ones that break away. The four remaining profiles, clustered between 62 and 60, follow the same logic, each strong on two pillars and honest about the third.

At 62 comes Neuropixels Opto, a new instrument rather than a discovery, scoring evidence 94, impact 51, and novelty 67. It is a hair-thin silicon probe, seventy micrometers wide and a centimeter long, that packs 960 recording sites together with two sets of fourteen light emitters delivering blue and red light along the shaft. The combination lets a single device both listen to neurons firing and switch specific neurons on or off with light, at precise depths, deep inside the brain where earlier tools could not reach. Built by the Allen Institute with University College London, the University of Washington, and IMEC, its value is as enabling technology: the Allen team reports it has already begun overturning long-held assumptions about how the cortex is wired. The moderate impact score is the method noting that a tool's influence is proven by the discoveries it produces over time, which are only just starting to arrive.

Next, at 61, is large-scale discovery of protein energy landscapes, with pillars of evidence 96, impact 54, and novelty 66. Most of the recent revolution in biology has been about predicting a protein's fixed shape. This work goes after what that picture leaves out: proteins breathe, flickering into rare higher-energy shapes that matter for how they function, and a new mass-spectrometry method measured the energy of those fleeting motions across 5,778 protein domains in parallel. The striking finding is that this hidden flexibility varies even between proteins that share the same fold and the same overall stability, a difference that static structure prediction cannot see. Coming from the Rocklin lab at Northwestern, it matters because a protein's motion, not just its shape, governs how it binds, catalyzes, and can be designed, and measuring that at scale for the first time gives protein engineers a dimension they were previously blind to.

At 60 is Helios, a 98-qubit trapped-ion machine, the top-scored result in quantum computing this year, scoring evidence 89, impact 53, and novelty 61. The persistent problem in the field is that quantum machines usually get less accurate as they get bigger, and errors pile up faster than qubits are added. Quantinuum's Helios uses 98 trapped ions with all-to-all connectivity, meaning any qubit can interact directly with any other, and reports two-qubit gate fidelity of 99.921%, accurate enough to hold performance as the system scales rather than degrade. From those 98 physical qubits it assembles 48 error-corrected logical qubits, and it ran random-circuit tasks placed beyond the practical reach of classical simulation. It matters because staying above the fidelity threshold that error correction requires is the gate every path to a useful quantum computer has to pass through. The score's ceiling at 60 is appropriate: this is real progress on the hardest bottleneck, not yet a machine doing useful work no other machine can.

The last of the leaders, also at 60, returns to AI: Evo 2, a genome foundation model across all of life, with the highest impact of this final group at evidence 97, impact 84, and novelty 59. It is a large model trained not on text but on DNA, over 9.3 trillion nucleotides drawn from across the tree of life, able to read a million bases of context at once and resolve the sequence one letter at a time. Without any task-specific training it can predict whether a genetic variant is harmful, reaching over 90% accuracy on clinically important BRCA1 mutations, and it can generate coherent genome-scale sequences from scratch. The comparatively lower novelty of 59 is the method being precise: this is a powerful second-generation model that scales up an established idea, which is different from inventing one, even as its impact score of 84 recognizes how broadly useful a general-purpose genomics model is. Built at the Arc Institute, it is the clearest sign that the foundation-model approach now reaches into the language of biology itself.

Step back from the ten and a through-line connects the bookends. The year's highest-scored result and one of its lower-ranked leaders are both AI systems pointed at discovery, one proposing hypotheses, the other reading genomes, and both illustrate the shift from software that summarizes what is known to software that proposes what to try next. That shift is easiest to see in the productized form of the co-scientist, now offered to researchers as a hypothesis-generation tool, walked through here on a real immunology problem.

None of these ten clears the scale's upper reaches, and that restraint is deliberate. Each is strong on the pillars that make it real and new, and honest about the pillar it has not yet earned, whether that is broad impact, long-term durability, or proven use. The ten anchor the index, but they are only the peak of a much larger and more uneven distribution. What that distribution looks like once you widen the lens to entire fields, where some sciences run hot and others quietly lag, is where the numbers get more interesting still.

4. Field by field: where each science stands

The eight fields in the index do not score evenly, and the gap between the top and the bottom is wider than most people expect. Biomedicine averages 50.7 across its 27 breakthroughs, while Space and Astronomy averages 37.2 across 12. That is a spread of almost fourteen points, and the reflex is to read it as a verdict on which sciences are healthier or working harder this year. It is not. Every field in the index clears the same high evidence bar, because Frontier only admits work that is already corroborated and traceable to a primary record. Averaged across all 101 entries, evidence sits at 90.7 out of 100 regardless of field. So evidence cannot be what separates biomedicine from space. Something else is doing the sorting.

What sorts the fields is the other two pillars, impact and novelty, and neither is a measure of effort or care. They are structural. Impact asks how directly a result reaches the world, through a clinic, a factory, or another field that builds on it. Novelty asks how far the result sits from everything that came before it. Read that way, the field ranking below is not a scoreboard of quality. It is a map of where genuinely new and genuinely consequential work happened to concentrate in 2026, and the map has a clear shape once you know to look for those two pillars rather than the evidence they all share. You can see how each pillar is defined on the public methodology page.

The bars fall into three loose bands: a leading pair (biomedicine and neuroscience, both near 50), a broad middle (AI, materials, physics, quantum, and climate, clustered between 42 and 46), and space alone at the bottom. Those bands are not accidents of a single year. Each reflects something durable about how the field turns careful work into results that are both new and consequential, and the eight subsections that follow explain, field by field, exactly what that something is.

Biomedicine and Health

Biomedicine leads the index, and it leads for a reason that is easy to state and hard to fake: its results can be measured in a human body. The field carries 27 entries, the most of any, at an average of 50.7, and its ceiling is the complete diploid human genome benchmark at a score of 68, a piece of foundational infrastructure that maps the 15.3% of the genome earlier benchmarks left out. But the field's strength is its depth, not just its peak. A single base-editing infusion cut LDL cholesterol by up to 62%, an absolute drop of 78 mg/dL, in an early trial. A one-letter edit switched fetal hemoglobin back on to treat sickle cell. An individualized mRNA vaccine left 11 of 14 triple-negative breast cancer patients relapse-free for up to six years.

Those outcomes score high on both the pillars that separate fields. Impact is high because a cholesterol number or a relapse count is real-world reach you can put a figure on. Novelty is high because base editing, in vivo delivery, and personalized neoantigen vaccines are genuinely new combinations of ideas, not refinements of old ones. The field also spans machine-learning biology, from the Evo 2 genome model to protein energy landscapes measured at scale, which broadens the base of strong scores rather than resting everything on one result. This is why biomedicine sits at the top: it is the field where solid, new, and consequential most often coincide.

Neuroscience

Neuroscience finishes a close second, averaging 49.6 across 14 entries, and its position rests on a different pillar than biomedicine's. Where biomedicine wins on measurable impact, neuroscience wins on novelty: its best 2026 work consists of first-of-their-kind maps and instruments that did not exist before. The field's top entry, a cortical output channel for perceptual categorization, scores 67 and shows that only one deep-layer cortical cell type develops categorical responses as an animal learns, a code that is present during a task and absent during passive listening in the same neurons on the same day. Alongside it sits the first brain-and-cord connectome of an adult fly, roughly 160,000 neurons wired at synapse resolution, and the Neuropixels Opto probe, a hair-thin shank carrying 960 recording sites and on-shank light emitters.

These results earn strong evidence and novelty scores, yet the field averages a step below biomedicine, and the reason is honest rather than critical. A new map of a circuit, or a new tool to record from one, moves understanding forward before it moves the world. The impact pillar registers that gap: the cortical channel scores 52 on impact against 64 on novelty, because its real-world reach is still one or more discoveries away. That is not a weakness of the science. It is a feature of a field whose payoff is measured first in knowledge and only later in therapy, which places neuroscience high on the board but just shy of the field whose results already reach patients.

AI and Computing

AI and Computing holds the single highest score in the entire index and still ranks only third by field average, and that apparent contradiction is the most instructive fact in this section. The field's top entry, an AI co-scientist that proposes and tests its own hypotheses, reaches 70, the ceiling of the whole 101-entry index, on the strength of a rare 91 / 85 / 74 across evidence, impact, and novelty. It is one of the few results that scores high on all three at once, because a working AI-for-science system is both novel in itself and impactful through every other field it accelerates. Yet AI as a whole averages 45.9 across 13 entries, well below that peak.

The explanation is that AI is the most bimodal field in the index. A handful of entries, the co-scientist and the Kolmogorov-Arnold framework that recovers symbolic physical laws (impact 93), sit near the top, while a longer tail of incremental method work pulls the average down. The honest reading of the field's high ceiling is also a caution: the co-scientist is best understood as a faster discovery loop, generating and ranking hypotheses in narrow, data-rich domains, with human researchers and instruments still verifying every result. That framing, high peak and grounded average, is exactly what an evidence-based score should produce for a field moving this fast, and it is why AI leads the middle band without matching the two fields above it.

Materials and Energy

Materials and Energy averages 45.6 across 10 entries, placing it near the top of the middle band, and its top result is the cleanest illustration in the whole index of how impact and novelty pull in opposite directions. The field's leader, a study that used ultrafast X-rays to capture a proton-coupled electron transfer, scores 65 on the back of a novelty of 68, because resolving where an electron redistributes and exactly when a proton attaches (at roughly 460 picoseconds) with atomic-site specificity is a genuinely new kind of measurement. But its impact score is just 34, the lowest of any top-of-field entry, because the result is foundational rather than applied. It changes what physical chemists can see, not yet what anyone can build.

That split is characteristic of the field, not a quirk of one paper. Materials science advances by first making the invisible visible, and those measurement breakthroughs carry high novelty and modest immediate reach. The field's more applied entries, such as a 4D-printed metamaterial that reshapes on heating to retune its microwave absorption (score 55), sit closer to the world and score more evenly across the pillars. Taken together, materials scores respectably because its best work is inventive, while its average stays out of the leading pair because so much of that invention is upstream of application, waiting for the impact to arrive.

Physics

Physics is the smallest field in the index, just 6 entries, and its scores are the most compressed, running from a maximum of 51 down to 27 at an average of 42.8. Its top entry, a bulk-resolved measurement of the g-wave order parameter in the altermagnet CrSb, pins down the internal structure of altermagnetism, a recently recognized third class of magnetic order that sits between the familiar ferromagnets and antiferromagnets. It is careful, well-evidenced work on a real frontier of condensed-matter physics, and the index scores it as important within its field rather than as a result that reshapes the wider world.

That gap between depth and score is structural, and it is the fairest way to read physics in this index. Fundamental physics clears the evidence bar easily, since its results are among the most rigorously checked in all of science. But the two differentiating pillars work against a high number here. The impact pillar, tuned to real-world and cross-field reach, registers a specialized order-parameter measurement modestly, and the novelty pillar, which rewards atypical combinations of ideas, reads incremental progress on a known question as exactly that. A small field of deep, narrow results will cluster in the low forties almost by construction. The compression of physics scores is not the index missing the physics; it is the index declining to inflate foundational work into something it does not yet claim to be.

Quantum

Quantum averages 42.1 across 9 entries, and it is the field where the distance between engineering achievement and Frontier Score is most worth explaining. The field's leader, Helios, a 98-qubit trapped-ion machine, is a serious milestone by any measure: 48 error-corrected logical qubits, all-to-all connectivity, and 99.921% two-qubit fidelity, accurate enough that error rates begin to fall rather than rise as the system grows. It scores 60, with a pillar profile of 89 / 53 / 61. The evidence and novelty are strong. The impact score of 53 is what holds the total down, and deliberately so.

The reason is that quantum computing's value is still largely promissory. Below-threshold operation, where a logical qubit's error rate improves as you add hardware, is the load-bearing result the field has been chasing, and Helios reaches it. But broadly useful fault tolerance, the point at which these machines solve problems classical computers cannot, remains a 2029 to 2035 horizon on the field's own roadmaps. An evidence-based score cannot credit impact that has not yet arrived, so quantum sits in the middle band despite real progress. This is the index working as intended: it separates a genuine engineering breakthrough from the world-changing application that breakthrough is meant to enable, and scores each for what it has actually demonstrated.

Before the last two fields, it is worth pausing on a dimension the average scores hide. A field's rank tells you how new and consequential its work is, but not how settled the underlying evidence is, and those are different questions. Half of the index is anchored on science that is already firm, while the rest is corroborated for what it currently claims yet still establishing its long-run standing. A result can score high and still be an early signal: the diploid genome benchmark scores 68 and is newly published, so its confidence is marked accordingly. The chart below shows how the whole index divides on that axis.

Read alongside the field averages, the confidence split is a reminder to hold two thoughts at once. A high score means a result is well-evidenced, genuinely new, and consequential relative to everything else this year. It does not promise that the result will still look the same in five years. The 13 early-signal entries are not weaker science; they are more recent or more contested claims whose corroboration is still accumulating, and several sit in the fields that follow.

Climate and Environment

Climate and Environment averages 42.1 across 10 entries, tying quantum in the middle band, and it gets there by a route that is almost the mirror image of quantum's. Its top result, a study finding pesticide residues altering the biodiversity of European soils, scores 52 on a profile that runs strong on impact (69) but softer on novelty (55). Residues turned up at 70% of the 373 sites surveyed across 26 countries and ranked as the second strongest driver of soil biodiversity after soil properties itself. That matters, and the impact score says so.

The field's ceiling stays at 52, though, and no climate entry breaks out the way the leaders in biomedicine or AI do. The structural reason is that much of the best climate science documents and quantifies rather than introduces a wholly new mechanism or tool, and the novelty pillar, by design, reads careful measurement of a worsening trend as incremental. A finding that compound drought-and-heat events have risen nonlinearly since the early 2000s is vital for policy and unmistakably solid, yet it extends a known trajectory rather than opening an unexpected one. That is the honest shape of climate in the index: high evidence, real and rising impact, and a novelty score that reflects how much of the field's value lies in measuring the world precisely rather than in surprising it.

Space and Astronomy

Space and Astronomy sits at the bottom of the field ranking, averaging 37.2 across 12 entries, and its position is the most structurally determined of any field in the index. The reason is not a shortage of remarkable observations. It is that space science is observation-rich and experiment-poor: you cannot run a controlled trial on the early universe, only interpret the light it sends. The field's top entry reads JWST's "little red dots" as young black holes in glowing cocoons and scores just 48, precisely because that reading is one of several. A competing 2026 hypothesis proposes the same objects are pulsating supermassive stars, with masses hedged at roughly 100,000 to a million times the Sun. When the leading result in a field is an unsettled interpretation, both its impact and its evidence-confidence are scored with appropriate caution.

That caution is why the field trails, and it is the right call rather than a slight. Space produces genuinely beautiful results, and the year offers a clear example: JWST detecting iron dust, silicon carbide, and organic PAHs in the dwarf galaxy Sextans A, at just 3 to 7% of solar metallicity, making it the lowest-metallicity galaxy known to contain PAHs. It is a foundational finding about how the early universe made its dust, and it is exactly the kind of work that scores solidly on evidence while the field's headline debates keep its averages low.

Put the eight fields back together and the ranking stops looking like a league table and starts looking like a map of pillars. Biomedicine leads because its results reach the body and score on impact and novelty at once. Space trails because its results, however striking, are often interpretations awaiting confirmation. Everything between is a specific balance of how new a field's work is and how far it reaches. The full per-field breakdown, every entry and every pillar, is downloadable in the open dataset, so the shape described here can be checked against the numbers rather than taken on trust.

5. What the scores actually reveal

An opinion list tells you what a handful of editors admired. An evidence-based score tells you something an opinion cannot: where the difficulty actually sits. Once every entry is broken into three independent pillars, Evidence, Impact, and Novelty, the index stops being a leaderboard and becomes a diagnosis of the frontier. The single most important pattern in the 2026 data is not which result won the year. It is that the three pillars behave completely differently from one another. Evidence sits high and clusters tightly. Impact swings across almost the entire scale. Novelty sits stubbornly low. Read together, those three facts say more about the state of science than any ranking of individual papers ever could, because they separate the reasons a result earns its place from the bare fact that it did.

Across all 101 breakthroughs, the average Evidence score is 90.7, while average Impact is 63.6 and average Novelty is just 52.0. The plain reading is that the science in the index is almost always solid. That is partly by construction, and it is important to say so out loud: Frontier only admits work that is already well-sourced and corroborated, so a high evidence floor is the admission bar, not a discovery about the world. What the numbers then reveal is what happens above that floor. If nearly everything admitted is solid, then solidity cannot be what makes something rank. The differentiator has to live elsewhere, and the data points to exactly where it lives.

The gap between the bars is the finding. Evidence at 90.7 is nearly a constant across the index: in the top entries it moves only from the mid-80s to the high-90s, a band of roughly a dozen points. Impact and Novelty are where entries pull apart. Impact in particular is the widest-ranging pillar of the three, running from near-zero for a narrow mechanistic result into the high-90s for a tool that reshapes a whole field. So the honest answer to what separates a 70 from a 30 is almost never the quality of the evidence. It is whether the work is also genuinely consequential and genuinely new. The frontier is not short on solid results. It is short on solid results that are new and consequential at once.

Concrete cases make that spread vivid. The axon-intrinsic regeneration brake, a precise account of how injured nerves hold themselves back, carries near-perfect evidence yet one of the lowest impact scores in the whole index, because its reach today is narrow and mechanistic. The ultrafast X-ray view of a proton-coupled electron transfer scores 65 overall on exceptional evidence and high novelty, but only low impact, since it resolves a reaction step rather than changing a technology. Against those sits an AI co-scientist with an impact score in the mid-80s. Same tight evidence band, entirely different consequence, and the score is built to see the difference.

This is also why no entry reaches 100, and why that ceiling is deliberate rather than an accident of a young dataset. A perfect score would demand a result that is simultaneously ironclad, field-changing, and wholly unprecedented, and that survives a reproducible scoring of all three at once. The year's highest-scored entry, that same co-scientist proposing and testing its own hypotheses, reaches 70, with strong evidence, high impact, and high novelty, and yet it maxes none of the three. Its evidence base is still firming up, not settled. That a system this striking stops well short of the ceiling is not the scale being harsh. It is the scale being honest about how rare a total breakthrough really is.

The confidence of the evidence base tells a parallel story about hype versus settled science. Of the 101 entries, 50 are Settled, 38 are Firming up, and only 13 are Early signal. Half the index is corroborated, mature science, and barely one entry in eight rests on a single early result. For a domain with a reputation for overpromising, that is a reassuring distribution: the index is dominated by work that has already held up, not by press releases. But there is a tension worth naming plainly. Several of the highest-scored entries are among the least settled, because the newest and most consequential work has simply had less time to be replicated. Freshness and certainty pull in opposite directions, and the score refuses to hide that.

The data also shows where the frontier is dense and where it is thin, and the two axes do not always agree. Biomedicine is both, with 27 entries and the highest field average at 50.7: it is where the most work is happening and where that work scores best. Neuroscience sits close behind at 49.6 across 14 entries. Space and astronomy is well-populated with 12 entries but the lowest average at 37.2, because much of the year's most exciting sky is still interpretive and unsettled, competing hypotheses rather than closed cases. Physics is the thinnest field in the index at just six entries. Density of activity and depth of evidence are separate things, and keeping them separate is precisely what a scored view buys you.

An honest index has to be pressure-tested against itself, so here are the three places this reading is most open to challenge.

  • Novelty may be undercounted. The novelty pillar rewards atypical combinations of prior ideas, following long-standing citation research. A result that is genuinely new but built from conventional methods can still score modestly, so a low novelty average partly reflects how novelty is measured, not only how much exists.
  • The high evidence floor is a selection effect. Evidence sits near 90 because corroboration is the price of entry. The pillar describes the admitted set, not all of science, so "evidence is rarely the bottleneck" is true inside the index, not a claim about research everywhere.
  • The ceiling is a calibration choice. A top score of 70 reflects a scale deliberately built so that 100 stays aspirational. Read it as a conservative young method, not as a verdict that no 2026 result was field-changing.

None of these caveats overturn the pattern; they sharpen how to read it. Evidence is rarely the differentiator, both because science that survives corroboration tends to be sound and because Frontier admits only such science in the first place. What remains genuinely scarce, on any reasonable measure, is work that is solid, new, and consequential together. That is the real signal buried in the 2026 scores, and it is exactly the signal a ranked list of admired papers cannot give you, because such a list never separates why a result earned its standing from the mere fact that someone decided it had.

6. The outlook: 2026 into 2027

A score is a snapshot, and the honest question a snapshot invites is where each line is heading. The 2026 index points to three arcs worth watching into 2027, and the discipline of an evidence-scored view is to describe them without the hype that usually attaches to them. In each case the near-term reality is more modest and more useful than the headline: a faster discovery loop rather than an autonomous scientist, gene editing moving from first-in-human toward pivotal trials rather than curing everything at once, and quantum machines getting more accurate as they grow rather than suddenly breaking encryption. The value of the index is that it lets each claim be checked against where the evidence stands today, instead of where a roadmap wishes it stood.

None of these arcs are predictions in the speculative sense. Each is anchored to work already in the index or to concrete, dated roadmaps from the groups doing the work. That distinction matters, because the gap between what a technology demonstrates and what it routinely delivers is where most forecasting quietly goes wrong. The three arcs below were chosen precisely because their next steps are already scheduled and measurable: trials with enrollment targets, machines with named successors and dates, deployments with named partners. When the next result lands, you will be able to place it on the same scale used here and see whether the arc actually bent the way its proponents expected it to.

The first arc is AI for science, and here the index is a useful corrective to the loudest version of the story. The year's top entry is an AI system in the AI and computing field that generates, debates, and ranks its own hypotheses, and its 2026 milestone was genuine. But the honest near-term shape is a discovery loop, not a discovery agent: the system proposes and prioritizes in narrow, data-rich domains, while humans and instruments still verify every result. Nature's own forward view for the year names the rise of AI scientists among the defining developments while holding to that same caution about verification.

The clearest 2027 signal for this arc is deployment breadth, not raw capability. Google DeepMind has begun bringing its co-scientist to all seventeen US Department of Energy National Labs under a program it calls Genesis, alongside an in-house claim of compressing "years to days." Treat that slogan as a hypothesis to test, not a finding: the meaningful metric will be replicated discoveries that hold up outside the lab that made them, scored on the same evidence, impact, and novelty pillars as everything else in the index. If the loop is real, its results will enter the ranking and move a field's numbers.

The second arc is base editing reaching the clinic, the most consequential biomedical thread of the year. Two entries anchor it: a single infusion that cut LDL cholesterol up to 62% by editing the PCSK9 gene in the liver, and a base edit that switched on fetal hemoglobin to treat sickle cell disease. Both are early-stage but durable, and their honest arc is first-in-human to pivotal. The cholesterol program, in particular, is advancing toward a Phase 2 trial after dose-escalation, the point at which promise is tested against hard clinical endpoints.

The live frontier here is in vivo editing, changing genes inside the body rather than in cells edited outside it and returned. One group has published a dated roadmap targeting a regulatory submission by mid-2026, first human dosing in the second half of 2026, and a late-stage trial in late 2027 for in vivo blood-stem-cell editing in sickle cell and beta thalassemia. Dates like these are what turn a demonstration into a schedule you can hold to account, and they are the honest antidote to talk of gene editing "curing disease" in the abstract.

The proof that editing can be tailored to one person already exists, which is what makes this arc more than incremental. A bespoke n-of-1 in vivo therapy was designed, manufactured, and administered to treat a single patient's rare genetic disease. The open 2027 question is therefore not whether personalized in vivo editing is possible, but whether it can become repeatable and affordable rather than a heroic one-off that only a handful of patients will ever reach. That is the metric to watch, and it is the kind of claim the index is built to score once the evidence matures.

The third arc is quantum scaling, where the index deliberately rewards accuracy over qubit count. The year's top quantum entry, Helios, a 98-qubit trapped-ion machine, matters less for its size than for staying accurate as it grew, reporting 99.921% two-qubit fidelity and 48 logical qubits distilled from 98 physical ones. The load-bearing milestone for the field is below-threshold operation: logical error rates that fall as the machine scales rather than climb. That property, not the headline qubit number, is what a serious reader should keep their eye on.

The dated roadmaps set realistic expectations for the two years ahead. Named successors, Sol around 2027, targeting near 100 logical qubits, and Apollo in 2029, reaching for the hundreds-of-logical-qubits regime, while a competing effort targets a fault-tolerant machine by 2029: roughly 200 logical qubits built from about 10,000 physical ones, using error-correcting codes that cut overhead sharply. These are ambitious but specific, and specificity is what lets an outside observer judge later whether a milestone was hit or missed rather than merely reframed.

The consistent message across independent 2026 outlooks is sobriety, and it is worth taking seriously. Expect lab-scale logical qubits through 2026 and 2027, with broadly useful, fault-tolerant machines still on a 2029 to 2035 horizon rather than imminent. For the quantum field as a whole, that caution is the point: beyond-classical demonstrations are genuinely real now, and general usefulness is genuinely years out, and an evidence score exists precisely to keep those two facts from blurring into one another in the public conversation.

What unites the three arcs is a shape, not a slogan. In each, a striking demonstration has already landed, and the real work of 2027 is repetition, scale, and durability: more replicated discoveries, more treated patients, more accurate machines running longer. That is the unglamorous middle of every real breakthrough, the stretch an opinion list tends to skip past and an evidence score is built to sit inside. When the next milestone arrives, it will enter the same index, be scored on the same three pillars, and either move its field's numbers or fail to. The outlook, in other words, is checkable, which is the only kind of outlook worth publishing.

7. How to read the index yourself

The point of an evidence-scored index is that you do not have to take its word for anything. Every claim in this piece traces back to a primary record, and the index behind it is open in the fullest sense: you can download it, inspect the method, and check any number yourself. That is the whole difference between a ranking you are asked to trust and a ranking you are able to verify. This closing section is a short, practical guide to doing exactly that, both with Frontier's own pages and with any breakthrough claim you happen to meet in the wild for the rest of the year.

Start with the live index itself, a single ranked view of all 101 breakthroughs, sortable and searchable, with each row carrying its Frontier Score and the three pillars behind it. The full methodology is public at version 1.2.0, so the way Evidence, Impact, and Novelty are computed is not a black box you have to accept on faith. The data sources, up to 13 independent public feeds, are named individually rather than gestured at. And the entire index is an open dataset you can hold in your own hands.

That openness is deliberate and load-bearing. The dataset lives at /data under a CC BY 4.0 license, downloadable as CSV and JSON, free to reuse in your own analysis as long as you attribute the source. If a figure in this article surprised you, that is the page where you can pull the record, trace it to its primary source, and decide for yourself whether the number holds. A score you can re-derive is a fundamentally different object from a score you are simply handed, and this one is built to be re-derived.

It also helps to remember that the index is living, not frozen. It is refreshed weekly, and today's 101 entries cover research published from October 2024 onward, concentrated across 2025 and 2026. Rankings will shift as evidence accumulates, as early-signal results either firm up or fade, and as new work is admitted. Treat any single reading, including this one, as a timestamped view of a moving picture rather than a final verdict, and check the live pages when precision matters.

Those pages tell you how Frontier reads a result. The more portable skill is reading any breakthrough claim, whether from a press release, a headline, or a friend, with the same discipline. Five questions do most of the work.

  • Is it corroborated? Look for at least one independent source confirming the core result. A single unreplicated paper is a signal, not a settled fact.
  • Does magnitude match stage? A 62% cholesterol drop in a small early trial is both promising and provisional. Match the strength of a claim to the stage of its evidence.
  • Is it new, or confirmatory? Both are valuable, but they are different things. Ask whether the work combines ideas in a new way or confirms an existing one more firmly.
  • How far does it reach? A result can be rock-solid yet narrow. Ask who benefits, how soon, and whether the reach is mechanistic or real-world.
  • Can you trace every number? If a striking figure has no primary record behind it, treat it as a headline, not a finding, until one appears.

Used together, those questions are simply the three pillars translated into plain language: is it real, is it new, does it matter, and can you check it. That is the entire philosophy of the index, and it is deliberately unglamorous. It will not tell you a result is mind-blowing; it will tell you a result is corroborated, consequential, and traceable, or that it is not yet, and it will show its working either way. The aim is not to end argument about the year's science but to ground it in something shared.

Frontier is built by Yuma Heymans (@yumahey), whose background is in large-scale search across roughly a billion online profiles and in consulting rigor, a lineage that shows plainly in the index's insistence that every figure trace to a source. The instinct throughout is the same one that guides good evidence work anywhere: prefer the primary record to the confident summary, and build a system where anyone can check the claim rather than a system that asks to be believed.

So read the dataset, challenge the method, and hold the next "breakthrough of the year" you encounter to the same standard used here: not whether it is admired, but whether it is solid, new, consequential, and checkable. That standard is the only real defense against a science conversation driven by who shouted loudest, and it is the reason this index exists at all.

Figures in this piece reflect the Frontier Index as of August 2026 (methodology version 1.2.0); the index refreshes weekly, so current scores and rankings may differ.

The Frontier Brief

The week's frontier, scored.

Every Monday, the highest-scoring new breakthroughs in the index, each with its score and a one-line read on why it matters. What you get

One email a week. The top new breakthroughs, scored. Unsubscribe anytime.

The Latest Scientific Breakthroughs of 2026, Scored and Ranked | Frontier