Medicine is the field where the gap between a headline and a result does the most damage. A cancer vaccine that "trains the immune system," a gene therapy that "cures" a disease, a weight-loss pill that "rivals injections": each of these phrases was true somewhere in the last eighteen months, and each was also, somewhere else, wildly oversold. The problem is not that the news is fake. The problem is that a press release and a landmark trial read the same in a feed, and the reader is left to guess which is which.
Frontier was built to close that gap. Every entry in its public index carries one number, a Frontier Score from 0 to 100, assembled from three measured pillars: how strong the evidence is, how much impact the result carries, and how novel it actually is. The score is not an opinion. It is computed from up to thirteen independent public data sources and every figure behind it is clickable back to the primary record. This piece is a tour of what that scoring says about the latest medical breakthroughs, drawn entirely from the biomedicine section of the live index, which currently holds 27 published breakthroughs in health and the life sciences.
Two things are worth saying up front, because they frame everything below. First, the highest-scoring medical result in the index reaches 68 out of 100, not 100. That is not a failure of the year. A perfect score would require a discovery that is at once ironclad in its evidence, field-changing in its impact, and genuinely without precedent, and almost nothing clears all three bars at once. Second, the medical breakthroughs cluster at unusually high evidence scores and more modest novelty, which tells you something real about where medicine is right now: this is a moment of powerful methods being applied and validated in humans, more than a moment of conceptual invention. The tools were invented over the last decade. 2025 and 2026 are when they reached patients.
What follows is not a list of everything that happened in medicine, which no article could be, but a scored reading of the results that cleared the index's bar, organized into the handful of themes that actually carry the year: gene editing moving inside the body, artificial intelligence learning to design biology, cancer vaccines proving they can last, obesity medicine going oral and combination, cells and organs becoming therapies, and the human genome finally being read in full. For each, the interesting question is not only what happened but how well it is proven, and that is the question the score is built to answer. Every figure quoted below is drawn from the live index and links back to its primary record, so nothing here has to be taken on trust, including the scores themselves.
How to read a medical breakthrough
Before the results, the method, briefly, because it is the whole point. A medical claim can be exciting for three separate reasons, and a good ranking has to keep them apart.
The Evidence pillar asks how well-established the result is: the size of the trial, whether the finding is corroborated across independent sources, how much citation and replication weight sits behind it, whether it is a peer-reviewed paper or a preprint. The Impact pillar asks how much it matters: the clinical benefit, the size of the population it could help, the momentum it has already gathered in the literature. The Novelty pillar asks how new it is: whether it opens a mechanism nobody had shown before, or applies a known method to a new target. A result can be high-evidence and low-novelty (a rigorous confirmation of an expected effect) or high-novelty and lower-evidence (a startling first-in-human with a handful of patients). The score does not collapse those into a single verdict. It shows you the shape.
That distinction matters most in medicine, where the temptation to overweight novelty is strongest. A "first-ever" is thrilling; a phase 3 trial in 3,127 people is boring; and yet the second is often the one that changes care. Frontier's scoring deliberately rewards evidence, which is why several of the flashiest results below sit in the middle of the ranking rather than the top. The full method is documented, and reproducible, on the methodology page, and every source that feeds a score is listed on the sources page.
A worked example makes the logic concrete. Two of the results below are, in a human sense, equally astonishing: a complete diploid human genome benchmark and a CRISPR therapy built for a single baby in about six months from diagnosis. Told as stories, the baby wins every time. But they score 68 and 42. The gap is not a judgment about which matters more to the people involved. It is a statement about evidence: the genome benchmark is validated to near-perfection across the entire diploid sequence and underpins an entire field, while the one-baby therapy is, by its nature, a single patient with weeks of follow-up measured against his own baseline. Both are real. One is proven at scale and one is a first. The score keeps those apart so that a reader, or an AI answering a question about the year's medicine, does not mistake a moving anecdote for settled science, or dismiss a foundational result because it lacks a human face. That is the entire job.
Here is how the 27 biomedicine breakthroughs actually score, ranked highest to lowest.
The scores run from 68 down to 41, a tight band that is itself informative: these are all serious results, and the differences between them are differences of degree, not of kind. What follows walks through them by theme, because the year's medicine tells a small number of very large stories.
Why medicine is the hardest field to score
It is worth dwelling for a moment on why a health index needs to be more careful than an index of, say, physics or astronomy. In most fields, the distance between a result and its consequences is long and abstract. In medicine it is short and personal. A person reads that a cancer vaccine works and wonders whether it could save someone they love. A person reads that a pill matches the injection and changes what they ask their doctor for. The scoring has to survive that reading, because the cost of overselling is not a corrected citation. It is a false hope, or a delayed treatment, or a patient demanding something that is years from approval.
That is why the evidence pillar carries the most weight in medicine, and why it is built from signals that are hard to fake: the size of the trial, whether it was randomized and controlled, whether independent groups have cited and built on it, whether it cleared peer review or is still a preprint. A single dramatic case, however moving, cannot score as high as a randomized trial in thousands of people, and the ranking below reflects that at every turn. The most futuristic results, the ones that make the best headlines, tend to sit in the middle of the table, held there by the honest thinness of their human evidence. The least glamorous, a large obesity trial or a genome benchmark, tend to rise, because the evidence behind them is deep.
There is a second reason medicine is hard to score: the caveats are load-bearing. In many fields a limitation is a footnote. In medicine a limitation is often the whole clinical story. A therapy that lowers a biomarker but has not yet been shown to prevent the disease the biomarker predicts; a trial whose follow-up is too short to see whether the effect lasts; a comparison defined after the fact rather than by the trial's design: these are not quibbles, they are the difference between a promise and a proof. Frontier records those caveats on every breakthrough page, in the same place as the claims, and the score is calibrated to them. When you see a spectacular result scored in the low 40s below, the caveat is the reason, and it is named.
With that framing, here is the year's medicine, one theme at a time.
Gene editing leaves the laboratory
For a decade, CRISPR and its descendants were a laboratory triumph looking for a clinic. The story of the last eighteen months is that they arrived in the clinic, in vivo, edited inside a living person's body rather than in cells removed, engineered, and returned. Four results in the index mark that transition, and together they are the single most consequential medical theme of the period.
The clearest is a single base-editing infusion that cut LDL cholesterol by up to 62%, which scores 67, the second-highest in all of biomedicine, with an evidence pillar of 97. The therapy, VERVE-102, is a one-time in-vivo base editor that delivers its editing enzyme and guide to the liver inside a GalNAc-targeted lipid nanoparticle and switches off the PCSK9 gene. Across six dose cohorts, 35 participants received it with at least 28 days of follow-up and no dose-limiting toxic effects. PCSK9 protein fell in a dose-dependent way, from 51% at the lowest dose to 88% at the highest, and LDL cholesterol dropped up to 62%, an absolute reduction of 78 mg per deciliter at the top dose. Critically, the effect held out to at least a year in 15 participants. High LDL is currently managed with drugs taken for life; an 88% cut in PCSK9 from a single infusion points toward one-and-done prevention. The score is not higher because, as the trial's own authors note, this is a phase 1 safety study, not a trial powered to show it prevents heart attacks, and a gene edit is permanent, so long-term safety beyond the current follow-up is still unestablished. That honesty is exactly what the evidence pillar is measuring.
A companion result attacks the same problem through a different gene. One CRISPR dose that cut a cardiovascular risk gene by nearly 90% describes CTX310, an in-vivo CRISPR-Cas9 therapy given intravenously to knock out ANGPTL3 in the liver, mimicking the natural loss-of-function variants that lower blood lipids. In 15 participants with uncontrolled lipid disorders on maximal therapy, a single dose reduced ANGPTL3 by a mean of nearly 80% at the higher doses, with triglycerides falling up to 84% and LDL cholesterol up to 87%. It scores lower, at 45, and the reason is instructive: the follow-up was only at least 60 days, response varied widely between individuals, the lowest doses did little, and serious adverse events occurred in two participants, including one death 179 days after a low dose whose relationship to the therapy is unclear. Same mechanism class, same organ, a more cautious score, because the evidence is thinner and the safety signal less clean. The ranking is doing its job.
The third result moves editing from prevention to cure. Base-edited stem cells that switch on fetal hemoglobin to treat sickle cell disease scores 57. Ristoglogene autogetemcel uses an adenine base editor to alter the HBG1 and HBG2 promoters in a patient's own stem cells, switching hemoglobin production from sickle hemoglobin back to protective fetal hemoglobin without making double-strand DNA breaks. Across 31 patients followed for a mean of 6.6 months, fetal hemoglobin rose above 60% of total hemoglobin and sickle hemoglobin fell below 40%, durably resolving the chronic anemia. This is a mechanistic step beyond the nuclease approach used by the first approved therapies such as Casgevy: base editing rewrites single DNA letters without cutting both strands. The novelty pillar (65) reflects that. But the caveats are heavy, and the score reflects those too: this was an unplanned interim analysis with short follow-up, the therapy still requires harsh myeloablative conditioning (87% of patients had significant adverse events and one died of idiopathic pneumonia syndrome), and it demands stem-cell collection and a prolonged hospital stay.
The fourth is the most extraordinary and, tellingly, not the highest scored. A CRISPR therapy built for one baby scores 42, with the index's near-maximal evidence pillar of 97 but a novelty of 46. After a newborn was diagnosed with severe CPS1 deficiency, a disorder carrying an estimated 50% mortality in early infancy, the team built a lipid-nanoparticle base-editing therapy targeting the child's specific mutation and, following regulatory approval, dosed him at roughly seven and eight months of age, going from diagnosis to treatment in about six months. It is the first medicine ever designed, manufactured, and dosed for a single patient's unique mutation, a template for the thousands of individually rare genetic diseases that share no common therapy. Why does it score in the low 40s? Because it is, by construction, a single patient with short follow-up, measured against his own baseline during intercurrent illnesses rather than a controlled comparison. It is a profound proof of feasibility, and Frontier scores feasibility honestly rather than romantically.
The one-baby result is worth pausing on, because it points at a change in what a medicine can be. Almost the entire pharmaceutical system is built around diseases common enough to justify a development program: a drug costs a fortune to design, test, and manufacture, so it has to treat many people to make sense. That logic leaves out the roughly thousands of ultra-rare genetic diseases, many caused by a mutation unique to a single family, for which no commercial drug will ever be built. The CPS1 therapy is the first concrete demonstration that a bespoke in-vivo editor can be assembled, approved, and dosed for one child on a clinical timescale, about six months from diagnosis. If that process can be turned into a repeatable template, with the editing platform held constant and only the guide changed per patient, it would open a category of medicine that the current economics simply cannot reach. The score of 42 is not a comment on that promise. It is a measurement of the evidence for this one child so far: real, early, and singular. The promise and the proof are different things, and the index is built to hold them separately, which is precisely what lets you take the promise seriously without overstating the proof.
Put the four together and the shape of the clinical gene-editing wave is clear. The lipid-related edits achieve the largest, cleanest numbers:
The common thread is delivery. Each of these therapies had to get an editing machine into the right organ inside a living body, and the workhorse is the lipid nanoparticle, the same delivery format the mRNA vaccines made famous, now carrying an editor instead of an antigen. The path looks like this.
It is worth being precise about what changed, because "gene editing" has been in the news for a decade and it is easy to miss the specific shift. The first approved editing therapies, such as Casgevy for sickle cell disease, are ex vivo: a patient's cells are removed, edited in a laboratory, and infused back, a process that requires stem-cell collection, weeks of manufacturing, and harsh chemotherapy to make room for the edited cells. That is why the sickle-cell result above, remarkable as it is, still carries a heavy caveat load and a score of 57 rather than higher: the editing chemistry advanced, but the procedure around it is still brutal. The lipid-nanoparticle results are different in kind. VERVE-102 and CTX310 are in vivo: the editor is infused directly and finds its own way to the liver, with no cells removed and no conditioning. That is the address change, and it is what turns a gene edit from a bone-marrow-transplant-scale ordeal into something closer to an ordinary infusion.
The liver is first for a reason worth understanding, because it explains both the promise and the current limits. Lipid nanoparticles naturally accumulate in the liver, and the liver makes many of the proteins that drive common chronic disease, PCSK9 and ANGPTL3 for cholesterol among them. So the first wave of in-vivo editing targets liver-made proteins behind cardiovascular risk, where the biology is well understood and the readout (a blood lipid level) is immediate. The harder frontier, reaching other organs, the brain, muscle, the immune system, is still ahead, and the nanoparticle route through skull immune cells discussed later is one early attempt at it. What to watch, in scoring terms, is whether these lipid-lowering edits move from changing a biomarker to preventing an event: the trials so far show LDL falling, not heart attacks prevented, and the jump from one to the other is exactly the jump the evidence pillar is waiting on.
That is the story of gene editing in this period: not a new idea, but a new address. The editors work in the organ that matters, from a single infusion, with effects that last. The scores say, correctly, that the evidence is early and the populations are small. They also say this is real.
Artificial intelligence learns to design biology
If gene editing is the year's clinical story, AI-designed biology is its methodological one, and it is the most crowded theme in the index: nine of the 27 biomedicine breakthroughs involve a generative or foundation model designing or predicting biology directly. This is where novelty scores run highest, and where the evidence pillar most often carries a caveat, because a designed protein that works at the bench is not yet a drug.
The anchor is Evo 2, a genome foundation model across all of life, scoring 60. It is a DNA foundation model trained on 9 trillion nucleotides spanning every domain of life, with a 1-million-token context at single-base resolution and 40 billion parameters, making it the largest and most general biological sequence model to date. Without task-specific fine-tuning it predicts the functional impact of genetic variation, from noncoding pathogenic mutations to clinically significant BRCA1 variants, and it generates coherent genome-scale sequences. That BRCA1 detail matters clinically: when a patient's genetic test turns up a variant of uncertain significance, a change whose effect nobody can yet call, the result is often anxiety without action, and a model that can predict whether such a variant is likely benign or damaging speaks directly to one of the most common and frustrating problems in genetic medicine. The team released the weights, the code, and the OpenGenome2 dataset, so the field can build on it, which is part of why its evidence and impact pillars sit as high as they do. The score's honesty is in the "so what": zero-shot variant prediction is promising but not a clinical verdict, and generative genome design raises biosecurity questions the authors had to screen for.
Alongside it, AlphaGenome reads a megabase of DNA at single-base resolution, scoring 55 with an impact pillar of 89, one of the highest in the field. A single model reads 1 megabase of surrounding context and predicts thousands of functional genomic measurements at single-base resolution in one pass, covering gene expression, chromatin accessibility, histone marks, transcription-factor binding, contact maps, and splicing at once. It matched or beat the strongest specialist models on 25 of 26 variant-effect evaluations. Most disease-associated variation sits in the non-coding genome, and interpreting it has meant juggling many narrow tools; consolidating that into one model is a genuinely practical advance. Crucially, scoring all the modalities together let AlphaGenome reconstruct the mechanism of clinically relevant variants near the TAL1 oncogene, explaining how a change acts rather than merely flagging that it might matter, which is the difference between a red mark on a report and a testable hypothesis a clinician can act on. The evaluations, the authors stress, are computational benchmarks, so predicted effects still need experimental confirmation.
Then the protein designers, a University of Washington-led cluster that reads like a catalog of a solved problem being generalized. RFdiffusion2 designs enzymes directly from catalytic geometry (score 53) generates protein scaffolds straight from the geometry of catalytic functional groups, clearing all 41 active sites on a benchmark where prior methods managed 16, and finding active enzymes after testing fewer than 96 sequences each. A companion result, designing zinc enzymes from scratch that rival nature's catalytic speed (score 50), used the same method to build metallohydrolases reaching a catalytic efficiency of 53,000 per molar per second with a turnover of 1.5 per second, with the crystal structure of the best design closely matching its computational model. Enzymes designed from scratch for multistep chemistry (score 42) went further, building serine hydrolases with a full Ser-His-Asp catalytic triad on five folds unlike any natural enzyme, with crystal structures matching the models to within one angstrom.
The design revolution is not limited to enzymes. BindCraft designs working protein binders on the first try (score 49, impact 87) reuses AlphaFold2's weights to reach nanomolar binders with experimental success rates of 10 to 100%, without high-throughput screening, producing working binders against targets as hard as CRISPR-Cas9 and birch allergen. Custom-designed proteins that read specific DNA sequences (score 45) solve a long-open problem, building small proteins programmed to recognize a chosen DNA target and to repress or activate neighboring genes in both bacterial and mammalian cells. And an AI-designed gene editor that works in human cells (score 45) mined 26 terabases of genomes to train protein language models that generated OpenCRISPR-1, a Cas9 sitting roughly 400 mutations from anything in nature that matches or beats standard SpCas9. Finally, a virtual cell model that predicts how cells react to being perturbed (score 41, the lowest in biomedicine, in part because it is a preprint) trained a transformer called State on over 100 million perturbed cells, with a cell-embedding component trained on 167 million cells, and improved discrimination of perturbation effects by more than 30%, even flagging strong effects in cell contexts it never saw in training.
Read together, these are three different attacks on the same prize: being able to specify a biological function and get a molecule or a prediction, without an experimental screen. The gene editor invented from scratch, OpenCRISPR-1, sits about 400 mutations from any natural protein and yet edits human cells as well as the standard tool, which means functional editing enzymes can be generated rather than discovered in microbes, sidestepping the usual problem that natural editors often work poorly when moved into human cells. The designed DNA readers solve a problem, building a protein that recognizes an arbitrary chosen DNA sequence, that has resisted decades of effort, and they work as gene regulators inside living cells. The virtual cell aims at the largest prize of all: predicting how a cell will respond to an intervention before you run it at the bench, which if it holds up would compress the front end of drug discovery. Each is scored honestly for where it is, proof of concept at the bench rather than a validated tool in the clinic, and each is a real step toward biology as something you design rather than something you find.
What is easy to miss, reading these as separate papers, is that they are one program maturing. A cluster of them come from the same lineage of protein-design work (the RFdiffusion and AlphaFold-derived methods), and they are no longer isolated demonstrations that a model can produce a plausible structure. They are hit rates: RFdiffusion2 finding active enzymes after fewer than 96 sequences, BindCraft reaching working binders one design at a time, the zinc enzymes matching their crystal structures to the computational model. A hit rate is what turns a research curiosity into an engineering discipline. When a chemist can specify a reaction mechanism and get a candidate enzyme, or specify a target and get a nanomolar binder, without a high-throughput screen, the economics of building a new biologic change. The bottleneck stops being "can we design it" and becomes "can we validate it," which is a much better problem to have.
The foundation models point the same way from the prediction side. Evo 2 and AlphaGenome are, in effect, general-purpose readers of the genome's functional consequences, and both were built by teams (Arc Institute and Stanford for Evo 2, Google DeepMind for AlphaGenome) with the compute to train at a scale individual labs cannot match. That concentration is itself a caveat the scores implicitly weigh: a single large model that beats 25 of 26 specialist tools is a consolidation of capability, and consolidation puts more of the field's power behind access and compute rather than distributing it. The open release of Evo 2's weights and dataset cuts the other way, toward broad access, which is part of why its evidence and impact pillars are so high. The tension between concentration and openness is going to define this theme, and it is the kind of thing an evidence-based index can track over time rather than assert once.
The honest limit, the one that keeps every result in this theme out of the top of the ranking, is that a design that works in a dish is not a therapy in a person. OpenCRISPR-1 edits human cells in the lab; it is not in a patient. The designed binders reduce allergen binding in patient samples; they are not an approved drug. The virtual cell predicts perturbation responses better than its predecessors; it is a preprint. None of that diminishes the achievement. It locates it. The design revolution is real and it is early, and the scores say both.
The pattern across all nine is the same, and the scores capture it precisely: very high novelty and impact, evidence held back by the gap between a bench demonstration and a validated therapeutic. AI can now design a working enzyme, a working binder, a working editor. Whether any specific design becomes a medicine is a separate, slower question, and the scores refuse to pretend otherwise. You can see the split directly by plotting evidence against impact for the marquee results.
Cancer vaccines get personal, and durable
The idea of a personalized cancer vaccine, one built from the specific mutations in a single patient's tumor, has been circling for years. What changed in this period is duration: the demonstration that the immunity such a vaccine induces can last for years, not months, in some of the hardest cancers to treat.
Individualized mRNA vaccines that evoke durable T cell immunity in triple-negative breast cancer scores 56. Fourteen patients with triple-negative breast cancer, a subtype with few targeted options, received an individualized neoantigen mRNA vaccine after surgery and standard therapy. In nearly all of them, high-magnitude, mostly new T cell responses appeared and stayed functional for several years, splitting into ready-to-act cytotoxic cells and stem-cell-like memory cells. Eleven of the fourteen remained relapse-free for up to six years. Of the three who recurred, one had the weakest vaccine-induced response and then achieved complete remission on subsequent anti-PD-1 therapy. This is durable, vaccine-induced immunity documented over years, which is why the evidence pillar reaches 96.
Its counterpart tackles an even harder target. A personalized mRNA vaccine that builds years-long immunity against pancreatic cancer scores 42, but carries the index's single highest evidence pillar, 99. Autogene cevumeran targets each patient's own tumor mutations in pancreatic ductal adenocarcinoma, one of the deadliest and least-mutated cancers, given after surgery alongside chemotherapy and a PD-L1 antibody. At a median follow-up of 3.2 years, patients who mounted a vaccine-induced T-cell response had recurrence-free survival still not reached at the median, versus 13.4 months in non-responders, and the vaccine-induced CD8+ T-cell clones had an estimated average lifespan of 7.7 years. Why the modest overall score despite that evidence pillar? Because, as the study itself is careful to say, this is a small phase 1 trial and the survival comparison is between responders and non-responders, a split defined after the fact, which cannot prove the vaccine caused the benefit. It is a spectacular signal that is not yet proof, and the score holds both truths at once.
Both vaccines share a design worth understanding, because it explains why they took so long and why the durability result matters so much. A neoantigen vaccine is not an off-the-shelf product. The tumor is sequenced, its unique mutations are identified, the ones most likely to be visible to the immune system are chosen, and a bespoke mRNA vaccine is manufactured for that one patient, all on a clinical timescale. The technology to do that at speed is the same mRNA-and-lipid platform that BioNTech built for COVID vaccines, which is why BioNTech appears on both of these papers. The open question was never whether you could make such a vaccine. It was whether the immune response it induced would last, or fade within months as so many cancer immunotherapies do. The answer, across both the breast-cancer and pancreatic-cancer results, is that in the patients who respond, the T cells persist for years, with the pancreatic study estimating an average clone lifespan of 7.7 years. That is the finding that moves the field, and it is also why the caveat about who responds is so important: in both trials, the patients who mounted a strong response did well, and the ones who did not are precisely where the next generation of vaccine design has to improve. The scores in the mid-40s to mid-50s hold that shape exactly: a real, durable effect in responders, not yet a proven cause of survival across everyone.
The metabolic era arrives in pill and combination form
No area of medicine has moved faster into public consciousness than obesity treatment, and two results in the index show where the field is going after the first generation of GLP-1 injectables: oral delivery and multi-mechanism combinations.
An oral GLP-1 pill that delivers injectable-level weight loss scores 45, with an impact pillar of 86. In a phase 3 trial of 3,127 adults with obesity but not diabetes, once-daily oral orforglipron produced a mean weight change of -11.2% at the top dose over 72 weeks, versus -2.1% for placebo, with 54.6% of patients losing at least 10% of body weight and 18.4% losing at least 20%. The significance is not the magnitude, which is in the range of injectables, but the form: orforglipron is a non-peptide small molecule that can be manufactured and taken as a simple pill, and existing GLP-1 drugs are costly, supply-constrained injectable peptides. A cheaper daily pill could widen access dramatically. The score reflects that this is a single trial in people without diabetes, with gastrointestinal side effects driving more discontinuation than placebo, and with long-term and cardiovascular outcomes still to come.
The combination result scores higher. Bimagrumab plus semaglutide, alone or in combination scores 55, with an evidence pillar of 97. In a double-blind phase 2 trial, 507 adults were randomized across four arms for 48 weeks. The high-dose combination reduced body weight by 17.8 kg, versus 14.2 kg for semaglutide alone, 9.3 kg for bimagrumab alone, and 3.3 kg for placebo. Bimagrumab is an antibody targeting type II activin receptors, designed to cut fat while promoting muscle, a distinct mechanism from GLP-1, and pairing the two adds to what semaglutide achieves on its own. The trials read cleanly side by side.
The two results point at the two frontiers that matter for obesity medicine now that efficacy is largely solved. The oral pill is about access: the first generation of GLP-1 drugs are injectable peptides that are expensive and have been chronically supply-constrained, and a small-molecule that can be made cheaply and swallowed could reach populations that injectables never will. The combination is about body composition: a persistent worry with rapid GLP-1 weight loss is that a meaningful share of the lost weight is muscle rather than fat, and bimagrumab's activin-receptor mechanism is aimed squarely at that, trimming fat while sparing or building muscle. Neither result is conceptually new, which is why the novelty pillars sit around 50, and neither pretends to be. What they are is well-proven answers to the two questions that follow a successful drug class: can more people get it, and can it be made to work better. In a field where the underlying efficacy is already established, those are the questions worth answering, and answering them in randomized trials of hundreds to thousands of people is why these score as high as they do.
The metabolic story is the quiet inverse of the AI-design story. Here the novelty is low (semaglutide is established, activin-receptor blockade is not new), the evidence is high (large, randomized, placebo-controlled), and the impact is enormous simply because of how many people obesity affects. The scores in the mid-40s to mid-50s are exactly right for rigorous, high-population, incremental advances. Not every breakthrough is a first. Some are just very well proven, and worth a great deal.
Cells and organs become medicine
A fourth theme takes the body's own building blocks, its immune cells, its islets, and eventually whole organs from another species, and turns them into therapies. These are among the most futuristic results in the index, and their scores land in the low-to-mid 40s precisely because the futurism outruns the human evidence so far.
CAR-T cells made inside the body scores 43, with an impact pillar of 87. Conventional CAR-T therapy needs weeks of bespoke manufacturing plus lymphodepleting chemotherapy, confining it to a few specialized centers. This team packaged anti-CD19 CAR mRNA into lipid nanoparticles targeted to CD8+ T cells, reprogramming the cells directly inside the body. A single in-vivo dose controlled tumors in humanized mice and drove durable B-cell depletion in nonhuman primates, and notably the B cells that returned afterward were predominantly naive, a signature consistent with an immune-system reset rather than mere depletion. That detail is why the interest runs beyond cancer: a transient, off-the-shelf CAR-T that resets the B-cell compartment is being pursued as a potential treatment for autoimmune diseases driven by misbehaving B cells, a far larger population than any single cancer. The approach has already drawn 258 citations and a field-weighted citation impact of roughly 123, a measure of how fast the field has seized on it. The score stays in the 40s because the evidence so far is in mice and monkeys, not humans, and the depth, durability, and dosing control of an in-vivo response in people are exactly what the first human trials, only now beginning, have to establish.
Manufactured islet cells reverse type 1 diabetes in a first trial scores 42. Zimislecel is an allogeneic islet-cell therapy grown from stem cells and infused once into people with type 1 diabetes and no detectable insulin production. Among the 12 participants in the efficacy analysis, 10 became fully insulin-independent at one year, and every participant was free of severe hypoglycemia, with detectable C-peptide confirming the manufactured islets engrafted and released insulin in response to meals. Restoring insulin independence from a scalable cell source removes the donor-tissue supply ceiling that has always limited islet transplants. The sober counterweight, reflected in the score: recipients still require lifelong immunosuppression, two participants died (from cryptococcal meningitis and from progression of preexisting dementia), and the cohort is small and the analysis interim.
The most striking of the three is an immune map of a pig kidney in a living person, scoring 49. Using transcriptomics, proteomics, metabolomics, and imaging, researchers tracked the immune response of a living patient carrying a CRISPR-engineered pig kidney with up to 69 gene edits, part of a program that reached a record 271-day human xenograft survival. Early T-cell rejection appeared within a week and was reversed with intensified immunosuppression, while rising donor-derived cell-free DNA tracked the rejection and fell with treatment, a promising noninvasive biomarker. The study already carries a field-weighted citation impact near 138, underscoring its landmark status as formal FDA-authorized trials begin. It is a single patient, so the patterns may not generalize, but it is among the first detailed mechanistic roadmaps of a gene-edited animal organ functioning inside a human being.
What unites these three, and separates them from the gene-editing results, is that they are replacing or reprogramming whole populations of cells, not rewriting a single gene. That is a larger intervention with a larger uncertainty, which is why the scores sit in the 40s despite high impact pillars (87 for in-body CAR-T, 77 for the islet therapy, 64 for the pig kidney). The shortage they address is real and enormous: type 1 diabetes destroys the insulin-producing islets and donor tissue is scarce; organ transplant lists run far ahead of donor supply; CAR-T works but reaches only patients who can travel to a specialized center for weeks of manufacturing. Each result attacks the supply ceiling directly, with a manufactured cell source, an animal organ, or an in-body reprogramming that skips the factory. And each carries the same honest limit, named in its score: the human evidence is a first trial, a single patient, or an animal model, and the immunosuppression, durability, and scale questions are open. These are the results most likely to look either prophetic or premature in five years, and the ranking is deliberately cautious about which.
Finishing the human genome, and reading what it says
The last major theme is the least clinical and, by the scores, the most solid: the completion and interpretation of the human genome itself. Two of the four highest scores in all of biomedicine belong here, because this is work of extraordinary evidentiary rigor.
The single highest-scoring medical result in the index, at 68, is a complete diploid human genome benchmark for personalized genomics. Researchers built a telomere-to-telomere benchmark for the diploid HG002 genome that reaches near-perfect accuracy, with no detectable errors, across 99.4% of the complete sequence. It adds 701.4 Mb of autosomal sequence plus both sex chromosomes, covering 15.3% of the genome that prior benchmarks simply omitted, and shows that assembly-based sequencing beats read-mapping by an order of magnitude in the hardest regions. Standard resequencing inherits the biases of the reference it maps to, so duplicated and structurally polymorphic regions go uncalled and cannot even be measured; a complete diploid truth set finally lets methods be scored where they actually fail. The word diploid is doing real work here: earlier complete-genome efforts assembled a single composite sequence, but a person carries two copies of every chromosome, one from each parent, and disease often hides in the difference between them. This benchmark resolves both haplotypes and includes a diploid annotation of 39,144 protein-coding genes across them, with de-novo assembly making just one error per 100 kilobases across 99.9% of benchmark regions. The evidence pillar is 96, and the result is foundational to everything else in clinical genomics: it is the ruler against which the accuracy of every sequencing method is now measured, including in the very regions where clinical tests most often fail silently.
Its companion, nearly complete genomes from 65 people that resolve the genome's hardest regions (score 44), sequenced 65 diverse human genomes into 130 haplotype-resolved assemblies, closing 92% of previous assembly gaps, completely assembling 1,246 centromeres, and resolving medically important loci like the immune MHC and the SMN1/SMN2 region behind spinal muscular atrophy. It lifted short-read genotyping quality high enough to detect 26,115 structural variants per individual, expanding what disease-association studies can examine.
Two more results read the genome for what it reveals about us. Ancient DNA reveals pervasive directional selection across West Eurasia scores 61, and it is a genuinely surprising result. Distinguishing sustained, fitness-driven change in the genome from the noise of migration and population mixing has been the central obstacle to using ancient DNA for evolutionary biology rather than just history. By building a method that tests for consistent allele-frequency trends over time and applying it to 15,836 ancient genomes (10,016 newly sequenced), the team found many hundreds of alleles under strong directional selection across the past ten millennia, overturning the prior picture in which such hard sweeps were rare. They estimated selection coefficients at 9.7 million variants, and documented shifts, on the scale of modern variation, in allele combinations that today predict lower body fat, schizophrenia risk, and cognitive measures. The authors are careful about the last point, and so is the score: the trait predictors were measured in industrialized societies, so how these signals relate to what was actually adaptive in the past remains open.
And mapping the genetic landscape across 14 psychiatric disorders (score 52) drew on 1,056,201 cases to find five underlying genomic factors that account for the majority of each disorder's genetic variance, around 66% on average, and map to 238 pleiotropic loci. The finding with teeth is that disorders long debated as distinct, schizophrenia and bipolar disorder among them, share almost all of their genetic signal, with very few disorder-specific loci, and that the factors map onto distinct cell types: a schizophrenia-and-bipolar factor enriched in excitatory neurons, an internalizing factor spanning depression, PTSD, and anxiety tied to oligodendrocyte biology. That points toward a classification of mental illness grounded in neurobiology rather than symptom checklists, and toward therapies aimed at commonly comorbid presentations rather than single diagnostic labels.
Three more results resist the themes above but score well on evidence, and each is worth a moment. Large-scale measurement of protein energy landscapes (score 61) addresses a blind spot created by the protein-design revolution itself: we can now predict a protein's resting structure with high accuracy, but the rare, higher-energy states that actually govern function, aggregation, and immune response have stayed invisible. Using multiplexed hydrogen-deuterium exchange mass spectrometry, the team measured the fluctuation energies of 5,778 protein domains in parallel, revealing hidden variation even between sequences that share a fold, and producing exactly the kind of large quantitative dataset that machine-learning models of protein stability have lacked. It is, in effect, the training data the next generation of design tools needs.
Nanoparticles that hijack skull immune cells for CNS drug delivery (score 50) attacks the problem that defeats most brain therapies: the blood-brain barrier blocks the majority of systemically delivered drugs. By loading drugs into nanoparticles that commandeer the immune cells resident in the skull and route them through skull-meninges microchannels directly into the brain, the approach reaches lesions without crossing the barrier, improving outcomes in preclinical stroke models with early clinical support. And a single-cell map of the sensory neurons that drive fracture healing (score 49) profiled the nerve cells that innervate bone before and after fracture, showing that removing that innervation cripples repair and naming FGF9 as a major regulator of bone healing, a concrete new target for the common, costly problem of fractures that do not mend.
The reason the genome results score so well deserves a sentence of its own, because it runs against the usual intuition that the exciting news is the clinical news. Completing the genome is infrastructure, and infrastructure is where evidence is strongest and consequences are longest. Every clinical genomics test, every disease-association study, every attempt to read a patient's DNA for risk, inherits the quality of the reference it is measured against. For twenty years that reference was incomplete, missing exactly the duplicated and repetitive regions where much disease-relevant variation hides. The diploid benchmark and the 65-genome pangenome close that gap, and they do it with the kind of exhaustive validation (near-perfect accuracy across 99.4% of the sequence, 1,246 fully assembled centromeres) that pushes an evidence pillar into the high 90s. These are not the results that will make next year's headlines. They are the results that everything in next year's headlines will quietly depend on.
The institutions behind the frontier
One thing an evidence index makes visible, almost as a side effect, is where this work comes from, and the pattern across the 27 biomedicine results is worth naming. Three kinds of institution recur.
The first is the large public research complex, and in this data it is overwhelmingly American. The top-scoring result, the diploid genome benchmark, is a collaboration of the National Institutes of Health, the National Human Genome Research Institute, the National Institute of Standards and Technology, Johns Hopkins, and UC Santa Cruz. The ancient-DNA selection map runs through the Broad Institute and Harvard. The one-baby CRISPR therapy came out of Children's Hospital of Philadelphia and Penn, with the Innovative Genomics Institute and Berkeley. This is the slow, expensive, publicly funded infrastructure that produces the highest-evidence results, and it is not an accident that it clusters at the top of the ranking.
The second is the design and foundation-model lab, where a small number of groups are producing a disproportionate share of the novelty. The protein-design results (RFdiffusion2, the zinc enzymes, the multistep hydrolases, the designed DNA readers) trace to a University of Washington-centered lineage with MIT and the Howard Hughes Medical Institute. The genome foundation models come from Google DeepMind (AlphaGenome) and the Arc Institute with Stanford (Evo 2, and the State virtual cell). BindCraft and the AI-designed gene editor come from the SIB Swiss Institute of Bioinformatics and EPFL. These are the groups turning biology into an engineering discipline, and they concentrate capability in a way the scores flag as both a strength and a risk.
The third is the company partnered with a clinic, which is where the clinical results live. BioNTech is on both cancer-vaccine papers, extending its mRNA platform from infection to oncology. Eli Lilly appears on the metabolic and the LDL-editing work. CRISPR Therapeutics is behind the ANGPTL3 edit. And a distinct cluster of the neuro-delivery and stroke work comes from Chinese institutions, including Tsinghua, Capital Medical University, and Beijing Tian Tan Hospital, a reminder that the geography of the frontier is broader than the top of this particular table suggests. Tracking that distribution over time, which institutions and which countries are producing the highest-evidence medicine, is the kind of question the open dataset exists to let anyone answer for themselves.
What to watch over the next year
Because Frontier records the caveat alongside every claim, it doubles as a watchlist: the limitation on each result is usually a precise statement of the experiment that would raise its score. Read that way, the biomedicine index points at a handful of specific things to watch.
Does a biomarker become an outcome? The in-vivo lipid edits, VERVE-102 and CTX310, have shown they can lower cholesterol dramatically from a single infusion. The result that would move them toward the top of the ranking is a trial large enough and long enough to show they prevent heart attacks and strokes, not just lower a number. That is a multi-year readout, and it is the single most consequential thing to watch in this whole set.
Does an animal model become a human one? In-body CAR-T has produced striking results in mice and monkeys, and first-in-human trials of related platforms are only now beginning. If in-vivo reprogramming matches the potency and persistence of manufactured CAR-T in people, cell therapy stops being a specialized-center procedure and becomes something closer to an injection, with implications well beyond cancer into autoimmune disease.
Does a design become a drug? The protein-design and foundation-model results have shown they can generate working enzymes, binders, and editors at useful hit rates. The step that would lift their evidence pillars is one of these designs completing the long walk through preclinical validation and into a human trial. Watch BindCraft and the AI-designed editor, whose outputs are closest to therapeutic use.
Does a responder effect become a population effect? Both cancer vaccines work impressively in the patients who mount a strong immune response. The next generation has to close the gap for the patients who do not, and the trials to watch are the larger, randomized ones now underway that will test whether the durability seen in triple-negative breast cancer and pancreatic cancer holds across everyone, not only responders.
In every case, the thing to watch is the same in structure: a result with high novelty and impact but an evidence pillar held back by early or narrow data, waiting on the specific trial that would prove it. The live biomedicine page is where those score changes will show up first, as the evidence arrives and the scores move with it.
What the pattern of scores actually says
Step back from the individual results and the ranking tells a coherent story about the state of medicine, one that no single headline could.
The evidence pillars are extraordinarily high, and the novelty pillars are not. Across the featured results, evidence scores cluster in the 90s while novelty sits in the 50s and 60s. That is the signature of a translational moment: the field is not inventing many wholly new concepts, it is taking methods proven over the previous decade (base editing, mRNA vaccines, GLP-1 agonism, protein design, long-read sequencing) and validating them in humans with rigorous trials. That is arguably the more valuable phase. Ideas are cheap; evidence in patients is expensive and rare, and it is what actually changes care.
Nothing reaches 70, and that is honest. The top medical score is 68. A perfect score would demand a result that is simultaneously ironclad, transformative, and unprecedented, and the results that are most transformative (in-body CAR-T, xenotransplant, one-baby CRISPR) are precisely the ones whose human evidence is thinnest, while the results with the strongest evidence (large obesity trials, genome benchmarks) are incremental rather than unprecedented. The very structure of the scores encodes the central tension of medical progress: the boldest ideas are the least proven, and the best-proven advances are rarely the boldest.
The delivery format is quietly the hero. Look across the themes and the same technology keeps reappearing in different costume: the lipid nanoparticle. It carries the base editor to the liver in VERVE-102, the CRISPR machinery in CTX310, the CAR instructions to T cells in the in-body CAR-T, and it is the descendant of the same particle that carried mRNA in the COVID vaccines and now carries the neoantigen vaccines. A large share of the year's clinical progress is not a new biological idea at all. It is a delivery vehicle maturing to the point where older ideas (edit this gene, express this antigen, reprogram this cell) finally reach the right place in the body. That is invisible if you read the results one at a time, and obvious the moment you line them up by evidence, which is exactly what an index is for.
Scores are a starting point, not a verdict. A number between 41 and 68 does not tell you whether a therapy is right for a particular person; no index can, and none should try. What the score does is tell you how much weight the current evidence can bear, and point you at the primary record so you can form your own view. A high score is an invitation to take a result seriously. A modest score on a dramatic result is an invitation to read the caveat before you believe the headline. Used that way, the ranking is not a substitute for judgment. It is scaffolding for it.
The frontier is defined by recency, on purpose. Frontier is a recency-led index: a result's standing reflects the edge of what is currently being established, and as a finding settles into routine practice it is, in a sense, no longer frontier, it is medicine. That is why this reads as a map of 2025 and 2026 rather than an all-time list, and why the timeline below matters. It also means the scores you see are a snapshot of an active argument, not a monument. As the trials that these results are waiting on report out over the next year, some of these numbers will rise and some will fall, and the live index will move with them. A static article can only show you the frontier as it stands today; the point of the underlying index is that it does not stand still.
The timing is unmistakable. Plotting the biomedicine results by quarter shows the wave building through 2025 and cresting in early 2026.
The concentration in early 2026, seven of the twenty-seven in a single quarter, is not a statistical artifact so much as a convergence: several long-running programs (the genome foundation models, the in-vivo lipid edits, the large obesity trials) happened to report their landmark results within months of each other, which is what a field looks like when a decade of method-building matures at once. The recency tilt also reflects how the index works, weighting the current edge, so the picture will keep shifting toward the newest quarter as fresh results are scored.
And every one of these results comes with its own honest limits, because Frontier records the caveats as carefully as the claims. The sickle-cell therapy still needs harsh conditioning. The pancreatic vaccine's survival comparison was defined after the fact. The oral GLP-1 pill has only one trial and known gastrointestinal side effects. The in-body CAR-T is still in mice and monkeys. Reading the scores is not a substitute for reading those caveats; it is a way in to them. That is the difference between a breakthrough index and a hype cycle: the index tells you where to be skeptical, and hands you the primary record so you can check.
What the evidence score is actually made of
It is fair to ask what "high evidence" concretely means, because a score is only as trustworthy as the signals underneath it. In medicine, Frontier's evidence pillar is assembled from measurable, public facts about a result rather than any editorial impression of it. Three kinds of signal do most of the work.
The first is how much the scientific community has taken the result up, measured not by raw citation count (which favors old and large-field papers) but by field-weighted citation impact, a normalized figure where 1.0 is the average for a paper of the same field, type, and age. The signal is visible directly in two of the results above: the pig-kidney immune map carries a field-weighted citation impact near 138, and the in-body CAR-T sits around 123 with 258 citations, numbers that mean the field seized on these papers at more than a hundred times the normal rate. That uptake is part of why both score respectably on evidence and impact despite being early-stage, and it is a signal a human editor would struggle to weigh consistently across 27 papers.
The second is the nature of the study itself: a randomized, controlled, blinded trial in thousands of people carries more evidentiary weight than a single-arm study in a dozen, and a peer-reviewed paper in a top journal more than a preprint. That is why the oral GLP-1 trial in 3,127 randomized adults and the bimagrumab combination in 507 hold strong evidence pillars, and why the State virtual cell, a bioRxiv preprint, scores lowest of the 27: not because the work is weak, but because it has not yet cleared peer review, and the score says so.
The third is clinical translation: whether a result has registered clinical trials, whether it has reached patients, whether public funders and regulators have engaged with it. A laboratory demonstration and an authorized human trial are different evidentiary objects, and the pillar treats them differently. All of this is drawn from public data sources, the same ones listed on the sources page, which is what lets the score be reproduced and checked rather than trusted. The point is not that the number is perfect. The point is that it is made of things you can look up, and every one of them is linked from the breakthrough's own page.
How to use this, and verify every number
Everything above is a snapshot of a living resource, not a static list. The biomedicine field page shows these 27 breakthroughs as they stand today, and it updates weekly as new results are ingested and scored, so it is always current in a way this article is not. Each breakthrough has its own page where every figure is clickable back to the primary source, from the trial paper to the citation databases behind the score, so you can check the claim rather than trust it. The full index across all fields lives on the Frontier homepage, and the way the scores are computed is documented on the methodology page.
If you work with this data, the entire scored index is available as a free, open dataset in CSV and JSON, licensed CC BY 4.0, on the data page. And if you would rather have the frontier come to you, the Frontier Brief is a weekly email of the highest-scoring new breakthroughs across every field, medicine included.
The reason to read medicine this way, rather than through the feed, comes down to what the two formats do to your attention. A feed rewards the most surprising claim, which is systematically the least proven one, because surprise and evidence pull in opposite directions: the more established a result, the less it startles. An index inverts that pressure. It puts the best-proven results where you will see them and marks the dramatic ones with the caveat that keeps them honest, so that over a year of reading you end up calibrated rather than whiplashed. You come away knowing that in-vivo gene editing genuinely works in the liver and is still early, that AI can now design working proteins and has not yet designed an approved drug, that obesity medicine is going oral, and that the human genome is finally complete enough to read where it used to go dark. None of that is a headline. All of it is true, and the scores are how you can tell.
That is the whole ambition of an evidence-scored index, and it matters most exactly here. In medicine the reader is rarely a spectator. They are a patient, or a caregiver, or a clinician, or an investor, or a researcher deciding what to build on, and every one of them is better served by a number they can trace than by an adjective they have to trust.
For the wider picture beyond medicine, the companion report The Latest Scientific Breakthroughs of 2026, Scored and Ranked applies the same evidence scoring across all eight fields of science. Medicine is where the stakes of getting the score right are highest. It is also, on this evidence, where the frontier is moving fastest.