Back to insights
Translational research20 August 2026Research note

Drug discovery is an attrition system

A quantitative account of why therapeutic programmes fail and how human-relevant models, disciplined assays, and explicit decision gates can improve attrition without claiming to eliminate it.

Institutional analysis982 wordsBy Ram Labs ResearchEvidence reviewed 20 August 2026
Principal finding

The rational objective is not to maximize candidate throughput. It is to reduce expensive late failure by improving the quality of therapeutic hypotheses, measurement systems, human relevance, and stop decisions at every stage.

~90% clinical attrition

NCATS states that nearly 90% of promising treatment candidates entering clinical trials fail.

Evidence[1]
13.8% Phase I to approval

Estimated overall probability across 15,102 drug-development programs in a peer-reviewed analysis using 2000–2015 data.

Evidence[2]
3.4% oncology probability

Lowest therapeutic-group Phase I-to-approval estimate in the same analysis; estimates vary by indication and method.

Evidence[2]
$1.14bn median capitalized R&D

Base-case estimate including failed trials and opportunity cost for medicines approved from 2009–2018; 95% CI $0.89bn–$1.48bn.

Evidence[3]

Attrition is information, not merely waste

A therapeutic programme tests a chain of propositions: the target contributes causally to disease; a modality can alter it; exposure reaches the relevant tissue; the intervention changes biology at a tolerable dose; and that change improves a meaningful outcome in a defined population. Failure at any link can terminate the programme. The system becomes inefficient when weak propositions survive too long, endpoints cannot distinguish mechanisms, or negative evidence is obscured by selection and sunk cost.

NCATS reports that nearly 90% of promising candidates entering clinical trials fail. A large peer-reviewed analysis of 15,102 programmes beginning Phase I estimated a 13.8% overall probability of approval, with marked variation: its estimate was 3.4% for oncology and 33.4% for infectious-disease vaccines. Methodology and time period matter, so no percentage is a timeless industry constant. The stable conclusion is that uncertainty is structured by stage, modality, indication, and evidence quality.

Evidence[1][2]

Start with a falsifiable therapeutic hypothesis

Target selection should state the causal claim, intervention direction, patient subgroup, expected biomarker response, efficacy endpoint, and foreseeable safety mechanism. Evidence from human genetics can strengthen causal inference by treating naturally occurring variants as perturbations, but it is not sufficient on its own. Variant effect, lifelong exposure, tissue specificity, pleiotropy, and pharmacological modulation can differ. Orthogonal evidence from genetics, pathology, perturbation experiments, and clinical observations should converge before scale is mistaken for confidence.

The programme should record what evidence would refute the hypothesis. A biomarker may demonstrate target engagement without establishing clinical benefit. Rescue experiments can test specificity, but only if reagents and models are characterized. A target that works in one molecular subtype should not be generalized to a heterogeneous diagnosis. Decision memos should expose contradictory evidence and alternative mechanisms, giving governance committees a traceable basis to continue, redesign, or stop.

Evidence[5]

Assay quality sets the ceiling for computation

High-throughput screening can evaluate enormous libraries, but throughput multiplies the properties of the measurement. An unstable assay produces more precisely catalogued noise. Before screening, teams need controls, dynamic range, repeatability, plate-position analysis, reagent provenance, interference testing, and a clear relationship between the measured signal and mechanism. Hits require orthogonal confirmation, concentration-response analysis, counter-screens, and identity and purity checks. The objective is a reproducible effect with a defensible causal interpretation.

Machine learning can prioritize compounds, predict properties, extract structure, and integrate multi-omic evidence. Its valid claim is bounded by the training distribution and experimental feedback. Retrospective enrichment is not a prospective hit-rate improvement; a docking score is not binding; binding is not cellular activity; cellular activity is not exposure or benefit. Prospective, blinded challenge sets and experimentally measured baselines are necessary to determine whether a model adds information beyond established methods.

Evidence[1]

Human relevance should enter before the clinic

Conventional two-dimensional cultures and animal models can be informative while failing to reproduce human tissue architecture, metabolism, immune context, or chronic exposure. NCATS is developing human cell-based tissue chips and organoid approaches to improve prediction, not to declare one universal replacement. Model choice should follow the question: a liver system for metabolism, a barrier model for transport, patient-derived cells for a genotype, or an immune-competent model for inflammatory effects.

Qualification requires reference compounds, inter-laboratory reproducibility, exposure matching, endpoint validity, and a stated domain of applicability. A complex model is not automatically a better model; it can add uncontrolled variance and obscure mechanism. The useful comparison is prospective: does evidence from the new model improve candidate ranking or detect a known liability beyond the current battery? A programme should publish negative and discordant results because those define the boundary where the model should not govern decisions.

Evidence[1]

Cost estimates are scenarios, not natural constants

A JAMA analysis of 63 medicines approved from 2009 through 2018 estimated median capitalized research-and-development investment of about $1.14 billion in its base case after accounting for failed trials and opportunity cost, with a 95% confidence interval of roughly $0.89–$1.48 billion. Estimates changed under alternative assumptions. These values describe a selected set of companies and medicines; they do not establish the cost of any future programme or justify a price.

The older observation labeled Eroom’s Law found that inflation-adjusted approvals per billion dollars of R&D spending had fallen roughly eightyfold from 1950, halving about every nine years. It is a historical productivity diagnosis, not a physical law. Better portfolio accounting separates cash spend, time cost, shared platform investment, failed-program allocation, and post-approval evidence. The operational target is not a single industry cost number but earlier, better-calibrated decisions with preserved evidence about why candidates failed.

Evidence[3][4]

An evidence ledger for programme decisions

Each stage gate should bind a claim to evidence and an action. Discovery records the target rationale and competing hypotheses. Hit-to-lead records assay validity, chemical identity, selectivity, and developability. Candidate nomination records pharmacology, toxicology, exposure margins, manufacturability, and translational biomarkers. Clinical gates record dose rationale, protocol deviations, subgroup evidence, and benefit-risk interpretation. Data, code, models, samples, and decisions need persistent identifiers and versioned provenance.

Portfolio metrics should reward information quality: replication rate, assay failure detected before screening, proportion of candidates with human-relevant evidence, prospective model enrichment, time to decisive experiment, and fraction of stops made before expensive stages. Raw molecule count and model benchmark scores invite local optimization. The serious research-lab posture is to make uncertainty explicit, run the experiment most likely to change the decision, and stop when the therapeutic hypothesis no longer survives. Attrition cannot be abolished; its timing and evidentiary value can be engineered.

Evidence[1][2][5]
Research boundary

Scope and limitations

Success rates depend on dataset construction, programme and indication definitions, censoring, period, and sponsor behavior. The 13.8% estimate uses 2000–2015 data and is conditioned on entering Phase I. Cost estimates are model-dependent and based on disclosed data from a limited sample. Human genetics, organoids, tissue chips, and AI may improve decisions but do not guarantee translation. This article is a research-governance framework, not a prediction of any programme’s success or medical or investment advice.

Evidence base

References

Source review: 20 August 2026. Quantitative values retain their original definitions, periods, and boundaries.

  1. 01
    Advance Development of and Access to More Treatments

    National Center for Advancing Translational Sciences · 2025

    ncats.nih.gov
  2. 02
    Estimation of clinical trial success rates and related parameters

    Biostatistics · 2019

    pmc.ncbi.nlm.nih.gov
  3. 03
    Estimated Research and Development Investment Needed to Bring a New Medicine to Market, 2009–2018

    JAMA · 2020

    jamanetwork.com
  4. 04
    Diagnosing the decline in pharmaceutical R&D efficiency

    Nature Reviews Drug Discovery · 2012

    doi.org
  5. 05
    Validating therapeutic targets through human genetics

    Nature Reviews Drug Discovery · 2013

    www.nature.com
  6. 06
    Our Impact on Drug Discovery and Development

    National Center for Advancing Translational Sciences · 2025

    ncats.nih.gov