Loading
Fetching the latest research
Retrospective study
We built two independent cohorts of drug assets whose Phase 3 result is now known, scored each early publication blind to that result at full criterion granularity, and validated the candidates that survive against a held-out split.
How to read this. Retrospective and hypothesis-generating. Two independent candidate signals now clear a pre-registered, held-out validation bar; neither is a confirmed predictor yet.
Drug development runs on a brutal arithmetic. Roughly 90% of programs that enter clinical testing fail, and the best large-sample estimate puts the odds of a drug making it from Phase 1 to approval at about 14%. The failures that hurt most come last: more than half of Phase 3 failures trace back to lack of efficacy, meaning the drug simply didn’t do what the earlier evidence said it would.
A single pivotal trial costs a median of around $19 million and often several times that, and capitalized estimates of a full development program run to $2.6 billion. Every one of those late failures was a bet placed years earlier, on an early human-trial paper someone judged good enough to build on.
We also know that the published record those bets rest on is shaky. Bayer scientists reported being able to reproduce only about a quarter of published findings they tried to build drug programs on. Amgen managed to confirm 6 of 53 landmark cancer papers. The Reproducibility Project: Cancer Biology found replication effect sizes 85% smaller than the originals. One estimate puts US spending on irreproducible preclinical research at $28 billion a year. Ioannidis made the statistical case for why this happens two decades ago, and it has aged well.
The field’s response has been rigor infrastructure. Reporting standards like CONSORT and ARRIVE, the NIH rigor and reproducibility framework that grew out of Landis et al.’s call for transparent reporting, risk-of-bias instruments like Cochrane RoB 2. Alongside them, a newer forensic toolkit that checks the paper text itself: GRIM recomputes whether reported means are arithmetically possible, statcheck re-derives reported p-values, and Carlisle’s method has flagged non-random baseline data across thousands of randomized trials.
Here’s the gap. All of that machinery grades papers. A separate literature predicts trial outcomes, but from asset-level features: indication, sponsor, molecule properties, trial design (Wong, Siah and Lo on success rates; PrOCTOR on toxicity-driven failure). Almost nobody has connected the two and asked, blinded and at full criterion granularity: do the specific things a rigor review catches in the earliest human-trial publication carry information about what happens in Phase 3?
That’s the question this study tests. Not a collapsed grade. A single number is convenient, but averaging away which criterion tripped is exactly what would hide a real, narrow signal inside a noisy one.
We sourced drug-asset trajectories from ClinicalTrials.gov, deliberately scanning both the completed and terminated slices of the registry. A cohort built only from completed, published trials is biased toward success, because failed programs are written up less often and less promptly. From several thousand screened trajectories, each asset had to have an early human-trial publication, a later definitive Phase 3 result, at least 0.7-confidence identity agreement between the two, and correct temporal order.
Every analysis runs twice, on two independently curated cohorts from the same eligibility pipeline. The narrow cohort requires the early paper to report a favorable efficacy claim: 99 programs, 40 later failed, 59 later succeeded. The broadened cohort admits any early result: 150 programs, 66 failed, 84 succeeded. Both are reported side by side rather than picking whichever looks better.
Broadened cohort
Any early result — 66 later failed, 84 later succeeded
Narrow cohort
Favorable early result required — 40 later failed, 59 later succeeded
To the outcome
Kaimen Rigor saw only the early paper, never the Phase 3 result
Kaimen Rigor scored each early paper blind, frozen at the decision point, with no access to the Phase 3 result, the registry outcome, or any later corrections or retractions. The engine ran in full audit mode at a single pinned version, and it is deterministic by design: two fully independent scoring passes over a held-out smoke cohort, bypassing every cache, produced a between-run intraclass correlation of 1.00 with zero flipped flags.
Nothing tested here was discovered in this cohort. Every criterion is a fixed, published part of the product before the analysis runs: eight reporting dimensions with their sub-checklists (premise, design, biological variables, ethics, key resources, statistics, data and code availability, transparency), the deterministic forensic checks described above, and a regrouping of every finding into six kinds of error, from numerical inconsistency to overstated conclusion to copyediting. Several hundred individually tested criteria per cohort, each run through the same univariate test. No curated shortlist chosen after looking at outcomes.
Read purely as a grader, the corpus is typical of the clinical literature. Of 144 scoreable papers in the broadened cohort, 46 drew a critical rating flagging them for priority expert review, 86 landed in the weak-to-adequate middle, and 12 reached strong or exemplary. Fewer than half came back with a clean overall verdict.
Seven individual criteria showed nominally higher odds of later Phase 3 failure when present in the early paper, with point estimates from 2.4× up to 10×. These are pre-correction associations, and most do not survive false-discovery correction on their own. Two of the seven also cleared an independent sensitivity check, the strongest form being reappearance in the other, independently built cohort’s own robustness pass. The other five are single-cohort, single-test leads, which is precisely the distinction a sensitivity check exists to draw.
Odds of a later trial failure, top individual candidates associated with HIGHER risk
Each row is one individual criterion whose presence in the early paper raises the odds of later failure, plotted with its 95% interval against the no-effect line at 1 — pooled across both cohort definitions and re-sorted so no row maps back to a specific one. Criteria associated with lower risk are tested identically but omitted here for focus.
Individual criteria are a blunt instrument at this cohort size: wide intervals, many comparisons. So criteria that assemble into a coherent kind of error were also tested together, as L1-regularized logistic models inside 5-repeat stratified nested cross-validation, then checked on a temporal holdout, a 70/30 split where the test slice is strictly later in time than the training slice. The pre-registered bar: held-out AUC above 0.60 with a 95% lower bound above 0.50, plus a 1,000-shuffle permutation test.
One composite criterion family clears that bar in the narrow cohort: out-of-fold AUC in the high 0.5s, permutation p under 0.02, and a held-out AUC of 0.73 whose interval clears 0.5. A different family clears the same bar in the broadened cohort, at a similar held-out AUC with permutation p near 0.01. Multiplicity is handled with Benjamini-Hochberg q-values computed both within families and globally.
Temporal-holdout discrimination of the two validated composite criterion families
Criterion family A — narrow cohort
29 held-out programs (9 failures) · 95% CI 0.54–0.90 · permutation p = 0.014
The real receiver-operating-characteristic curves of the two validated composite families, each over its cohort’s temporal holdout — the ~30% of programs whose outcomes are strictly later in time than everything the model was built on. The dashed diagonal is chance (auROC 0.5). On the precision side of the same predictions, average precision is 0.56 against a 0.31 failure base rate for family A, and 0.55 against 0.40 for family B. The jagged shape is honest for held-out slices this size — which is also why the confidence intervals matter more than any local wiggle, and why these are candidates for prospective validation rather than a deployed predictor. The two panels are different models on different cohorts, deliberately not overlaid.
Interested in which criteria these are?
Therapeutic area doesn’t explain it. Raw failure rates do vary by area against the 44% overall rate, but leave-one-area-out testing shows no consistent generalization from indication mix alone; the largest area performs at almost exactly chance when held out, and the most extreme-looking areas are the smallest samples. Every candidate composite was also re-tested against a matched-pairs subset (nearest by indication, publication year, outcome year); a family that reverses or collapses under matching is treated as nonportable.
One honest wrinkle
Rigor grades the evidence package, not the drug. A clean paper can sit under a molecule that fails for biology, and a flawed paper can sit under one that works; this study is consistent with both. The value lives in specific, individually validated criteria, not in a collapsed score. And a temporal holdout plus a permutation test is a real upgrade over an in-sample read, but it still makes these candidates for a dedicated prospective validation study, not a deployed predictor.
Still, the stakes make even a candidate signal worth chasing. At a decision point where the downstream trial costs eight figures and more than half of efficacy bets go wrong, a triage signal doesn’t need to be an oracle to matter. It needs to survive validation.
The fastest way to judge the work is on an asset of your own. Request a demo and we’ll center it on a real BD or R&D decision you bring.
Criterion family B — broadened cohort
45 held-out programs (18 failures) · 95% CI 0.50–0.84 · permutation p = 0.010