Methodology
How Kaimen Rigor evaluates evidence.
Enough to judge the method, without publishing the prompts or orchestration behind it. Every report is built to be checked, and every finding traces back to its source.
Confidence, materiality, and decision impact are never one number.
Detection confidence
How strong is the internal evidence that this finding is real?
Scientific materiality
Could this change the scientific interpretation of the result?
Decision impact
Could this change a clinical, regulatory, or commercial decision?
Collapsing these into a single score hides exactly the distinction a decision-maker needs.
The method, section by section.
- Sources accepted
- Publications and preprints (as PDF, or resolved by DOI, PMCID, or URL), supplementary appendices, protocols and statistical-analysis plans, trial-registry records, and relevant regulatory documents. Related prior studies are retrieved to contextualize claims.
- Finding taxonomy
- Findings are typed: numerical and statistical consistency, protocol-to-publication concordance, claim-to-evidence support, reporting and ethics completeness, citation integrity, and reproducibility. Each carries a severity.
- Evidence-retrieval process
- Consequential claims are extracted with their exact source text and location, then cross-checked across the assembled documents and against external references (for example, Crossref and OpenAlex) so each claim is evaluated against the evidence that bears on it.
- Calculation & consistency checks
- Reported statistics are recomputed (statcheck-style), and internal quantities like denominators, counts, confidence intervals, and table totals are checked for consistency. Declared data and code links are probed for liveness and basic content consistency where present.
- Materiality framework
- Every finding is graded by whether it could change interpretation or a decision, separating typographical and reporting issues from findings that move conclusions.
- Confidence framework
- Detection confidence reflects the strength of the internal evidence for a finding. It is not a probability of misconduct, and it is deliberately kept separate from materiality.
- Human-adjudication process
- Consequential candidate findings are routed to independent human review. Authors or responsible parties are offered an opportunity to respond before a public record is finalized.
- Version control
- The Kaimen Rigor version and finding taxonomy are stamped on every report. Public reports are versioned and retain the underlying evidence trail, so findings can be corrected, updated, or withdrawn as new evidence emerges.
- Known limitations
- The system is precision-first: it favors defensible candidate findings over exhaustive recall. Some flaw classes, for example image-based integrity issues or problems that require raw participant-level data, are outside its current scope. Candidate findings are not conclusions.
- Model & methodology update policy
- Model-stack and methodology changes that affect output are versioned; material changes are documented so results remain attributable to a specific release.
Measured, and reported honestly.
Two tracks. A prospective accuracy study on curated corpora with human adjudication, where candidate, human-verified, and author- or journal-confirmed findings are reported separately and headline rates are withheld until a corpus is fully adjudicated. And a retrospective outcome study, whose full design and results are written up on the blog.
Track one · Prospective accuracy study
A prospective study of Adcurare across 50 influential recent clinical-trial reports, each read for candidate numerical inconsistencies, reporting omissions, and cross-document discrepancies. The figures below stay withheld until all 50 are adjudicated - the first rule of the design rather than a gap in the page.
- Trial reports evaluated
- Pending
- Candidate findings generated
- Pending
- Independently human-verified
- Pending
- Confirmed by author or journal
- Pending
- Findings that affected interpretation
- Pending
- Positive predictive value after adjudication
- Pending
- Median time per review
- Pending
How it is reported
- No error rate is published until all 50 trials are systematically adjudicated.
- Candidate, human-verified, author-confirmed, and corrected findings are reported separately.
- Trials with unresolved findings are not named.
- Sampling methodology and limitations are disclosed.
On the blog · Retrospective outcome study
Does the rigor of an early trial paper predict its Phase 3 outcome?
We scored the early paper of 166 assets whose Phase 3 outcome is now known (60 later succeeded, 106 later failed), blind to that outcome. The overall rigor score does not, on its own, separate the two, near-chance discrimination. Several individual criteria are associated with later failure, and one clears multiple-comparison correction. The write-up carries the full cohort construction, blinding, statistical model, sensitivity analyses, and limits.
Read the study