Ventures · Blog
Auditing Claims, Not Counting Papers
On the object of scientific scrutiny, and the business that follows from it
When an organization must act on a scientific claim, it asks how strongly the literature supports it and weighs the studies that agree against those that dissent. That procedure counts publication as verification, and it misplaces the object of scrutiny. The unit of evaluation is the scoped claim, not the paper, and the work worth doing is the audit of the warrants by which evidence bears on it. Organizations already pay for judgments like these, in an unreliable form.
The question that misplaces its object
Organizations routinely act on scientific claims whose validity is uncertain. A biomarker is said to predict response to a therapy. A model is said to generalize beyond its benchmark. A method is said to control for a confound that would otherwise undermine the inference it supports. The customary response is to ask how much the published literature supports the claim: assemble the relevant studies, weigh the ones that confirm it against the ones that dispute it, and form an overall impression of where the evidence lies.
That procedure treats each paper’s conclusion as if publication had already done the verifying. A careful reader will discount for study quality, but the conclusions are still aggregated without being re-audited. Publication does not do the verifying,1 and a procedure that assumes it does is scrutinizing the wrong object.
Publication is evidential, not constitutive
A published paper has cleared a threshold: a few referees’ overall judgment, at one time, under the conventions of one venue. That is real information about the claim. It is not a determination that the claim holds up when its support is examined piece by piece.
We know this because investigators have run fresh studies to test large samples of published findings, and a substantial fraction do not hold. A coordinated effort in psychology replicated roughly a third of the effects it re-tested. An industrial team in preclinical oncology confirmed only six of fifty-three landmark results.2 Peer review also agrees with itself only weakly. When the same submissions are placed before independent committees, the accept-and-reject decisions diverge far more than confidence in the system would predict.3 What publication certifies and what a decision requires are different things, and much costly error accumulates in the distance between them.4
A test score records that a student passed a particular test on a particular day. It does not constitute the ability, and if the test was easy it says little about the ability at all. Publication is a record of the same kind: the claim passed one test, and a weak one. Publication is evidential rather than constitutive. Counting such records asks how probable the claim is, given the tests it has passed. An audit asks something else. A claim earns credit by surviving tests severe enough that, were it false, they would likely have exposed it,5 and the audit’s work is to administer such tests at the level of the warrant and report the failures and survivals it finds. No count of publications performs that work, because publication records no such test.
The unit is the claim; the object of audit is the warrant
If publication does not verify, then the literature is not a tally to be summed, and something else has to be the unit of evaluation. The common instinct is to fix the evaluation of science by building a better home for papers: a better venue, a better review process. That mistakes the unit. The right unit is the claim, stated precisely enough that one can say what would count for it, what would count against it, and what would count as a limit on its scope. “This biomarker predicts response to this therapy” is not yet auditable, because it does not say for whom, measured how, over what period. “Under this population, this assay for the biomarker, this operationalization of response, and this outcome window, the biomarker predicts response” is auditable. Scope is what gives a claim a determinate answer.
Once the claim is the unit, each paper stops being a vote and becomes an evidence artifact: a piece of material attached to the claim so that its bearing can be tested. The thing to test is the warrant. A warrant is the link that is supposed to explain why a given piece of evidence bears on the claim at all.6 For the biomarker claim, one warrant is that the study measured the biomarker the way the claim construes it. Another is that the study handled the confound rather than assuming it away, and that its adjustments did not create the opposite problem: conditioning on a collider, a common effect of two variables, manufactures an association that was never there. Another is that the study population falls within the claim’s stated scope, rather than merely near it. Another is that the paper supports the target claim itself, and not a weaker proxy that resembles it.
None of these questions is answered by the fact of publication. A paper may support the claim, qualify it, narrow it, or fail to bear on it at all, and which of these obtains is a finding the audit must establish. The paper does not carry it in by having appeared in print. Papers do not vote; their evidential contribution has to be earned through warrants that survive audit.
Decomposition before verdict
The standard way to evaluate a body of evidence is holistic review: read the material, form an overall impression, and record it. An expert’s overall impression is often perceptive, but it is exposed to a failure it is not structured to prevent. An evaluator forms a global impression early and then recruits considerations in its support. And when two evaluators reach different impressions, the disagreement is hard to interpret, because it attaches to no identifiable point.
The remedy is a matter of order. Overall judgment stays; it moves to the end. Assess the components independently first, each under a named criterion: was the confound handled, does the population match the scope, does this warrant hold. The decision literature calls this the use of mediating assessments. Complex judgments tend to become more reliable when the intermediate questions are scored separately first, because a global impression then has no chance to silently contaminate each component.7 Under this discipline, an audit becomes a structured object. Each assessment stands on its own evidence, with its own confidence and its own limitations. Only afterward is holistic judgment brought to bear, in a synthesis that cites the assessments it integrates and states which parts of the claim hold, which do not, and where the chain of evidence gives way. The verdict remains an act of expert judgment; what has changed is that it comes last, after the components are fixed, so it can no longer shape them. Disagreement also becomes legible. Because the criteria are named and shared, the same question means the same thing across audits, and two evaluators who differ now differ about an identifiable, recurring thing that a decision-maker can act on.
One might object that none of this is new. A careful reviewer already decomposes; she already asks whether the confound was handled and whether the population matches. True, but she does so privately, so no one else can examine the reasoning. She does so tacitly, so it cannot be contested at the point where it went wrong. She does so variably, since the next reviewer decomposes differently or not at all. And she does so disposably, since the reasoning is discarded once the verdict is written. Making the decomposition explicit moves the judgment from a place where it is hard to examine to a place where it can be.
The scrutiny people already pay for
Organizations already spend to evaluate claims like these. A company weighs whether to license an academic finding. A fund weighs whether a startup’s technical result holds. A foundation weighs whether a program’s evidence justifies renewal. A journal or university faces a contested result. In each case, someone pays experts to answer whether a claim survives scrutiny.
The answer comes back as holistic expert judgment. It is fragmented across advisors. It is delivered as a verdict. It rests on reasoning the buyer cannot examine, contest, or reuse. A biotech deciding on a biomarker license pays a specialist to read the literature and return a recommendation, and that recommendation is a global impression with nothing beneath it, the undisciplined form of the judgment decomposition exists to reform. The market for evaluating claims exists and is funded. It is simply served in an unreliable form.
What a commissioned audit delivers
In a commissioned audit, a sponsor brings a claim that bears on a decision. The claim is scoped: stated with the population, measures, and conditions under which it is to be evaluated. The relevant papers, datasets, and models are attached as evidence. The evaluation is then decomposed into warrant-level assessments, each authored, scoped, rated for confidence, and open to challenge. A synthesis integrates them into a stated audit status: which parts of the claim hold, which do not, and where the evidence gives way. What the sponsor receives is a map of the claim’s support, the warrant-level findings and the synthesis together, that it can act on and interrogate.
The economics follow the same shape. Sponsors fund audits. Reviewers are compensated to produce warrant-level assessments within their competence. The platform composes those assessments into audits and, over time, into a standing record. For the biotech, this replaces a go/no-go memo with rated, contestable assessments of each warrant from the biomarker example, integrated into a conclusion the company can interrogate.
Why the audit compounds
A holistic report with nothing beneath it does most of its work the moment it is delivered. Its conclusion rests on judgment that was never structured for reuse, so when the conclusion is doubted there is little to salvage. A decomposed audit does not expire that way. Each warrant assessment is an authored, dated, contestable object. When new evidence arrives or a verdict is revised, the individual assessments largely persist; what changes is mostly their combination. The finding that a study’s population lies outside a claim’s scope stays true, and stays useful, after the headline conclusion reverses.
So the audit accumulates. It deepens as more evaluators test more of its warrants, and the record of why a claim stands or fails is kept. What accrues is a corpus of scrutiny: scoped claims, tested warrants, and the reasoning that connects them. A competitor cannot reproduce that corpus merely by hiring capable people, because the asset is the accumulated work itself. A consultancy resells expert hours. An audit graph is owned, and it compounds.
What the argument assumes
The first assumption is commercial: that decomposed scrutiny can command the budget the holistic form already holds. If it cannot, the market stays where it is.
The second is that warrants decompose cleanly enough to assess in isolation. Some claims separate into independent links. Others are holistic: their warrants interact, so that assessing the components one by one loses what mattered. If too many consequential claims resist decomposition, the method collapses back into the upfront global impression it was meant to discipline.
The third is the supply of evaluators. Scoped warrant assessment asks respected experts to do focused work. The work must be worth their time and legible to their standing, or the audits will be rigorous in design and empty in practice.
The fourth is a matter of stance. The observation that publication certifies less than we ask of it can sound like disregard for the literature, and the evaluators the platform needs are the likeliest to hear it that way. If the distinction between publication and verification cannot be held in the open, the framing loses the participation it requires.
None of these is settled by argument. Each is a question that a live audit, run for a real sponsor on a real claim in computational biology or biomedical machine learning, is positioned to answer, and C-SQD is being built to run it.
Footnotes
-
Nor could anything do it conclusively. Popper’s argument that no finite body of evidence verifies a general empirical claim (The Logic of Scientific Discovery, 1959) is why the standard throughout this essay is survival of scrutiny, not proof. The complaint against publication is not that it fails to achieve certainty but that it is mistaken for the scrutiny it does not perform. ↩
-
The Open Science Collaboration, “Estimating the Reproducibility of Psychological Science,” Science 349 (2015), replicated roughly a third of the effects when they were re-tested with new samples; C. G. Begley and L. M. Ellis, “Raise Standards for Preclinical Cancer Research,” Nature 483 (2012), confirmed six of fifty-three landmark preclinical results in fresh experiments. Both are replication studies — new data testing published findings — in the sense standardized by the National Academies’ Reproducibility and Replicability in Science (2019), notwithstanding the older “reproducibility” in the first title. The fields and methods differ; the lesson — that publication is a weak predictor of a claim’s survival — does not. ↩
-
On the limited inter-reviewer reliability of peer review, see D. V. Cicchetti, “The Reliability of Peer Review for Manuscript and Grant Submissions,” Behavioral and Brain Sciences 14 (1991); and the 2014 NeurIPS consistency experiment, in which two independent program committees reviewed the same submissions and disagreed on a large share of accept-and-reject decisions. ↩
-
J. P. A. Ioannidis, “Why Most Published Research Findings Are False,” PLoS Medicine 2 (2005). The argument concerns base rates and study design rather than misconduct, which is precisely why publication status cannot stand in for the survival of a claim’s warrants. ↩
-
The severity requirement is Deborah Mayo’s: a claim is well tested only to the extent that it has survived a probe that would likely have uncovered its falsity. See Error and the Growth of Experimental Knowledge (University of Chicago Press, 1996) and Statistical Inference as Severe Testing (Cambridge University Press, 2018). The audit’s object is not the probability of a claim given the records in its favor but the discovery of where its warrants fail such probes. ↩
-
The distinction between a claim, the data adduced for it, and the warrant licensing the move from one to the other is Stephen Toulmin’s, The Uses of Argument (Cambridge University Press, 1958). Auditing the warrant, rather than counting the evidence, is the shift at issue here. ↩
-
The use of mediating assessments derives from the decision-quality literature: D. Kahneman, O. Sibony, and C. R. Sunstein, Noise: A Flaw in Human Judgment (2021), and D. Kahneman, D. Lovallo, and O. Sibony, “A Structured Approach to Strategic Decisions,” MIT Sloan Management Review (2019). The core finding is that scoring intermediate assessments independently, before forming an overall judgment, reduces the error introduced when a global impression is allowed to shape each component. ↩