How We Evaluate Evidence

Evidence Should Be Able to Change the Conclusion
A conclusion becomes more trustworthy when it survives serious attempts to challenge it. WIN therefore treats evidence not as material gathered to defend a preferred answer, but as information that can strengthen, weaken, revise, or overturn an explanation.
This requires more than collecting facts. Evidence must be examined for relevance, reliability, independence, context, limitations, alternative interpretations, and the strength of the connection between what is observed and what is being claimed.
The objective is not perfect certainty. In complex human systems, perfect certainty is often impossible. The objective is disciplined judgment: knowing what the evidence supports, what it does not support, how confident the conclusion should be, and what remains unresolved.
Start With the Quality of the Evidence
Not all evidence deserves equal weight.
A claim may rest on direct observation, controlled research, administrative data, historical records, expert analysis, testimony, anecdote, inference, or other forms of information. Each can be useful for particular purposes, but each has different strengths and limitations.
WIN asks whether a source is relevant to the question, whether the information can be independently checked, whether important context is missing, whether the measurement is appropriate, and whether there are plausible reasons the evidence could be distorted or incomplete.
The strength of a conclusion should reflect the strength of the evidence supporting it.
Consider Source Independence
Ten sources repeating the same original claim are not necessarily ten independent pieces of evidence.
WIN examines whether apparently separate sources ultimately depend on the same underlying study, dataset, witness, report, assumption, or information chain. Repetition can increase visibility without increasing evidentiary strength.
Independent evidence is especially valuable when different methods or sources converge on the same conclusion.
Distinguish Correlation From Causation
Two things occurring together does not by itself establish that one caused the other. They may share another cause, influence one another in both directions, occur together by coincidence, or reflect a selection or measurement effect.
Causal claims therefore require additional reasoning. Timing, mechanism, comparison cases, competing explanations, dose or intensity effects, natural experiments, controlled studies where appropriate, and consistency across different forms of evidence can all matter.
WIN matches the strength of causal language to the strength of the available evidence.
Look for Evidence Against the Explanation
A serious test of an idea asks what we would expect to observe if the idea were wrong.
Contradictory cases, failed predictions, unexpected outcomes, and evidence favoring another explanation are not inconveniences to be hidden. They are information.
WIN deliberately examines evidence that could weaken important conclusions. This reduces confirmation bias—the tendency to notice and remember information that supports an existing belief while discounting information that challenges it.
Ask What Would Prove Us Wrong
Falsifiability means that a claim should, where possible, expose itself to evidence that could show it to be mistaken.
If no possible observation could count against a claim, the claim becomes difficult to evaluate objectively.
For WIN, “Prove Us Wrong” is therefore not rhetorical. Important explanations identify assumptions, expected results, boundary conditions, and evidence that would require reconsideration. A framework that cannot be corrected cannot reliably improve.
Replication and Convergence Matter
A single result can be important without being conclusive.
Confidence generally becomes stronger when relevant findings can be reproduced, when independent evidence converges on the same explanation, and when a conclusion survives examination using different methods, populations, conditions, or sources where those comparisons are appropriate.
Failure to reproduce a result does not automatically prove the original finding false, but it can materially reduce confidence and requires investigation.
Population and Context Matter
Evidence from one population, location, institution, historical period, or set of conditions does not automatically apply unchanged to another.
WIN examines whether the people and circumstances represented by the evidence are sufficiently relevant to the claim being made. Differences in age, incentives, institutions, culture, technology, economic conditions, implementation, selection, or other variables may limit generalization.
Conclusions are kept proportionate to the populations and conditions actually supported by the evidence.
Confidence Should Match the Evidence
Evidence rarely divides neatly into “true” and “false.”
Some conclusions are strongly established; others are well supported but incomplete; some remain plausible hypotheses; and some questions remain genuinely unresolved.
WIN keeps those distinctions visible. Confidence rises when independent evidence converges, predictions succeed, competing explanations weaken, and results remain durable. Confidence falls when evidence conflicts, assumptions fail, important variables were omitted, findings do not reproduce, or better explanations emerge.
Uncertainty is not a defect when it accurately reflects the state of knowledge. False certainty is more dangerous because it can prevent correction.
Distinguish Absence of Evidence From Evidence of Absence
Failing to find evidence for a claim does not always establish that the claim is false.
The evidence may be difficult to observe, the available sample may be too small, the measurement may be insensitive, or the relevant information may not yet exist.
At the same time, repeated failure to find an expected effect can become meaningful when the effect should have been detectable under the conditions examined.
WIN therefore asks not merely whether evidence was absent, but whether the evidence should reasonably have been present and detectable if the claim were correct.
Examine Measurement Quality
Evidence can be only as reliable as the methods used to produce it.
WIN considers whether important concepts were defined clearly, whether measurements actually represent what they claim to measure, whether data were collected consistently, whether missing information could alter the result, and whether the measurement process itself could introduce bias or error.
Precise numbers do not compensate for measuring the wrong thing.
Separate Statistical Significance From Practical Importance
When statistical analysis is relevant, a detectable difference is not automatically an important difference.
WIN considers the magnitude, practical consequence, uncertainty, context, and durability of an observed effect rather than treating a statistical threshold by itself as proof that a result is meaningful.
The practical importance of an effect depends on the question being examined.
Consider Conflicts of Interest Without Using Them as Proof
Financial, professional, ideological, institutional, or personal interests can create incentives that deserve consideration when evaluating evidence.
A conflict of interest does not automatically make a finding false, just as the absence of an obvious conflict does not make a finding true.
WIN examines the evidence and methodology themselves while also considering whether relevant incentives, funding relationships, selective reporting, or undisclosed interests could have influenced how evidence was produced or presented.
Competing Explanations Must Be Compared
Finding evidence consistent with one explanation is not enough when the same evidence is also consistent with plausible alternatives.
WIN compares competing explanations and asks which accounts for the evidence with the fewest unsupported assumptions, which better explains contradictory cases, which produces useful expectations, and what additional evidence could distinguish among them.
This is particularly important when examining proposed Real Root Causes. A deeper-sounding explanation is not automatically a better causal explanation.
Match the Claim to the Evidence
Evidence supporting a narrow conclusion does not automatically justify a broad one.
WIN distinguishes statements such as “this occurred in the observed group,” “this association appeared under these conditions,” “this intervention may have contributed to the result,” and “this factor causes the outcome generally.”
Those statements make progressively different claims and may require progressively stronger evidence.
Historical Evidence Requires Context
Historical evidence can reveal recurring patterns, institutional failures, incentives, unintended consequences, and long-term effects that are difficult to observe in short studies.
It also requires caution. Different historical periods may involve different technologies, institutions, populations, laws, incentives, economic conditions, and social environments.
WIN uses historical evidence to inform present reasoning without assuming that superficially similar events necessarily have identical causes or consequences.
Anecdotes Can Reveal Questions Without Settling Them
Individual experiences and case examples can reveal possibilities, identify failure modes, expose unexpected consequences, or suggest questions worth investigating.
They usually cannot establish how common an effect is or whether the same explanation applies broadly.
WIN therefore treats anecdotes as potentially useful evidence while keeping conclusions proportionate to what those observations can actually establish.
Revision Is Part of the Research Process
Changing a conclusion because better evidence became available is not a failure of the research process. Refusing to change a conclusion after its factual basis has weakened would be the greater failure.
WIN therefore treats important conclusions as revisable. New evidence can refine definitions, change confidence levels, expose previously unseen interactions, reveal unintended consequences, or require an explanation to be replaced altogether.
Significant revisions should be documented sufficiently to preserve institutional understanding of what changed and why.
The durable element should not be loyalty to a particular answer. It should be loyalty to a process capable of learning.
Continue Exploring
Evidence becomes especially valuable when expectations fail. The next page examines how mistakes, unsuccessful interventions, and unexpected outcomes can reveal weaknesses in an explanation and improve future reasoning.