The P-Value, Start to Finish
What question is a p-value actually answering?
Start with the question it does NOT answer, because this is the most common misreading of a p-value anywhere it's used, not just in sports betting: a p-value is not "the probability this system is real" or "the probability the edge exists." It's the answer to a narrower, more precise question:
A small p-value means the answer is "very unlikely" — which is evidence the no-edge assumption is probably wrong, i.e. evidence of a real edge. It is indirect evidence, not proof, and it says nothing about why the edge exists or whether it will hold going forward. That distinction matters enough that it gets its own callout at the end of this article.
How it's calculated, start to finish
Set up the null hypothesis and count decided games
Every push is thrown out first — a push is neither a win nor a loss,
so it isn't a trial in this test at all. Let n = the number
of DECIDED games (wins + losses only) and k = the number
of those that were wins. The null hypothesis, written formally:
p₀ is the baseline you're testing against — 0.5 for "no edge at all" (a coin flip), or 0.5238 for "doesn't even beat the -110 vig" (see the break-even box below). This site defaults to 0.5 and lets you override it to any baseline in the query page's "P-value baseline" field.The binomial probability of any single outcome
If the null hypothesis is true (win probability really is
p₀ on every independent game), the probability of
landing on exactly k wins out of n decided
games follows the binomial distribution:
C(n, k) = n! / [k! × (n−k)!] is the binomial coefficient — the number of different orderings of wins and losses that produce exactly k wins out of n. This formula gives one single probability; the p-value needs to combine many of these.Sum every outcome "at least as extreme" as what you observed
This is the step that actually makes it a p-value rather than just a
probability. The exact two-sided binomial test (this is the real
name — the site computes this via SciPy's binomtest, the
standard implementation of the classical exact test) doesn't just look
at the one outcome you got; it sums the probabilities of every possible
outcome k′ from 0 to n that is at least as unlikely
as your actual result under the null hypothesis:
Compare the p-value to a threshold
There's no universal "correct" cutoff, but 0.05 is the conventional one across most of science, and this site follows the same convention (with 0.005 shown in a stronger color as "very significant," matching how the site's own p-value tiles are colored). Below the threshold: the record would be a rare event under pure chance, so the no-edge assumption looks wrong. Above it: the record is easily explained by chance alone, and should be treated as "not yet distinguishable from luck" — not as "proven fake," just not yet proven real.
Step 1 — H₀: p = 0.5, n = 100, k = 65.
Step 2 — the single-outcome probability of exactly 65 wins: P(K=65) = C(100,65) × 0.5¹&sup0;&sup0; ≈ 0.00086.
Step 3 — sum P(K=k′) for every k′ from 0 to 100 whose probability is ≤ 0.00086 (this includes both the far-right tail near 65-100 wins AND the mirrored far-left tail near 0-35 wins, since a 35-65 record would be exactly as extreme in the other direction) — the exact binomial calculation gives a two-sided p-value ≈ 0.0035.
Reading it: if this system truly had no edge (a real coin flip repeated 100 times), a record at least this lopsided would only happen about 0.35% of the time. That's well under the conventional 0.05 cutoff — real evidence this 65% win rate isn't just noise.
What a low p-value does NOT mean
This is the section that turns a p-value from a number you glance at into a tool you can actually trust — the site's own p-value note quotes part of this directly, but it's worth the full explanation:
- It doesn't account for how many queries you tried. Test 100 different field combinations against the same data and roughly 5 of them will show p ≤ 0.05 from pure chance alone, even if NONE of them have a real edge — this is the multiple-comparisons problem, covered in full in the main Learn library's "Avoiding Overfitting" lesson. A single query's p-value says nothing about how many candidates you tried before landing on it.
- It's not the same as economic significance. A huge sample can produce a tiny, statistically "significant" p-value for an edge so small it barely covers the vig. Always read the p-value together with ROI and profit, never alone.
- It assumes the games are independent. The math above treats
every decided game as its own coin flip. If a system is secretly betting
on strongly correlated games (e.g. the same team's whole season), the
effective sample size is smaller than
nsuggests, and the true p-value is less extreme than the formula reports. - It says nothing about the future. A p-value is entirely a statement about the sample you already have. It's not a forecast, and it doesn't adjust for whether the underlying market has changed since the data was collected (rule changes, a formerly-inefficient line getting sharper over time, etc.).
← Back to the Reading Your Results lesson · Also see: Quantifying Winning & Losing Streaks →