Quantifying ATS Winning & Losing Streaks
Setup: a streak is a sequence problem, not just a record
A won-loss record (say, 58-42 ATS) throws away information: it tells you
the total, but not the order. Two systems can have the identical
58-42 record while one alternates W/L almost every game and the other goes
on a 14-game win streak followed by a 9-game losing streak. The methods
below all use the order — treating each decided game as one term in a
chronological sequence of W's and L's, exactly the sequence the site's own
max_win_streak/max_losing_streak calculation walks
through game by game.
1. The Wald–Wolfowitz Runs Test
This is the direct test for "is this sequence actually random, or is
there real streakiness (or anti-streakiness) in it?" A run is a
maximal block of identical results — the sequence W W L L L W L W W W
has 5 runs: WW, LLL, W, L,
WWW. Under the null hypothesis that wins and losses are
independent and identically distributed (i.e. no real streak tendency,
just a fixed win rate applied randomly game to game), the number of runs
has a known distribution.
Let n₁ = number of wins, n₂ = number
of losses, n = n₁ + n₂, and R = the
observed number of runs.
Variance: σ²R = [2 × n₁ × n₂ × (2 × n₁ × n₂ − n)] / [n² × (n − 1)]
Test statistic: Z = (R − μR) / σR
R=18 observed runs.μR = 1 + (2×30×20)/50 = 1 + 24 = 25
σ²R = [2×30×20×(1200−50)] / [50²×49] = 1,380,000 / 122,500 ≈ 11.27 → σR ≈ 3.36
Z = (18 − 25) / 3.36 ≈ −2.08 — comfortably beyond the usual ±1.96 cutoff, so this record has meaningfully fewer runs (longer, clumpier streaks) than 50 independent coin-flip-style bets would typically produce. That's evidence of real streak behavior in the sequence itself — separate from, and complementary to, the system's own win-rate p-value (see the companion article on p-values).
2. Expected Longest Run (Schilling's Approximation)
This answers a different, very practical question: "given how often
this system wins, how long a streak should I expect to see purely
by chance over N games?" — the baseline your actual max_win_streak
number should be compared against before calling it remarkable. For
n independent trials each with success probability
p, the expected length of the longest run of successes is
given by an approximation formalized by Mark F. Schilling ("The Longest
Run of Heads," The College Mathematics Journal, 1990), building on
earlier work by Erdős and Rényi:
q = 1 − p, log1/p(x) = ln(x) / ln(1/p), and γ ≈ 0.5772 is the Euler–Mascheroni constant. (A small oscillating correction term is omitted here — negligible for practical bankroll/streak planning.)log2(100 × 0.5) = log2(50) ≈ 5.64
γ/ln(2) ≈ 0.5772/0.6931 ≈ 0.833
E[L100] ≈ 5.64 + 0.833 − 0.5 ≈ 5.97 games
In plain terms: a perfectly average, no-edge 50% system, run for 100 games, would be expected to produce a win streak of roughly 6 games just from chance. A system advertising a "10-game win streak!" over a similar sample isn't automatically special — Lesson 3 covers this exact trap (small-sample cherry-picking) in the main Learn library.
3. The Geometric Distribution (one streak's own length)
Once a streak has started, how long should any single run of wins
(or losses) last before it breaks? If each game is independent with win
probability p, the length of one streak follows a
geometric distribution:
k games before the first loss. Its expected value, E[k] = 1 / (1 − p), is the average length you'd expect ANY winning streak (not the longest one — see #2 above) to reach before ending.4. Conditional Win Rate After a Win vs. After a Loss (the "Hot Hand" Test)
This is the formal version of the question everyone actually means when they ask "does this system get hot?" — is the win probability actually different immediately after a win versus immediately after a loss? This is the same test structure Gilovich, Vallone & Tversky used in their famous 1985 "hot hand" basketball study (Cognitive Psychology, 17(3), 295–314), applied here to ATS results instead of made/missed shots. It's a two-proportion z-test:
p̂1 = win rate on games immediately following a win (n1 such games), p̂2 = win rate immediately following a loss (n2 such games), and p̂ is the pooled win rate across both groups combined.5. Maximum Drawdown (the dollar-based streak metric)
The four methods above all describe the win/loss sequence. Maximum drawdown, already on the query pages as its own tile, is the equivalent concept measured in real dollars against a flat-stake equity curve rather than in games:
Traders sometimes also compute the Calmar ratio (annualized return divided by maximum drawdown) as a single number combining profitability and downside risk — worth knowing the name if you ever see it elsewhere, though this site reports the two pieces (ROI and max drawdown) separately rather than pre-combined, so you can weigh them yourself.
← Back to the Portfolio Approach lesson · Next: How the P-Value Is Calculated →