The gambler's fallacy is one of the best-documented reasoning errors in behavioral science: after a run of reds on a roulette wheel, people feel — genuinely, intuitively — that black is "due." It isn't. Each spin is independent. The wheel has no memory of what it just did, and no obligation to balance itself out over the next few spins. The fallacy is believing otherwise, and it shows up far outside casinos. It shows up constantly in how teams talk about estimates.
What it looks like in delivery work
"We've missed the last four deadlines, we're due for one to land on time." "This team has been on a hot streak, surely the next one runs long." Both statements sound like reasonable pattern-recognition. Both commit the same error the roulette player commits: treating a sequence of past outcomes as evidence about the probability of the next one, when — unless something has actually changed about the work, the team, or the process — the next ticket's odds of running over or under its estimate are governed by the same underlying rate they always were, streak or no streak.
The chart above illustrates this directly. The true probability of any given ticket running over its estimate stays flat at 40% for the entire sequence — nothing about the process changes. A streak of three-in-a-row "over" tickets still happens periodically, purely by chance, the same way three reds in a row happen on a fair roulette wheel far more often than people expect. It's not a sign the streak is about to break. It's not a sign anything is wrong. It's what a constant 40% rate actually looks like when you watch it for twenty tickets instead of assuming it should look tidier than that.
Why this matters more in delivery work than at a casino
At a casino, the fallacy costs the person making the bet. In delivery work, it costs the estimate itself. A team that believes a streak of on-time deliveries means the underlying rate has actually improved — rather than checking whether it has — will keep making commitments based on a rate that hasn't changed, right up until a normal-probability bad streak arrives and the belief gets shattered all at once. The inverse is just as costly: a team convinced that four late deliveries in a row means "we're due" for a smooth one may quietly stop scrutinizing whatever's actually driving the lateness, because the fallacy offers a comforting story — luck will fix it — that doesn't require anyone to look at the data.
How to actually correct for it
Separate "the rate changed" from "the rate is just doing what rates do." A genuine improvement in delivery reliability shows up as a shift in the underlying data — lower cycle time variance, less WIP, fewer blocked tickets — not as a subjective feeling that a streak has gone on long enough. Before concluding the team has gotten better or worse, look for a specific, measurable cause. If there isn't one, the streak is very likely just noise around an unchanged rate.
Let a large enough sample answer the question instead of a short streak. Four tickets is not a reliable sample of anything. Forty is a much better one. This is exactly why Monte Carlo forecasting resamples from the team's actual historical distribution rather than extrapolating from whatever just happened — it's structurally immune to the "we're due" reasoning, because it doesn't treat the most recent handful of outcomes as more predictive than the full distribution behind them.
Track the rate itself, not the streak. An accountability log that records every forecast against its actual outcome is the practical antidote here — it turns "does it feel like we're due for a good one" into "what has our actual on-time rate been over the last fifty tickets," which is a question with a real, checkable answer instead of a gut feeling shaped by whatever happened most recently.
The honest version of "due"
There's a real phenomenon that sounds similar to the fallacy but isn't one: regression to the mean. If a team's true long-run rate is 70% on-time and it's currently on an unusually bad streak, the next several deliveries probably will look better than the streak — not because anything is "due," but because the streak itself was already an unlikely deviation from a stable average, and unlikely deviations don't tend to repeat. The distinction that matters is whether there's an actual stable average to regress toward. If there is, that's just how averages behave. If there isn't — if the process is genuinely noisy or has genuinely changed — there's nothing to be due for at all, in either direction.