Research guide · Sample size

Small samples, variance, and false confidence

A small sample is not automatically useless. It is simply easier for unusual events to dominate. The goal is to match confidence to the amount and quality of evidence rather than treating every average as equally stable.

Count opportunities, not just games

“Last five games” sounds consistent, but it can represent very different amounts of evidence. A starting pitcher may make five appearances with more than one hundred batters faced. A relief pitcher may face fifteen. A football receiver may run many routes but receive only a handful of targets. A bench player may appear in five games without playing meaningful minutes.

The relevant denominator depends on the question. Use plate appearances for hitting rates, attempts for shooting percentages, routes for target opportunity, faceoffs or shot attempts for some hockey questions, and minutes when playing time drives accumulation. Game count is a convenient label, not always the best sample-size measure.

Why short windows swing

An extreme result has more influence when there are few observations. One four-hit baseball game can dominate a five-game average. One overtime basketball game can add five minutes and several counting-stat opportunities. As the sample grows, each unusual event becomes a smaller share of the whole.

Simple illustrationIf a player records 2, 1, 0, 1, and 6 units across five games, the average is 2.0. Remove the six-unit game and the remaining average is 1.0. The headline average doubled because of one result.

This does not make the six-unit game fake. It means the estimate is sensitive. A sensitive estimate should be described with a wider range of uncertainty.

Selection bias can manufacture a trend

Researchers often choose a window after noticing the result: “since March 12,” “in the last seven,” or “when playing on Tuesdays.” If the starting point was selected because it creates an interesting pattern, the same data cannot independently prove that the pattern is meaningful.

Use standard windows before looking at the answer. If you routinely compare five, ten, twenty, and season views, you reduce the temptation to keep adjusting the date until a preferred story appears. Custom windows are still useful when tied to a real event such as a trade, injury return, coaching change, or lineup promotion.

Regression is not the same as reversal

When an unusually high conversion rate moves closer to a player’s longer baseline, that is often called regression toward the mean. It does not require the next performance to be bad. It means an extreme rate is unlikely to remain equally extreme without evidence of a lasting underlying change.

Separate volume from efficiency. Increased minutes, attempts, targets, or plate appearances can remain elevated because of a new role. An extreme shooting, scoring, or conversion percentage may move toward normal even while volume stays higher. The most defensible projection can combine the new opportunity with a less extreme efficiency assumption.

Use relevant baselines

A career average is not always the correct comparison. Skills develop, ages change, leagues change, and roles change. Choose the longest baseline that still resembles the current conditions. For a rookie, that may be only the current season. For a veteran who changed teams and roles, the most relevant baseline may begin after the move.

Opponent splits also require care. Ten career meetings can span years of roster, coaching, and skill changes. Head-to-head history can supply context, but a current role and current opponent environment usually deserve more weight than an old matchup record.

Look for agreement across independent signals

Confidence improves when different types of evidence point in the same direction. More minutes and more shot attempts are related but not identical. A lineup promotion and increased plate appearances reinforce each other. A target increase supported by more routes is stronger than a target increase caused by one broken play.

Avoid counting the same underlying fact several times. Points, field-goal attempts, and usage can all rise because minutes increased. Listing all three does not create three independent reasons unless each contributes distinct information.

Use precise language: “The short sample suggests a possible role change, supported by higher minutes and attempts. The scoring rate remains volatile.” That sentence is more informative than declaring the player hot.

Questions to ask before trusting a split