Pythagorean Expectation

Also known as Pythagorean win percentage · Pythagorean theorem of baseball · expected win percentage · Bill James Pythagorean · Pythagenpat · runs scored runs allowed win percentage · points for points against expected wins

W%=RSkRSk+RAkW\% = \frac{RS^{\,k}}{RS^{\,k} + RA^{\,k}}

Enter your known values, leave one input blank, and solves for the missing one. Try different units for next level excitement!

Learning zone

Bill James noticed, some time in the late 1970s, that a baseball team's win percentage was predicted remarkably well by its runs scored and runs allowed — better, in fact, than by its own record from the first half of the season. The formula he wrote down resembles the Pythagorean theorem when the exponent is two, and he named it accordingly. That resemblance is a pun. There is no geometry here and no theorem; it is an empirical curve that happens to have squares in it.

The exponent is fitted, not derived, and it is not a constant of nature. Nothing produces kk from first principles. It is chosen to make the curve match a large body of completed seasons, and it comes out differently in every sport, because the shape of the scoring distribution differs: around 2 in baseball, where James found it, and materially different in basketball, hockey and football. Later refinements make it depend on the scoring environment within a single sport — the Pythagenpat version sets the exponent from the run-scoring level of the season itself, since a high-scoring era and a low-scoring one bend the curve in different ways. Borrowing a kk from one sport and applying it to another is simply using the wrong curve, and it will be wrong in a plausible-looking way.

What the estimate is FOR is the gap between it and the real record. A team several wins above its Pythagorean estimate has been winning close games and losing blowouts — its runs were distributed inefficiently across the schedule, in the sense that they arrived where they were least needed. Historically, that pattern does not persist: teams above their estimate tend to fall back and teams below it tend to climb. This is a statement about the average of many teams and not a promise about any one of them, and it is regularly overstated. Some of the gap is bullpen quality, some is a manager's use of a roster in close games, and some — probably most — is luck.

The theoretical footing is real but partial. If you model each team's runs per game as a draw from a suitable skewed distribution and ask how often one exceeds the other, something of this shape falls out, which is why the fitted exponent lands near 2 rather than anywhere at all. Nobody derives the working formula that way; it remains a curve fitted to history and justified by how well it forecasts.

A last caution about ratios. Runs scored and runs allowed are season totals, so the estimate assumes the same team played all season. A midseason trade, a long injury or a September of substitutes all break that, and the answer will not know it. And the formula treats a 900-to-800 season exactly like a 90-to-80 one — only the ratio matters — which is convenient and is also why it says nothing whatever about how much evidence there is behind the numbers.

Pythagorean Expectation
W%=RSkRSk+RAkW\% = \frac{RS^{\,k}}{RS^{\,k} + RA^{\,k}}
W%RSRAk
Where
  • W%W\%= Expected win percentage
  • RSRS= Points scored
  • RARA= Points allowed
  • kk= Exponent
Missing one of these? Work it out first, then come back