Statistics formula solvers

Addition Rule (Mutually Exclusive Events)

P(A∪B)=P(A)+P(B)P(A \cup B) = P(A) + P(B)

ProbabilityStatisticsFor events that cannot both happen, the chance that either one occurs is simply the sum of their separate probabilities.

Bayes' Theorem (Two Hypotheses)

P(A∣B)=P(B∣A) P(A)P(B∣A) P(A)+P(B∣Ac) (1−P(A))P(A \mid B) = \frac{P(B \mid A) \, P(A)}{P(B \mid A) \, P(A) + P(B \mid A^{c}) \, \left(1 - P(A)\right)}

ProbabilityStatisticsUpdates a prior belief into a posterior after evidence arrives, weighing the true-positive rate against the false-positive rate.

Binomial Distribution Mean

μ=np\mu = n p

ProbabilityStatisticsExpected number of successes across n independent trials that each succeed with probability p — the mean of the binomial distribution.

Binomial Distribution Variance

σ2=np(1−p)\sigma^{2} = n p (1 - p)

ProbabilityStatisticsSpread of the number of successes across n independent trials, at its largest when the per-trial chance p sits at one half.

Binomial Probability

P(X=k)=(nk)pk(1−p) n−kP(X = k) = \binom{n}{k} p^{k} (1 - p)^{\,n-k}

ProbabilityStatisticsChance of exactly k successes in n independent trials that each succeed with the same fixed probability p, as with coin tosses.

Birthday Problem (All Distinct)

P=N!(N−n)!  N nP = \frac{N!}{(N - n)! \; N^{\,n}}

ProbabilityStatisticsChance that n independent picks from N equally likely options are all different, the engine behind the birthday paradox.

Chance of Meeting a Target Number on One Die

p=d−t+1dp = \frac{d - t + 1}{d}

Games & RatingsStatisticsThe chance that a single fair die of d faces shows at least the target number t. Counting the faces that succeed is the whole derivation: there are d − t + 1 of them, and the plus one is the step everybody drops the first time.

Chebyshev's Inequality

P≥1−1k2P \ge 1 - \frac{1}{k^{2}}

StatisticsAlgebraThe least possible fraction of ANY data set lying within k standard deviations of its mean — no bell curve assumed, no shape assumed at all. At k = 2 it guarantees three quarters.

Chi-Square Contribution of One Cell

χcell2=(O−E)2E\chi^{2}_{\text{cell}} = \frac{(O - E)^{2}}{E}

StatisticsAlgebraHow much a single cell of a contingency or goodness-of-fit table adds to the chi-square statistic, from its observed and expected counts.

Classical Probability

P=fnP = \frac{f}{n}

ProbabilityStatisticsProbability of an event as the number of favourable outcomes divided by the total number of equally likely outcomes.

Coefficient of Determination (R²)

R2=r2R^{2} = r^{2}

StatisticsAlgebraThe share of variation in y explained by a simple linear regression, obtained by squaring the correlation coefficient.

Coefficient of Variation

CV=sxˉCV = \frac{s}{\bar{x}}

StatisticsAlgebraRelative variability: the standard deviation expressed as a fraction of the mean, so spreads measured on different scales can be compared.

Cohen's d (Effect Size)

d=xˉ1−xˉ2spd = \frac{\bar{x}_1 - \bar{x}_2}{s_p}

StatisticsAlgebraStandardised effect size: the gap between two group means measured in pooled standard deviations rather than raw units.

Complement Rule

P(Ac)=1−P(A)P(A^{c}) = 1 - P(A)

ProbabilityStatisticsThe chance an event does not happen is one minus the chance it does, because every trial must end in one case or the other.

Conditional Probability

P(A∣B)=P(A∩B)P(B)P(A \mid B) = \frac{P(A \cap B)}{P(B)}

ProbabilityStatisticsThe chance of A once B is known to have happened, found by rescaling the overlap to the reduced sample space B.

Confidence Interval Lower Limit

L=xˉ−zσnL = \bar{x} - z \frac{\sigma}{\sqrt{n}}

StatisticsAlgebraThe lower bound of a confidence interval for a mean, pulling the critical z-value and standard error back from the sample mean.

Confidence Interval Upper Limit

U=xˉ+zσnU = \bar{x} + z \frac{\sigma}{\sqrt{n}}

StatisticsAlgebraThe upper bound of a confidence interval for a mean, adding the critical z-value times the standard error to the sample mean.

Degrees of Freedom (One-Sample t)

df=n−1df = n - 1

StatisticsAlgebraThe degrees of freedom used to look up a critical value for a one-sample t-test or confidence interval, one fewer than the sample size.

Elo Expected Score

EA=11+10 (RB−RA)/400E_A = \frac{1}{1 + 10^{\,(R_B - R_A)/400}}

Games & RatingsStatisticsThe share of a point a player is expected to take against a given opponent, worked out from the difference between their two ratings alone. Half a point means an even match; the curve rises toward one as the gap grows and never quite reaches it. A 200-point advantage is worth about 0.76, which is the one number most players already carry in their heads.

Elo Rating Change After a Game

R′=R+K (S−E)R' = R + K \, (S - E)

Games & RatingsStatisticsThe whole of an Elo update in one line: take the difference between what actually happened and what was expected, multiply by the K-factor, and add it to the old rating. Every point one player gains the other loses, so a closed pool's total rating never changes.

Expected Successes in a Dice Pool

E=n (d−t+1)dE = \frac{n \, (d - t + 1)}{d}

Games & RatingsStatisticsThe average number of dice in a pool of n that meet or beat a target number t. It is the per-die chance multiplied by the number of dice — the binomial mean, arrived at without ever needing the binomial distribution, because expectation adds whether or not the dice are independent.

Expected Sum of Several Dice

E=n (d+1)2E = \frac{n \, (d + 1)}{2}

Games & RatingsStatisticsThe long-run average total from rolling n identical dice of d faces each. One die averages the midpoint of its faces, and expectation adds, so n of them average n times that — which is why the answer lands on a half whenever the number of dice is odd.

Expected Trials Until First Success

E[X]=1pE[X] = \frac{1}{p}

ProbabilityStatisticsAverage number of independent attempts needed before the first success when each attempt succeeds with probability p.

Expected Value of a Bet

E=p W−(1−p) LE = p \, W - (1 - p) \, L

ProbabilityStatisticsAverage profit per play of a two-outcome wager that pays W with probability p and costs L the rest of the time, over many plays.

Expected Value of an Exploding Die

E=d (d+1)2 (d−1)E = \frac{d \, (d + 1)}{2 \, (d - 1)}

Games & RatingsStatisticsThe long-run average of a die that is rolled again and added whenever it lands on its highest face, with the rerolls themselves able to explode without limit. An ordinary six-sided die averages 3.5; the same die exploding averages 4.2, and the whole of that extra 0.7 comes from a geometric series that converges because each further explosion is six times rarer than the last.

Expected Value of the Higher of Two Dice

Emax⁡=(d+1)(4d−1)6dE_{\max} = \frac{(d + 1)(4d - 1)}{6d}

Games & RatingsStatisticsThe long-run average when two identical dice are rolled and only the larger is kept. On six-sided dice it is 161/36, about 4.47, against 3.5 for a single die — the extra is what an advantage on a roll is actually worth, and it is smaller than most people guess.

Expected Value of the Lower of Two Dice

Emin⁡=(d+1)(2d+1)6dE_{\min} = \frac{(d + 1)(2d + 1)}{6d}

Games & RatingsStatisticsThe long-run average when two identical dice are rolled and only the smaller is kept. On six-sided dice it is 91/36, about 2.53, against 3.5 for a single die. Added to the average of the higher die it gives exactly d + 1, and that identity is a complete proof that both formulas are correct.

General Addition Rule

P(A∪B)=P(A)+P(B)−P(A∩B)P(A \cup B) = P(A) + P(B) - P(A \cap B)

ProbabilityStatisticsThe chance that either of two events happens, correcting the simple sum by subtracting the overlap that would be counted twice.

General Multiplication Rule

P(A∩B)=P(A) P(B∣A)P(A \cap B) = P(A) \, P(B \mid A)

ProbabilityStatisticsChance that both events happen when the second depends on the first, as in drawing two cards without replacement from a deck.

Geometric Distribution (First Success)

P(X=k)=(1−p) k−1pP(X = k) = (1 - p)^{\,k-1} p

ProbabilityStatisticsChance that the first success in a run of repeated independent trials arrives exactly on trial number k, after k - 1 failures.

Glicko Expected Score Against One Opponent

E=11+10 −g (r−rj)/400E = \frac{1}{1 + 10^{\,-g \, (r - r_j)/400}}

Games & RatingsStatisticsThe Elo curve with the rating difference first shrunk by the opponent's uncertainty. When the opponent's rating is exactly known the two agree; when it is not, this prediction sits closer to an even result, because a rating you cannot trust cannot support a confident forecast.

Glicko g Attenuation Factor

g(RD)=11+3q2RD2π2,q=ln⁡10400g(RD) = \frac{1}{\sqrt{1 + \dfrac{3 q^{2} RD^{2}}{\pi^{2}}}}, \qquad q = \frac{\ln 10}{400}

Games & RatingsStatisticsThe factor by which a rating difference is discounted when the opponent's own rating is uncertain. It is one when the opponent is perfectly known and falls toward zero as their rating deviation grows, which is Glicko's way of saying that a prediction can only be as confident as the weaker of the two ratings involved.

Glicko Rating Deviation Growth While Idle

RD=RD0 2+c2tRD = \sqrt{RD_0^{\,2} + c^{2} t}

Games & RatingsStatisticsHow the uncertainty attached to a rating grows during a layoff. Glickman's insight was that a rating is a claim with an error bar, and that the error bar widens whenever a player stops producing evidence — so a 1900 who last played eight years ago is a far weaker claim than a 1900 who played last week.

Glicko Rating Update (One Game)

r′=r+q g (s−E)1RD2+q2g2E(1−E)r' = r + \frac{q \, g \, (s - E)}{\dfrac{1}{RD^{2}} + q^{2} g^{2} E (1 - E)}

Games & RatingsStatisticsGlicko's replacement for Elo's fixed K-factor: the step size is worked out from how well the rating was already known. A player with a wide rating deviation moves a long way on one result; a player whose rating rests on hundreds of games barely moves at all. The two terms in the denominator are exactly that trade — prior knowledge against the information one game carries.

Hypergeometric Probability

P(X=k)=(Kk)(N−Kn−k)(Nn)P(X = k) = \frac{\binom{K}{k} \binom{N - K}{n - k}}{\binom{N}{n}}

ProbabilityStatisticsChance of drawing exactly k successes in a sample of n taken without replacement from a population of N holding K successes.

Interquartile Range (IQR)

IQR=Q3−Q1IQR = Q_3 - Q_1

StatisticsAlgebraThe width of the middle half of a data set, from the first quartile to the third, and the spread measure box plots are built on.

Lincoln-Petersen Mark-Recapture Estimate

N=M CRN = \frac{M \, C}{R}

Ecology & BiodiversityStatisticsPopulation size estimated from two samples: mark M individuals, later catch C, and count how many of those carry a mark. Assumes a closed population, equal catchability and no lost marks — every one of those failing biases the answer, and the Chapman correction is what small samples actually use.

Log5 Matchup Probability

p=pA−pApBpA+pB−2pApBp = \frac{p_A - p_A p_B}{p_A + p_B - 2 p_A p_B}

Games & RatingsStatisticsPredicts a head-to-head result from two win percentages against a common field. A .500 side beats a .400 side exactly 60% of the time, an even pair gives 0.500 whatever their shared strength, and the whole thing is the Bradley-Terry paired-comparison model wearing different clothes.

Logistic Population Growth

N(t)=K1+K−N0N0 e−rtN(t) = \frac{K}{1 + \dfrac{K - N_0}{N_0}\,e^{-rt}}

Ecology & BiodiversityStatisticsPopulation at time t under density-dependent growth toward a carrying capacity K. Growth is nearly exponential while the population is small, slows through an inflection at K/2, and flattens as the resource runs out.

Margin of Error for a Mean

E=zσnE = z \frac{\sigma}{\sqrt{n}}

StatisticsAlgebraHalf-width of a confidence interval for a mean, built from the critical z-value, the standard deviation and the sample size.

Margin of Error for a Proportion

E=zp(1−p)nE = z \sqrt{\frac{p (1 - p)}{n}}

StatisticsAlgebraThe plus-or-minus quoted with a poll result, built from the critical z-value, the sample proportion and the number of respondents.

Median and Quartile Position

Lq=q (n+1)4L_q = \frac{q\,(n + 1)}{4}

StatisticsAlgebraWhich observation in a sorted list is the median or a quartile: the position, counting from the smallest. q = 2 gives the median, q = 1 and q = 3 the lower and upper quartiles.

Midrange

M=xmax⁡+xmin⁡2M = \frac{x_{\max} + x_{\min}}{2}

StatisticsAlgebraThe midpoint between the largest and smallest observations, a quick centre estimate computed from just the two extremes.

Multiplication Rule (Independent Events)

P(A∩B)=P(A) P(B)P(A \cap B) = P(A) \, P(B)

ProbabilityStatisticsWhen one event has no influence on the other, the chance that both occur is the product of their separate probabilities.

Normal Probability Density

f(x)=1σ2πe−(x−μ)22σ2f(x) = \frac{1}{\sigma \sqrt{2\pi}} e^{-\frac{(x - \mu)^{2}}{2\sigma^{2}}}

StatisticsAlgebraThe height of the normal bell curve at a given value, set by the distance from the mean in standard deviations.

Odds and Probability

O=P1−PO = \frac{P}{1 - P}

ProbabilityStatisticsConverts between a probability and odds in favour, the ratio of the chance it happens to the chance it does not.

One-Sample T-Test Statistic

t=xˉ−μs/nt = \frac{\bar{x} - \mu}{s / \sqrt{n}}

StatisticsAlgebraTests a sample mean against a claimed value when the standard deviation is estimated from the sample itself rather than known.

One-Sample Z-Test Statistic

z=xˉ−μσ/nz = \frac{\bar{x} - \mu}{\sigma / \sqrt{n}}

StatisticsAlgebraTests a sample mean against a claimed population mean when the population standard deviation is known, in standard-error units.

Outlier Lower Fence

LF=Q1−1.5 IQRLF = Q_1 - 1.5 \, IQR

StatisticsAlgebraTukey's lower cutoff for outliers: any observation below one and a half interquartile ranges under the first quartile is flagged.

Outlier Upper Fence

UF=Q3+1.5 IQRUF = Q_3 + 1.5 \, IQR

StatisticsAlgebraTukey's upper cutoff for outliers: any observation above one and a half interquartile ranges beyond the third quartile is flagged.

Percent Difference

PD=∣x1−x2∣(x1+x2)/2PD = \frac{|x_1 - x_2|}{(x_1 + x_2)/2}

StatisticsAlgebraCompares two measurements of equal standing by dividing their gap by their average, when neither counts as the accepted value.

Percent Error

PE=∣xmeas−xacc∣xaccPE = \frac{|x_{\text{meas}} - x_{\text{acc}}|}{x_{\text{acc}}}

StatisticsAlgebraHow far a measurement strays from the accepted value, expressed as a fraction of that accepted value for lab reports and calibration checks.

Performance Rating (Linear Approximation)

Rp=Ravg+400 (W−L)NR_p = R_{avg} + \frac{400 \, (W - L)}{N}

Games & RatingsStatisticsWhat a player's results in one event were worth, expressed on the rating scale: the average rating of the opponents faced, adjusted by 400 points for every whole point of score above or below an even split. This is the linear form used for quick estimates, and it is an approximation of a definition that is otherwise iterative.

Pielou's Evenness (J′)

J′=H′ln⁡SJ' = \frac{H'}{\ln S}

Ecology & BiodiversityStatisticsShannon diversity as a fraction of the largest value it could take for that many species, so a community with every species equally abundant scores 1 and a community dominated by one species scores near 0. A proportion, never a count.

Poisson Probability

P(X=k)=λke−λk!P(X = k) = \frac{\lambda^{k} e^{-\lambda}}{k!}

ProbabilityStatisticsChance of exactly k events in a fixed interval when events occur independently at a constant average rate lambda.

Pooled Standard Deviation

sp=(n1−1)s12+(n2−1)s22n1+n2−2s_p = \sqrt{\frac{(n_1 - 1) s_1^{2} + (n_2 - 1) s_2^{2}}{n_1 + n_2 - 2}}

StatisticsAlgebraCombines two sample standard deviations into a single estimate of common spread, weighting each by its degrees of freedom.

Predicted Value from a Regression Line

y^=a+bx\hat{y} = a + b x

StatisticsAlgebraReads a prediction off a fitted least-squares line for any chosen value of the predictor, given the intercept and slope.

Probability of At Least One Success

P=1−(1−p)nP = 1 - (1 - p)^{n}

ProbabilityStatisticsChance that at least one of n independent attempts succeeds, found as one minus the chance that every single attempt fails.

Pythagorean Expectation

W%=RS kRS k+RA kW\% = \frac{RS^{\,k}}{RS^{\,k} + RA^{\,k}}

Games & RatingsStatisticsEstimates the win percentage a team deserved from the points it scored and the points it allowed, rather than from the games it happened to win. Bill James named it for its resemblance to the Pythagorean theorem when the exponent is two; the resemblance is the only thing the two have in common.

Range (Max minus Min)

R=xmax⁡−xmin⁡R = x_{\max} - x_{\min}

StatisticsAlgebraThe simplest measure of spread in a data set: the distance from the smallest observation to the largest.

Regression Line Intercept

a=yˉ−bxˉa = \bar{y} - b \bar{x}

StatisticsAlgebraThe y-intercept of a least-squares line, fixed by the requirement that the line pass through the point of averages.

Regression Slope from Correlation

b=rsysxb = r \frac{s_y}{s_x}

StatisticsAlgebraThe least-squares slope of a regression line, recovered from the correlation coefficient and the two standard deviations.

Sample Size for a Mean

n=(zσE)2n = \left( \frac{z \sigma}{E} \right)^{2}

StatisticsAlgebraHow many observations a study needs to estimate a mean within a target margin of error at a chosen confidence level.

Sample Size for a Proportion

n=z2 p(1−p)E2n = \frac{z^{2} \, p (1 - p)}{E^{2}}

StatisticsAlgebraHow many respondents a survey needs to estimate a percentage within a target margin of error at a chosen confidence level.

Shannon Diversity Index (H′)

H′=−∑ipiln⁡piH' = -\sum_{i} p_i \ln p_i

Ecology & BiodiversityStatisticsDiversity of a community of up to four species from their proportional abundances, in nats. It rises with both the number of species and the evenness of their shares, and it is Shannon's information entropy with species where the symbols were.

Simpson's Index of Diversity (1 − D)

1−D=1−∑ipi21 - D = 1 - \sum_{i} p_i^{2}

Ecology & BiodiversityStatisticsProbability that two individuals drawn at random from a community of up to four species belong to DIFFERENT species. This page returns 1 − D, which rises with diversity; Simpson's own D = Σ pᵢ² is the dominance and runs the opposite way.

Species-Area Relationship (S = cA^z)

S=c AzS = c\,A^{z}

Ecology & BiodiversityStatisticsExpected number of species in an area, as a power law fitted to survey data. The exponent z is typically 0.20 to 0.35 for nested samples within a region and 0.25 to 0.45 for true islands; c is a fitted constant whose units depend on z and on the area unit it was fitted in.

Standard Error of a Proportion

SE=p(1−p)nSE = \sqrt{\frac{p (1 - p)}{n}}

StatisticsAlgebraHow much a sample percentage typically wanders from the true population proportion, given the proportion and the sample size.

Standard Error of the Mean

SE=σnSE = \frac{\sigma}{\sqrt{n}}

StatisticsAlgebraHow much a sample mean typically wanders from the true mean, shrinking with the square root of the sample size.

Two-Sample Z-Test Statistic

z=xˉ1−xˉ2σ12n1+σ22n2z = \frac{\bar{x}_1 - \bar{x}_2}{\sqrt{\dfrac{\sigma_1^{2}}{n_1} + \dfrac{\sigma_2^{2}}{n_2}}}

StatisticsAlgebraCompares two independent sample means when both population standard deviations are known, scaled by the combined standard error.

Variance and Standard Deviation

σ2=σ⋅σ\sigma^2 = \sigma \cdot \sigma

StatisticsAlgebraConverts between variance and standard deviation: the variance is the square of the standard deviation, and the deviation is its square root.

Variance of a Sum of Dice

σ2=n (d2−1)12\sigma^{2} = \frac{n \, (d^{2} - 1)}{12}

Games & RatingsStatisticsThe spread of the total from n identical dice of d faces. Because the dice are independent their variances simply add, which is why rolling several small dice gives a much tighter total than rolling one large die of the same average — the classic reason a designer chooses three six-sided dice over one twenty-sided one.

Weighted Mean of Two Groups

xˉ=n1xˉ1+n2xˉ2n1+n2\bar{x} = \frac{n_1 \bar{x}_1 + n_2 \bar{x}_2}{n_1 + n_2}

StatisticsAlgebraCombines the averages of two groups into one overall mean, weighting each group by how many observations it contains.

Z-Score (Standard Score)

z=x−μσz = \frac{x - \mu}{\sigma}

StatisticsAlgebraHow many standard deviations a value sits above or below the mean, turning any measurement into a comparable standard score.