Simpson's Index of Diversity (1 − D)

Also known as Simpson index · Gini-Simpson index · Simpson's diversity index · 1 - D · probability of interspecific encounter · PIE · Simpson dominance · Simpson's D

1D=1ipi21 - D = 1 - \sum_{i} p_i^{2}

Enter your known values, leave one input blank, and solves for the missing one. Try different units for next level excitement!

Learning zone

Edward H. Simpson published this in a one-page note in Nature in 1949 while working at Bletchley Park, and it asks a question you can act out: reach into the community twice at random, and what is the probability the two individuals are different species? Sum the squared proportions and you get D, the probability they are the same species — the chance of dominance. Subtract from one and you get the answer this page returns. For a community at 0.4, 0.3, 0.2 and 0.1, D is 0.16 + 0.09 + 0.04 + 0.01 = 0.30, and 1 − D is 0.70.

Which convention does this page return? 1 − D. That matters more than it should, because both quantities are called "Simpson's index" in the literature and they run in opposite directions. Simpson's own D goes UP as a community gets more dominated by one species, so a low D is a diverse site — which reads backwards to everyone the first time. To fix that, ecologists invented two complements: 1 − D, the Gini-Simpson index or "Simpson's index of diversity", which is what you get here; and 1/D, the reciprocal or Simpson's index of dominance, which runs from 1 to S and has the pleasant property of being an effective species count. Three names, three scales, one Greek letter. When you read a Simpson value in a paper, find out which one before you compare it to anything. The answer box on this page also reports D alongside, so you can quote either.

Simpson and Shannon rank sites differently, and the disagreement is structural rather than a rounding difference. Squaring a proportion crushes small numbers: a species at 1% contributes 0.0001 to D and is essentially invisible, while the same species contributes 0.046 nats to Shannon's sum, which is a real share of a typical H′. So Simpson is a measure of the common species and Shannon gives the rare ones a voice. Two sites with the same handful of abundant species and wildly different tails of rarities will look nearly identical to Simpson and clearly different to Shannon. Neither is wrong. Report both, and say which you preferred and why.

The same comparability warning as its neighbour applies, with one useful twist: because Simpson barely notices rare species, it is far less sensitive to sampling effort than Shannon is. If you are stuck comparing surveys of different sizes and cannot redo the fieldwork, Simpson is the less dishonest of the two. It is still not a licence to compare beetles in a pitfall trap against birds on a transect.

One footnote that catches people out. Simpson's original 1949 paper actually gives the index for sampling WITHOUT replacement, Σ n(n − 1)/(N(N − 1)), which is the exactly correct probability when you pull two individuals out of a finite sample and do not put the first one back. The Σ p² form on this page is the infinite-population limit of that, and the two agree to better than a percent once you have a few hundred individuals. Below about fifty, use the finite form; the difference is real and it runs in the direction of making a small sample look more dominated than it is.

Simpson's Index of Diversity (1 − D)
1D=1ipi21 - D = 1 - \sum_{i} p_i^{2}
one quadrat1 − Dtwo draws
Where
  • 1D1 - D= Simpson's index of diversity
  • p1p_1= Proportion of species 1
  • p2p_2= Proportion of species 2
  • p3p_3= Proportion of species 3
  • p4p_4= Proportion of species 4