Statistics Diagrams

Last updated:

Box Plot of Eleven Orders

Chart

Eleven order values as dots and as a box plot.

Eleven order values drawn as a box plot Eleven order values from 14 to 310 dollars are plotted as dots on a dollar axis, with ten of them bunched between 14 and 60 and one at 310. The mean of 57.18 dollars falls to the right of all but two orders. Below, the box plot draws a box from the first quartile at 23.50 to the third quartile at 44.50 with the median at 30, whiskers from 14 to 60, a fence at 76 dollars, and the 310 order as a separate outlier point. Eleven orders one dot per order mean $57.18 Box plot the same eleven orders median $30 fence: Q3 + 1.5 × IQR = $76 outlier $310 Q1 $23.50 Q3 $44.50 Box: the middle half of the orders, IQR = $21. Whiskers reach $14 and $60. $0 $40 $80 $120 $160 $200 $240 $280 $320 order value

The 68–95–99.7 Rule

Chart

Shares of a normal distribution within 1, 2, and 3 SD.

A normal curve with bands at one, two, and three standard deviations A symmetric bell curve centered on the mean. The darkest band, from one standard deviation below the mean to one above, holds about 68 percent of values. The band out to two standard deviations holds about 95 percent, and the band out to three holds about 99.7 percent. −3 SD −2 SD −1 SD mean +1 SD +2 SD +3 SD about 68% within 1 SD about 95% within 2 SD about 99.7% within 3 SD A normal distribution shaded by distance from the mean

Center of a Right-Skewed Distribution

Chart

Where the mode, median, and mean fall on a right-skewed curve.

Mode, median, and mean on a right-skewed distribution A curve rises steeply from zero to a peak and then trails off in a long tail to the right. A line at the peak marks the mode. A line to its right marks the median, with half the values on each side. A line further right marks the mean, pulled toward the long tail by the few very large values. A right-skewed distribution a floor at zero and no ceiling mode: the peak median: half the values on each side mean: pulled toward the long tail a few very large values form the long right tail 0 value, such as order value or page load time Left to right on the axis: mode, then median, then mean

A Bimodal Distribution

Chart

Cache hits and misses as two peaks, with the mean between.

A response-time distribution with two peaks and the mean between them A response-time curve has a tall, narrow peak on the left for fast cache hits and a lower, wider peak on the right for slow cache misses. The mean falls in the valley between the peaks, a response time that few requests actually have. Response times for a cached endpoint two populations mixed into one distribution cache hits: fast cache misses: slow mean: in the valley, where few requests fall response time

What Correlation Sees

Chart

Four scatter plots and the correlation coefficient of each.

Four scatter plots with their correlation coefficients Four scatter plots. Points along a rising straight line have a correlation near plus one. A shapeless cloud has a correlation near zero. Points along a strong arch also have a correlation near zero, because the coefficient only measures straight-line relationships. A tight cluster of unrelated points plus one far-off point has a strong positive correlation created by that single point. Straight line points hug a rising line r = +0.98 No relationship a shapeless cloud r = +0.02 Strong curve a clear arch, and r misses it r = +0.05 One extreme point no pattern except one point r = +0.73

Averages of Skewed Data Form a Bell Curve

Chart

Skewed orders, and averages of samples of 5 and 200.

Individual order values next to the distribution of sample averages On the left, individual order values rise steeply from zero and trail off in a long right tail. On the right, averages of many random samples of 5 orders are still somewhat skewed and spread widely around the true average. Averages of samples of 200 orders form a narrow, symmetric bell curve centered on the true average. Individual orders a store's full order history, heavily right-skewed $0 $40 $80 $120 $160 $200 order value true average Averages of repeated samples narrower and more symmetric as samples grow $0 $40 $80 $120 $160 $200 average order value of a sample true average samples of 5 orders: still skewed, wide samples of 200 orders: a narrow bell curve

Repeated 95% Confidence Intervals

Chart

Intervals from 20 samples, and the one that misses.

Twenty 95 percent confidence intervals around a true mean Twenty horizontal intervals, one per sample, are stacked beside a vertical dashed line at the true population mean. Each interval is centered on its own sample mean, so they shift left and right, and each sample's own spread sets its width. Nineteen of the twenty cross the true mean. One interval lies entirely to one side and misses it. 20 samples from the same population each sample gives its own 95% interval true population mean sample 1 sample 2 sample 3 sample 4 sample 5 sample 6 sample 7 sample 8 sample 9 sample 10 sample 11 sample 12 sample 13 sample 14 sample 15 sample 16 sample 17 sample 18 misses the true mean sample 19 sample 20 interval contains the true mean (19 of 20) interval misses it (1 of 20)

The p-value of the Checkout Test

Chart

Differences expected by chance, and how rare +0.6 points is.

The chance distribution of the checkout test difference with the p-value shaded A bell curve centered on zero shows how the difference between the two checkout pages would vary across repeated tests if the pages performed the same. Most of the curve sits between about minus 0.6 and plus 0.6 percentage points. The observed difference of plus 0.6 is marked, and both tails beyond 0.6 points in either direction are shaded, each holding about 1.8 percent. Together they make up 3.6 percent, the p-value of 0.036. If the two pages performed the same how the observed difference would vary across repeated tests −1.0 −0.6 −0.3 0.0 +0.3 +0.6 +1.0 difference in conversion rate, variant minus control (percentage points) observed: +0.6 1.8% 1.8% most tests land near zero Shaded: a gap of 0.6 points or more in either direction. Together, 3.6% of tests, so p ≈ 0.036.

Power and the Two Errors

Chart

Chance curves under no effect and a 0.6-point effect.

Overlapping outcome curves for no effect and a 0.6-point effect Two bell curves of the observed difference in conversion rate overlap. One is centered on zero, for pages that perform the same, and the other on plus 0.6 percentage points, for a variant that is 0.6 points better. A vertical significance bar sits at plus 0.56. The part of the zero-centered curve beyond the bar, about 2.5 percent, is a false positive, and a matching shaded tail below minus 0.56 makes up the rest of the 5 percent significance level. The part of the 0.6-centered curve below the bar, about 45 percent, is an effect the test misses, so the test's power is about 55 percent. The checkout test with 10,000 visitors per group the same test under two possible truths −0.6 0.0 +0.6 +1.2 observed difference in conversion rate (percentage points) significance bar: +0.56 if the pages perform the same if the variant is 0.6 points better one α tail: 2.5% β: missed, about 45% Power is the part of the right curve past the bar: about 55%. A larger sample narrows both curves and shrinks β. The 5% significance level is the two shaded α tails together, one on each side of zero.

Found this useful? Share it:

Share on LinkedIn