← All lessons
University physics · Feynman Vol. I · Ch. 6

Probability

This lesson follows chapter 6 of the Feynman Lectures, in which probability appears as a way of putting numbers on our guesses when the information at hand is incomplete. The simulations start with dice and coins, move on to counting paths and to a walker who steps at random, and end with the electron of the hydrogen atom, whose position can only be described by a probability density. Along the way, one question keeps coming back: how far can a result stray from what we expected without anything being wrong with the experiment? The original chapter can be read free of charge on the Caltech website.

  1. 1The die
  2. 2Thirty coins
  3. 3Counting paths
  4. 4The walker
  5. 5The fraction of heads
  6. 6Density
  7. 7Uncertainty
STEP 1

What does “a probability of 1/6” mean?

The bench holds three dice, A, B and C, and one of them has been chosen in secret to be loaded. On that one a six comes up once in three throws, and each of the other faces twice in fifteen. Each click throws all three dice together, and the histograms show the fraction of throws on which each face came up, beside a dashed line at 1/6. After ten throws the three histograms usually look equally lopsided, and a fair die showing three or four sixes is nothing unusual.

The line at 1/6 comes from no throw at all. It comes from looking at a well-made die and finding no reason to favour one face over another, and six faces with no favourite split the total into six equal shares. What that fraction promises is a proportion over many repetitions, not the outcome of any single throw: the observed fraction \(N_A/N\) wanders around 1/6 and tends to close in on it as the number \(N\) of throws grows, without ever standing still.

With 600 throws per die, the expected number of sixes is 100 for a fair die and 200 for the loaded one. A fair die seldom strays from 100 by more than about 20, and the gap between 100 and 200 is wide enough for the histograms, which at first were hard to tell apart, to start telling different stories.

The bars at the bottom answer a different question, namely which of the three is loaded. Before the first throw all three sit at 1/3, because nothing sets one die apart from the others. After every throw the bench asks how likely the observed counts would be if each die in turn were the loaded one, and whichever die piles up extra sixes gains credit. In our simulations, with 10 throws per die the guess passes 95% in fewer than one draw in ten, and with 100 it passes in nearly all of them.

The dice do not change during the experiment, and yet the probability that B is the loaded one rises and falls with every click. This only seems odd while we think of probability as a property of the die. Feynman reads it differently: a probability measures what we are entitled to expect when what we know is incomplete, so it shifts whenever that knowledge shifts. The button that reveals the loaded die makes the point plain, since whoever presses it now assigns 1 to one die and 0 to the other two, although nothing whatsoever has happened to the dice.

\(P(A) \approx \dfrac{N_A}{N}\)\(\langle N_A\rangle = N\,P(A)\)\(N\) is the number of throws, \(N_A\) is how many times outcome \(A\) came up, a six for instance, and \(\langle N_A\rangle\) is the expected number. For a six, \(P(A)\) is 1/6 on a fair die and 1/3 on the loaded die of the bench.

In Feynman: §6-1 Chance and likelihood ↗

Let's discuss

  • Throw 10 times, look at the guess and throw 10 more. Has the most suspicious die changed? Why does that not point to a fault in the bench?
  • Swap the dice, throw 30 times and try to pick out the loaded one from the histograms alone, before looking at the bars at the bottom. Repeat a few times. How often do you get it right?
  • With 600 throws per die, how many sixes do you expect from each? Check against the frequency of sixes in the statistics.
  • Press “Show the loaded die”. Has the probability that A is loaded changed for you? And for someone who did not see the screen?
Throws per die
Frequency of six (A · B · C)
Most likely guess
STEP 2

How many heads come up in 30 tosses?

Each game on the bench is 30 tosses of a fair coin, shown at the top as circles marked H, for heads, or T, for tails, and the number of heads in each game goes into the histogram below. The red dots over the histogram are the prediction of how many games should land on each number of heads. With few games the bars jump above and below the dots, and after a few hundred they follow them closely.

Before playing, the natural guess is 15, half of 30, and indeed no other number is more likely. Yet exactly 15 comes up in only 14% of games, and over a run of five games the chance of seeing no 15 at all is close to one half. The results spread around 15: about 80% of games fall between 12 and 18, and about 64% in the range from 13 to 17, which is the range \(15 \pm \sqrt{30}/2\) used in the statistics of the bench. Results further out, 9 heads or fewer, or 21 or more, turn up in about 4% of games with nothing wrong with the coin.

Where do these numbers come from? A sequence of 30 tosses is one of \(2^{30}\), just over a billion, all equally likely with a fair coin, and \(\binom{30}{k}\) of them have exactly \(k\) heads. For \(k = 15\) that is just over 155 million, and the ratio comes to 0.14. The same calculation with 100 tosses per game shows that the most likely value, 50, comes up in only 8% of games, while the range from 45 to 55 covers 73%.

This width is what allows us to suspect a coin at all. Someone who sees 18 heads in 30 has no grounds to accuse anybody, since almost one fair game in five gives 18 or more. A fair coin, by contrast, gives 24 heads less than once in a thousand games, and at that point suspicion begins to have some footing. What matters is whether the result fits inside the width that chance usually produces.

\(P(k) = \binom{N}{k}\,2^{-N}\)\(P(15) = \binom{30}{15}\,2^{-30}\)\(\approx 0.14\)\(N\) is the number of tosses per game, \(k\) is the number of heads and \(\binom{N}{k}\) is the number of sequences of heads and tails with exactly \(k\) heads.

In Feynman: §6-2 Fluctuations ↗

Let's discuss

  • Play 100 games of 30 tosses. In what fraction of them did exactly 15 heads come up, and what does the binomial predict?
  • Start again and play one game at a time. How many did it take before the first 15 appeared?
  • Switch to 100 tosses per game. Does the fraction of games with exactly half heads go up or down? And the fraction inside the range?
  • A classmate says their coin gave 19 heads in 30 and must therefore be loaded. What fraction of your fair games gave 19 or more?
Tosses per game
Games
Mean heads
Exactly 15
Between 13 and 17
STEP 3

In how many ways can 2 heads come up in 3 tosses?

In the Paths scenario, an upright version of Feynman's path diagram, each row of the figure is a toss, each step down to the right is a head, H, and each step down to the left a tail, T, with the letters from step 2. The number inside each node counts the paths that reach it from the top. With 3 tosses the bottom row reads 1, 3, 3, 1, and tapping the node for 2 heads lights up the three paths that lead there, THH, HTH and HHT.

With a fair coin, each of the 8 possible sequences of 3 tosses has the same chance, 1/8, and the probability of 2 heads is simply the fraction of sequences ending at that node, 3/8. The rule works for any number of tosses: there are \(2^N\) equally likely sequences, and \(\binom{N}{k}\) of them have exactly \(k\) heads. This is the calculation behind the 14% of games with 15 heads in the previous step.

Filling in the figure does not require listing a single sequence. A path can only reach a node from one of the two nodes just above it, either the one on the left, with a head on the last toss, or the one on the right, with a tail. So each number is the sum of the two sitting above it, and the whole figure is Pascal's triangle. With 6 tosses, the middle node collects 10 + 10 = 20 paths out of a total of 64.

The Galton scenario does the same count with balls, on a board like the one Francis Galton built in the nineteenth century. There are 10 rows of pins, and at each pin the ball goes left or right with chance 1/2, just like a coin. The bin it lands in records how many times it went right, that is, how many heads came up in 10 tosses. The middle bin collects close to a quarter of the balls, because 252 of the 1,024 paths end there, whereas each end bin has a single path.

With 200 balls the bins trace a bell that follows the red curve, and bins 4 to 6 together hold about 66% of the balls. No ball has to know about the curve for it to appear; it is enough that far more paths lead to the middle than to the edges. The two end bins, for their part, usually stay empty, and in about two thirds of sessions of 200 balls not one ball reaches them.

\(\binom{N}{k} = \binom{N-1}{k-1} + \binom{N-1}{k}\)\(\binom{N}{k} = \dfrac{N!}{k!\,(N-k)!}\)\(\binom{N}{k}\) is the number of sequences of \(N\) tosses with \(k\) heads, and \(N! = N\,(N-1)\cdots 2\cdot 1\). The first equality is the rule of adding up the two nodes above.

In Feynman: §6-2 Fluctuations ↗

Let's discuss

  • With N = 4, predict the numbers in the bottom row before looking. Which node gets the most paths?
  • With N = 6, tap the node for 0 heads. How many paths reach it, and what is the probability of six tails in a row?
  • Drop 200 balls and compare the middle bin with the prediction, about 49. Repeat a few times. How much does the number vary?
  • If the board had 20 rows, how many paths would lead to each end bin? And how many in all?
Scenario
Paths to the node
Balls at the bottom
STEP 4

How far does the random walker stray?

Each walker on the bench starts at \(D = 0\) and, at every step, moves one unit up or down according to the toss of a coin. With 200 walkers the traces open into a fan that follows the curves \(\pm\sqrt{N}\), and after 400 steps the measured \(D_{\text{rms}}\) lands between 18 and 22 in about 95% of sessions, close to \(\sqrt{400} = 20\). With 20 walkers the same measurement varies a good deal more, from 15 to 25 in nine sessions out of ten, and with a single walker it falls anywhere between 2 and 40 in the same proportion.

The walker is the coin of step 2 seen from another angle. If each head is a step up and each tail a step down, \(D\) is the number of heads minus the number of tails, \(D = 2k - N\). Each extra head moves the walker two steps, which is why the range \(15 \pm \sqrt{30}/2\) of step 2 is the same thing as \(D\) within \(\pm\sqrt{30}\).

The mean of \(D\) stays near zero throughout, because every path that climbs has an equally likely twin that descends. It tells us which way the group tends to drift, which is no way in particular, but says nothing about how far each walker wanders off. So the bench measures the root mean square, \(D_{\text{rms}} = \sqrt{\langle D^2\rangle}\), in which a walker 20 steps below the origin counts as much as one 20 steps above.

The plot of \(\langle D^2\rangle\) against \(N\) comes out almost straight, and the reason lies in a single step. A walker at \(D\) moves to \(D+1\) or to \(D-1\), and the square becomes \(D^2 + 2D + 1\) or \(D^2 - 2D + 1\). Averaging the two outcomes, the \(2D\) term cancels and \(D^2 + 1\) is left, so every step adds exactly 1 to the mean square, wherever the walker happens to be. Starting from zero, \(\langle D_N^2\rangle = N\), and the typical distance grows as \(\sqrt{N}\).

The same arithmetic describes a speck of dust jostled at random by the molecules of the air, which is Brownian motion, and the sum of many small sources of error in a measurement, each with a sign nobody can predict. The errors never cancel completely: the sum of \(N\) independent contributions of size 1 tends to have size \(\sqrt{N}\), smaller than \(N\) but growing without limit. For the same reason, a particle spreading at random needs four times as many steps to get twice as far.

\(\langle D_N^2\rangle = \langle D_{N-1}^2\rangle + 1\)\(= N\)\(D_{\text{rms}} = \sqrt{N}\)\(D_N\) is the distance from the origin after \(N\) steps, measured in steps, and \(\langle\;\rangle\) is the average over many walkers.

In Feynman: §6-3 The random walk ↗

Let's discuss

  • With 1 walker, press Walk several times. How often does it end up outside ±20?
  • With 200 walkers, what fraction ends up inside the curves \(\pm\sqrt{N}\)? The message below the bench tells you.
  • If each walker took 1,600 steps, what would you expect for \(D_{\text{rms}}\)?
  • Why is the mean of \(D\) of no use for saying how far the walkers get?
Steps N
Mean D
Measured Drms
√N
STEP 5

Why does the fraction of heads go to 1/2?

The bench tosses three fair coins 10,000 times each and follows the three sequences along an axis of \(N\) on a logarithmic scale, on which each division multiplies the number of tosses by ten. The upper plot shows the fraction of heads \(N_H/N\) of each sequence, and the lower one the excess of heads over half, \(N_H - N/2\). The fractions start scattered between 0 and 1 and close in around 1/2, while the excesses do the opposite and drift away from zero, often reaching a few dozen heads by the end.

The excess of heads is the walker of step 4 taking half-unit steps. Each head raises \(N_H - N/2\) by ½ and each tail lowers it by ½, so the excess equals \(D/2\) and its typical distance from zero, measured by the root mean square, is \(\tfrac12\sqrt{N}\). The fraction departs from 1/2 by that same excess divided by \(N\), and \(\sqrt{N}/2\) divided by \(N\) gives \(1/(2\sqrt{N})\), which shrinks as \(N\) grows. The excess grows, but the number of tosses dividing it grows faster.

The red curves on the upper plot mark this typical distance around 1/2, the solid one at \(\tfrac12 \pm 1/(2\sqrt{N})\) and the dashed one at twice that. Once \(N\) passes a few hundred, the fraction of a sequence tends to lie inside the solid curve about 68% of the time and inside the dashed one about 95%. At \(N = 10{,}000\) the dashed curve sits at 0.5 ± 0.01, too narrow to make out on a scale from 0 to 1, so from \(N = 1{,}000\) onwards a box in the corner of the plot magnifies the range from 0.47 to 0.53. All three sequences end inside the dashed curve in roughly 87% of sessions.

This tightening is what gives meaning to the definition of step 1, in which probability is the proportion of cases over many repetitions. No finite sequence fixes that proportion, but the band in which the fraction wanders narrows as \(1/\sqrt{N}\), and so the measured fraction converges while the counts that make it up keep drifting apart. Anyone watching only the excess might conclude that the coin is getting worse, and anyone watching the fraction sees the opposite, although both are looking at the same sequence.

The same calculation tells us what it costs to measure a probability. For the width \(1/(2\sqrt{N})\) to be 0.01, that is, 1% on a probability near 1/2, \(\sqrt{N}\) has to be 50, which takes about 2,500 tosses. This width, the root mean square of the deviation, is what is called the standard deviation, and with 2,500 tosses the measured fraction lies within 0.01 of the true one about two times in three. Getting to 95% requires 0.01 to be two widths, which calls for the 10,000 tosses on the bench, and every further digit of precision costs a hundred times as many tosses.

\(N_H - \tfrac{N}{2} \sim \tfrac12\sqrt{N}\)\(\dfrac{N_H}{N} - \tfrac12 \sim \dfrac{1}{2\sqrt{N}}\)\(N_H\) is the number of heads in \(N\) tosses, and \(\sim\) indicates the typical size of the deviation, measured by the root mean square like the \(D_{\text{rms}}\) of step 4. That size is the standard deviation.

In Feynman: §6-3 The random walk ↗

Let's discuss

  • Pause near \(N = 100\). What is the largest excess among the three sequences, and how far does the corresponding fraction stray from 1/2?
  • At \(N = 10{,}000\), how many of the three fractions ended within 0.5 ± 0.01? Repeat a few times.
  • A coin gave 5,200 heads in 10,000 tosses. How many widths \(1/(2\sqrt{N})\) is that, and would you be suspicious of it?
  • How many tosses would it take to know the probability of heads to within 0.1%, in the same one-standard-deviation sense?
Tosses N
Fraction of heads
Excess NH − N/2
STEP 6

What is the probability of landing in an interval?

In the Sum of steps scenario, each walker takes steps of random length, any value between \(-\sqrt{3}\) and \(+\sqrt{3}\) with equal chance, a range chosen so that the mean square of one step is 1, as with the unit steps of step 4. The bench adds up the steps of 5,000 walkers and draws a histogram of the final positions \(x\). With 1 step the histogram is a rectangle, with 2 it turns into a triangle, with 5 it already resembles the bell of the red curve, and with 30 the two can hardly be told apart.

The width grows as \(\sqrt{N}\) by the argument of step 4: at every step, the mean square of the position goes up by the mean square of the step, which here is 1, and after \(N\) steps \(\langle x^2\rangle = N\). The square root of that number is the standard deviation \(\sigma = \sqrt{N}\). What is surprising is the shape, because the step on the bench has nothing bell-like about it, and yet the sum of only a few of them already looks like a Gaussian of width \(\sigma\). This result, the central limit theorem, holds for almost any step shape, provided the steps are independent, have zero mean and have a finite mean square.

Once step lengths are continuous, the chance of ending at exactly \(x = 0\), or at any other number picked beforehand, is nil, and none of the 5,000 sums hits a fixed number in every decimal place. That is why the vertical axis shows not probabilities but a density \(p(x)\), with each bar as tall as the fraction of sums that fell in it divided by the width of the bar. The probability of landing in an interval is then the area under \(p(x)\) between its edges, and the whole area comes to 1.

The shaded band is that interval, and its edges can be dragged with a finger or the mouse. Between \(-\sigma\) and \(+\sigma\) the Gaussian predicts 68.3% of the sums, and between \(-2\sigma\) and \(+2\sigma\), 95.4%. With 1 step the count gives about 58% within \(\pm\sigma\), since the rectangle has no tails, with 2 steps about 65%, and from 5 onwards the gap to 68.3% is no larger than the noise in 5,000 sums.

In the Gas scenario, the curve is the speed density of nitrogen molecules, the Maxwell distribution. At 300 K the most probable speed, at the peak of the curve, is 422 m/s, the mean is 476 m/s and the root mean square, the rms speed, is 517 m/s, all a little above the speed of sound in air, about 350 m/s. The three differ because the curve is not symmetric. It starts at zero, since a speed close to nothing requires all three components of the velocity to be small at once, and it has a long tail on the side of high speeds.

Temperature shifts and widens the curve at the same time. The three speeds grow as \(\sqrt{T}\) and at 1,200 K they are twice those at 300 K, and since the area stays at 1, the peak comes down. The area is read just as for the sums of steps: the fraction of molecules between 400 and 600 m/s drops from 36% at 300 K to 13% at 1,200 K, while the fraction above 1,000 m/s climbs from about 1% to 42%.

\(P(a<x<b) = \int_a^b p(x)\,dx\)\(p(x) = \dfrac{1}{\sigma\sqrt{2\pi}}\,e^{-x^2/2\sigma^2}\)\(f(v) = 4\pi\left(\dfrac{m}{2\pi kT}\right)^{3/2}\)\(v^2\,e^{-mv^2/2kT}\)\(p(x)\) is the probability density of \(x\), \(\sigma = \sqrt{N}\) is the standard deviation of the sum of \(N\) steps with mean square 1, \(f(v)\) is the speed density, \(m\) is the mass of a molecule, \(k\) is Boltzmann's constant and \(T\) is the absolute temperature.

In Feynman: §6-4 A probability distribution ↗

Let's discuss

  • With 1 step, put the edges at \(\pm\sigma\). Why does the count stay below the Gaussian's 68%?
  • From how many steps on does the histogram stop resembling the shape of a single step?
  • With 30 steps, what fraction of the sums lands more than \(2\sigma\) from the origin, on either side?
  • In the gas, at what temperature does the most probable speed of nitrogen reach 844 m/s? And what happens to the fraction between 400 and 600 m/s?
Scenario
Steps per walker
Predicted by the Gaussian
Counted in the 5,000 sums
STEP 7

Can we know both where an electron is and how fast it moves?

In the Packet scenario, the curve on the left is the density \(p_1(x)\) of a particle's position, with the width \(\Delta x\) set by the slider, and the one on the right is the density \(p_2(v)\) of its velocity, with the smallest width \(\Delta v\) that quantum mechanics allows for that \(\Delta x\). With \(\Delta x = 10^{-10}\) m, roughly the size of an atom, the minimum \(\Delta v\) is 579 km/s for an electron, 315 m/s for a proton and \(5.3\cdot10^{-10}\) m/s for a dust grain of \(10^{-15}\) kg.

All three numbers come from one rule, which sets a floor of \(\hbar/2m\) under the product \(\Delta x\,\Delta v\), where \(m\) is the mass and \(\hbar \approx 1.05\cdot10^{-34}\) J·s is Planck's constant divided by \(2\pi\). This inequality is Heisenberg's uncertainty principle, in the form Feynman writes it, with the velocity standing in for the momentum \(mv\). Dragging the slider to the left narrows the position curve and widens the velocity curve in the same proportion, since dividing \(\Delta x\) by ten multiplies the minimum \(\Delta v\) by ten, and the axis scales change in jumps, from 1 to 2, 5 and 10, so that both curves still fit on the screen. The mass in the denominator accounts for the distance between the three particles: the proton has 1,836 times the mass of the electron, and the dust grain about \(10^{15}\) times.

For the dust grain, the rule holds but leaves no trace anyone could measure. Even with \(\Delta x\) of 1 µm, the size of the grain itself, the minimum \(\Delta v\) is \(5.3\cdot10^{-14}\) m/s, while thermal agitation in air at 300 K gives the grain speeds of around 2 mm/s, more than ten orders of magnitude higher. For an electron held inside an atom the situation is reversed, and a minimum \(\Delta v\) of hundreds of km/s means that no description in which the electron sits still at a known point can be right.

The two curves of the Packet are not a summary of sharply defined velocities that we have merely stopped tracking, as with the Maxwell curve of step 6, where each nitrogen molecule has its own speed and the density is a bookkeeping device forced on us by having far too many molecules to follow. For the electron on the bench, \(p_1(x)\) and \(p_2(v)\) are the fullest description quantum mechanics has to offer, with nothing sharper hidden underneath, and that is why the minimum width of \(p_2(v)\) does not shrink however carefully we measure.

In the Hydrogen scenario, each dot is a possible position of the electron in the ground state of the atom, and together they form the cloud Feynman uses to picture the atom at the end of the chapter. The dots are drawn in three dimensions and projected onto the screen, as in a photograph, so many of them appear closer to the proton than they really are. The histogram on the right uses the true distance \(r\) of each dot from the proton and follows the curve \(P(r)\), whose peak, the most probable radius, sits at the Bohr radius \(a_0 \approx 0.53\cdot10^{-10}\) m, while the mean radius is \(1.5\,a_0\). The mean of the dots approaches this value with a typical error that falls as \(1/\sqrt{N}\), about 0.09 \(a_0\) with 100 dots and 0.009 \(a_0\) with 10,000.

The cloud has width \(a_0\) in each direction, and by the rule the electron's \(\Delta v\) inside the atom cannot drop below about 1,100 km/s. In the ground state, \(\Delta v\) in each direction is close to 1,260 km/s, and the product \(\Delta x\,\Delta v\) is only some 15% above the minimum the rule allows.

\(\Delta x\,\Delta v \ge \dfrac{\hbar}{2m}\)\(P(r)\,dr = \dfrac{4r^2}{a_0^3}\,e^{-2r/a_0}\,dr\)\(\Delta x\) and \(\Delta v\) are the widths, measured by the standard deviation, of the position and velocity densities, and \(m\) is the mass of the particle. \(P(r)\,dr\) is the probability of finding the electron at a distance between \(r\) and \(r + dr\) from the proton: the density per unit volume, proportional to \(e^{-2r/a_0}\), times the volume \(4\pi r^2\,dr\) of the shell.

In Feynman: §6-5 The uncertainty principle ↗

Let's discuss

  • With the electron, take \(\Delta x\) down to \(10^{-12}\) m. What fraction of the speed of light is the minimum \(\Delta v\)?
  • Swap the electron for the proton without touching \(\Delta x\). Why does the minimum \(\Delta v\) fall by a factor of exactly 1,836?
  • What \(\Delta x\) would bring the minimum \(\Delta v\) of the dust grain up to 1 mm/s? Does that fit on the scale of the slider?
  • In Hydrogen, the histogram of \(r\) drops to zero near the proton. Why, then, is the projected cloud densest at the centre?
Scenario
Particle
Minimum Δv

Where we go next

The lesson ended with the electron described by a cloud of probability, and Feynman's next chapter, The Theory of Gravitation, goes back to ground where classical mechanics gets things right with remarkable precision: Newton's gravitation and the motion of the planets. It is also on the Caltech website, and the lesson in this track that goes with it is The Theory of Gravitation.