Why does pi appear in stats?
59sThe anecdote about the statistician's friend questioning pi's appearance in population statistics is relatable and sparks curiosity.
▶ Play Clip"Delivers on the promise of going beyond integral tricks, but the final connection to the Central Limit Theorem is deferred, leaving a slight gap."
This video explores the deep mathematical connection between the constant π and the normal distribution, moving beyond the standard integral trick to provide a more intuitive explanation. It begins with a classic proof showing why π appears in the Gaussian formula, then introduces the Herschel-Maxwell derivation to explain the function's origin, and finally hints at how this connects to the Central Limit Theorem.
The video opens with a story about a statistician and a friend, where the friend questions why π appears in the Gaussian distribution, setting up the central question of the video.
π appears in the normal distribution because the area under the curve e^(-x^2) is √π, which is needed to normalize the distribution so the total area equals 1.
The classic proof involves computing the integral of e^(-x^2) by bumping up to two dimensions, using polar coordinates to exploit circular symmetry, and then relating the volume to the original area.
The proof starts by considering the volume under the surface e^(-(x^2+y^2)) instead of the area under the curve, which simplifies the integral due to the factor 2r from the Jacobian.
The volume under the bell surface is π, and by slicing the volume in a different way, the original area is shown to be √π.
John Herschel (and later Maxwell) derived the Gaussian distribution from two assumptions: rotational symmetry and independence of x and y coordinates, leading to the function e^(-c r^2).
The derivation leads to a functional equation f(r) = f(x)f(y), which forces f to be an exponential function, specifically e^(c x^2), with c negative for normalization.
The video notes that the Herschel-Maxwell derivation presumes a multi-dimensional setting, but the Central Limit Theorem explains why the Gaussian arises in one-dimensional sums, and this connection is deferred to a future video.
The video provides a more intuitive explanation for π's presence in the normal distribution by deriving the Gaussian from first principles, but acknowledges that a full understanding requires connecting this to the Central Limit Theorem, which is left for a follow-up video.
What is the area under the curve e^(-x^2)?
The square root of π.
02:30
Why does π appear in the normal distribution?
Because the area under the curve e^(-x^2) is √π, which is needed to normalize the distribution.
02:30
What is the first step in the classic proof for the Gaussian integral?
Bump the problem up one dimension, considering the volume under the surface e^(-(x^2+y^2)).
05:17
What is the volume under the bell surface e^(-(x^2+y^2))?
π.
08:55
What are the two assumptions in the Herschel-Maxwell derivation?
Rotational symmetry (density depends only on distance from origin) and independence of x and y coordinates.
13:29
What functional equation does the Herschel-Maxwell derivation lead to?
f(r) = f(x)f(y), where r = √(x^2 + y^2).
16:54
What form does the function f take from the functional equation?
An exponential function, specifically e^(c x^2), with c negative for normalization.
20:33
π from the area under the curve
Explains the direct origin of π in the Gaussian formula, a key fact for understanding the distribution.
02:30Bumping up a dimension
A clever problem-solving technique that simplifies the integral, illustrating a general mathematical strategy.
05:17Herschel-Maxwell derivation
Provides a first-principles derivation of the Gaussian, making the appearance of π less surprising.
12:51Solving the functional equation
Shows how the functional equation forces an exponential form, a key step in the derivation.
17:11[00:00] effectiveness of mathematics in the natural sciences. but even more fun than the title is the way that he chooses to open it. who were classmates in high school, talking about their jobs.
[00:18] One of them became a statistician and was working on population trends. and the reprint started, as usual, with the Gaussian distribution. And the statistician explained to the former classmate the meaning of
[00:30] the symbols for the actual population, the average population, and so on. sure whether the statistician was pulling their leg. And what is this symbol over here?
[00:44] What is that? The ratio of the circumference of a circle to its diameter. Surely the population has nothing to do with the circumference of a circle.
[00:59] phenomenon of concepts in pure math seeming to find applications But I would like to stay focused on this particular anecdote
[01:11] and the question that the statistician's friend is getting at. explains the pi inside the formula for a normal distribution. And despite there being a number of other really great explanations online,
[01:24] see some links in the description, I cannot help but indulge in the pleasure of For one thing, there is a fun side note that I didn't learn until recently about how you can use this proof to derive the volumes of higher dimensional spheres.
[01:38] want to do is try to go beyond the classic proof. What I want to ask is, can we find an explanation that would satisfy their disbelief?
[01:50] about a function that was handed down to them on high. have anything to do with population statistics.
[02:02] Until we fully draw that connecting line, we should consider the task incomplete. have some of the backdrop here, because there we broke down the formula for a normal distribution, which is also called a Gaussian distribution.
[02:17] the basic function that describes the bell curve shape is e to the negative x squared. And the reason that pi showed up in the final formula was that the area underneath this
[02:30] curve works out, as you will see in a couple minutes, to be the square root of pi. that square root of pi to make sure that the area under the curve is one,
[02:42] which is a requirement before you can interpret it as a probability distribution. this gets mixed together with some of the other constants, but in its purest form that pi originates from the area underneath this curve.
[02:59] but I want to emphasize it's not the last step. To satisfy the question raised by that hypothetical statistician's friend, We need to also answer why is it that this function e to
[03:13] I mean, there are lots of different formulas you could write down that would give a shape that, you know, vaguely bulges in the middle and tapers out on either side. So why is it that this specific function holds such a special place in statistics?
[03:29] To phrase our goal another way, can we find a connection between the proof that shows why pi shows up and the central limit theorem, which, as we talked about in the last video, is the thing that explains when you can expect a normal distribution to arise in nature.
[03:44] let's dig into the classic and very beautiful proof. the tool for doing that is an integral.
[03:57] you might imagine approximating that area with many different rectangles under the curve, where the height of each such rectangle is the value of the function above that point, in this case, e to the negative x squared for a certain input x,
[04:11] We need to add up the areas of all these rectangles, and the use of that notation dx is kind of meant to imply you shouldn't think of any
[04:24] specific width, but instead you ask, as the chosen width for your rectangles gets thinner and thinner, what does this sum of all those areas approach? Of course, all of that is just notation unless you provide a way to answer that question,
[04:36] and the magic of calculus is that it provides just that, at least usually. You see, usually the procedure here would be to find some function whose derivative is equal to the stuff we have on the inside, e to the negative x squared.
[04:49] The problem is, for this particular function, it is It's a little weird and beyond the scope of what I want to talk about here,
[05:02] it is a well-defined function, you cannot express what that antiderivative is using all our usual tools, like polynomial expressions, trig functions, exponentials, So finding this area requires a bit of cleverness.
[05:17] And the first step to this trick is easily the most absurd. We start by bumping things up one dimension, so that instead of asking for the area under a bell curve, we ask for the volume underneath this kind of bell surface.
[05:32] Who ordered another dimension? other than to say, watch what happens when we just try it. In general, with hard problems, it's never a bad idea to try solving cousins of
[05:45] the problem, since that can help you get a little bit of momentum and insight. it takes in two different inputs, x and y, which we might think of as a point And the way to think about it is to consider the distance from that point to the origin,
[06:01] which I'll label as r, and then to plug in that distance to our original bell You might notice the lines I've drawn on this diagram make a right triangle. So, by the Pythagorean theorem, x squared plus y squared equals r squared.
[06:16] you can think in the back of your mind, that's really the The main thing to notice here is how this gives our function a kind of circular symmetry,
[06:29] in the sense that all of the inputs that sit on a given circle have the same output. it means it has a rotational symmetry about the z-axis. so for our question of computing the volume underneath the surface,
[06:46] and imagine integrating together a bunch of thin little cylinders underneath that Here, making this a little more quantitative, let's focus on just one of those
[06:58] cylindrical shells, where its area is going to be the circumference of that shell times on a soup can that we can unwrap into a rectangle. The circumference of the cylinder, which is the top side of that rectangle,
[07:12] And then the height of our cylinder, the other side of our rectangle, is the height of the surface at this point, which by definition is the value of our function associated with that radius, which like I said earlier you can think of as e to
[07:27] The real way you want to think about this is to give that cylinder a little bit of thickness, which we'll call dr, so that the volume that it represents is approximately that area we just looked at multiplied by this thickness dr.
[07:41] all of these different cylinders as r ranges between 0 and infinity. Or more precisely, we consider what happens as that thickness gets thinner and thinner,
[07:53] different thin cylinders that sit underneath that curve. You might think this is just a harder version of what we were looking at earlier,
[08:05] But actually something very helpful has happened. Now the stuff inside that integral, having picked up this term 2r,
[08:18] We can now apply the usual tactics of calculus. derivative of negative e to the negative r squared.
[08:30] We take that antiderivative and plug in the upper bound, or speaking a little bit more precisely, if you consider the limit of this
[08:43] and we subtract off the value of that antiderivative at the lower bound, So all in all, the whole integral just works out to be 1,
[08:55] which means all we're left with is that factor out in front, pi. Evidently, the volume underneath this bell surface is pi. because the surface has this intrinsic circular symmetry.
[09:10] As I said, throughout math, if you face a hard problem, And in this case, it's helpful not just for building intuition,
[09:22] but we can directly relate the three-dimensional graph to our two-dimensional graph by analyzing the volume in a second, different way. You see, the more general way to approach volumes underneath surfaces is to
[09:35] think of chopping it up into slices that are all parallel to one of the axes. For example, all these slices that are parallel to the x-axis. For example, this right here is a slice that corresponds to the plane y equals 0.
[09:48] and if we write out the function, this should actually make a lot of sense. You could just plug in y equals 0, but to help see what happens with other slices, we could also write our function as e to the negative x squared times e to
[10:03] It factors out nicely. specifically the number 1. So this is the same graph we've seen before, e to the negative x squared,
[10:16] meaning that the area of this slice is exactly the thing that we're looking for. What's nice is there's nothing really special about this particular slice. it corresponds to multiplying this curve by a different number.
[10:34] meaning its area is the same as our mystery constant, just scaled down by some number. Each one of these slices has the same basic shape,
[10:46] is not at all true for most two-variable functions. This is very much dependent on the fact that we were able to factor our function into one part that's just dependent on the y and another part that's just dependent on the x.
[11:02] here's another way we could phrase it. equals negative infinity up to infinity, where the term inside And when we multiply it by a little thickness dy,
[11:19] you might think of it as giving each one of those slices a little bit of volume. And remember, that term c sitting in front represents the thing we want to know, which itself is an integral, a suspiciously similar-looking integral.
[11:32] See, if we take the expression on the top and we factor out that constant c, because it's just a number, it doesn't depend on y, the thing we're left with, the thing we don't know.
[11:45] works out to be this mystery constant squared. it's just relating one thing we don't know to another thing we don't know,
[11:58] we know that it's equal to pi. the area underneath this bell curve, must be the square root of pi.
[12:10] It's a very pretty argument, but a few things are not entirely satisfying. something that just happened to work without offering much of a sense for how you Also, if we think back to our imagined statistician's friend,
[12:26] it doesn't really answer their question, which was what do circles have to do with Like I said, it's the first step, not the last, and as our next step, let's see if we can unpack why this proof is not quite as wild and arbitrary
[12:39] this function e to the negative x squared is coming from in the first place.
[12:51] John Herschel was this mathematician slash scientist slash inventor who really did all sorts of things throughout the 19th century. He made contributions in chemistry, astronomy, photography, botany,
[13:03] he invented the blueprint and named many of the moons in our solar system, and in the midst of all of this, he also offered a very elegant little derivation for the Gaussian distribution in 1850.
[13:15] kind of probability distribution in two-dimensional space. For instance, maybe you want to model the probability density for hits on a dartboard. What Herschel showed is that if you want this distribution to satisfy two pretty
[13:29] and even if you had never heard of a Gaussian in your life, you would be inexorably drawn to use a function with the shape e to the negative You do have one degree of freedom to control the spread of that distribution,
[13:44] and of course there's going to be some constant sitting in front to make sure it's normalized, but the point is that we're forced into this very specific kind of bell The first of these two properties is that the probability density around each
[13:57] point depends only on its distance from the origin, not on its direction. this would mean that you could rotate the board and it would make no difference for the distribution.
[14:12] Mathematically, this means that the function describing your probability distribution, well it can be expressed as some single variable function of the radius r.
[14:25] And just to spell it out, r is the distance between the point xy and the origin, the square root of x squared plus y squared. independent from each other, which is to say if you learn the x coordinate of a point,
[14:40] it would give you no information about the y coordinate. which describes the probability density around each point on the xy plane, can be factored into two different parts, one of which can be purely written in terms
[14:55] of x, this is the distribution of the x coordinate, I'm giving it the name g, and the other part is purely in terms of y, this would be the distribution for the But if you combine this with the assumption that things are radially symmetric,
[15:09] the behavior on each axis should look the same. So we could also write this as g of x times g of y, it's the same function. to the one we were looking at, the one that describes our probability
[15:25] To see this, imagine you were to analyze a point that was on the x-axis, Then the two distinct ways to express our function based on the two different
[15:38] properties tells us that f of r has to equal some constant multiplied by g of r. just up to some constant multiple. It would be really nice if we could just assume that that constant was one,
[15:53] And what I'm going to do, which might feel a little bit cheeky, What this means is that our answer is going to be a little bit wrong. distribution will be off by some constant factor.
[16:09] But that's no big deal, because in the end we can just rescale to make sure the area under the curve is one, like we always do with probability distributions. nice little equation purely in terms of the function f.
[16:24] If you have some point in the xy-plane, a distance r from the origin, then f of r tells you the relative likelihood of that point showing up in the More specifically, it gives the probability density of that point.
[16:39] but Herschel's two different properties evidently imply something kind of funny about it, which is that if we take the x and y coordinates of that point on the plane and evaluate this function on them separately, taking f of x times f of y,
[16:54] Or if you prefer, we could expand out the meaning of that distance r as the square root of x squared plus y squared, and this is what our key equation looks like. We're not solving for an unknown number.
[17:11] Instead, we're saying that the equation is true for all possible numbers x and y, and the thing we're trying to find is an unknown function. that satisfies this property, e to the negative x squared,
[17:26] and as a sanity check, you might verify for yourself that it does satisfy that. and to instead deduce what all of the functions are which satisfy this property.
[17:38] but let me show you how you can solve this one. which will be defined as our mystery function evaluated at the square root of x.
[17:52] Said another way, h of x squared is the same thing as f of x. the negative x squared will happen to be one of the answers, But again, we're pretending like we don't know that.
[18:08] The reason for doing this is that the key property for f looks a little bit nicer if because now what it's saying is if you take two arbitrary positive numbers and you add them up and evaluate h, it's the same thing as evaluating h on them separately
[18:22] In a sense, it turns addition into multiplication. take a moment to walk through why this forces our hand. As a next step, you might want to pause and convince yourself that
[18:36] this property also must be true if we add up an arbitrary number of inputs. think about plugging in a whole number, something like h of 5.
[18:51] this key property means that it must equal h of 1 multiplied by itself five times. I could have chosen any whole number n, and we'd be forced to conclude
[19:06] that the function looks like some number raised to the power n. And let's go ahead and give that number a name, like b for the base of our exponential. As a little mini exercise here, see if you can pause and take a moment to convince
[19:20] that if you plug in p over q to this function, it must look like this base b raised to the power p over q. that input to itself q different times.
[19:38] And then because rational numbers are dense in the real number line, continuous functions, this is enough to force your hand completely and say that h has to be an exponential function, b to the power x, for all real number inputs x.
[19:55] The way we defined h, it's only taking in positive numbers. functions as some base raised to the power x,
[20:08] mathematicians often like to write them as e to the power of some constant c times x. determine which specific exponential function you're talking about just
[20:20] makes everything much easier any time calculus comes wandering along your path. And so this means that our target function f has to look like e to the power of some constant times x squared.
[20:33] that was merely handed down to us from on high. Instead we started with these two different premises for how we wanted a distribution in two dimensions to behave, and we were drawn to the conclusion that the shape of
[20:47] the expression describing that distribution as a function of the radius away from the origin has to be e to the power of some constant times that radius squared. You'll remember I said earlier this answer will be off by a factor of a constant.
[21:00] and geometrically you might think of that as scaling it so that the volume under the Now you might notice that for positive values of this constant in the exponent c,
[21:13] so the volume under that surface would be infinite, You can't turn it into a probability distribution. constant in the exponent has to be a negative number,
[21:28] and the specific value of that number determines the spread of the distribution. who's most well known for having written down the fundamental equations
[21:40] for electricity and magnetism, independently stumbled across the same derivation. statistical mechanics and he was deriving a formula for the distribution for velocities of molecules in a gas, but the logic all works out the same.
[21:55] For you and me, if we view this as the defining property of a Gaussian, then it's a little bit less surprising that pi might make an appearance. After all, circular symmetry was part of this defining property.
[22:08] saw earlier feel a little bit less out of the blue. I mean, a key problem-solving principle in math is to use the defining features of your setup, and if you had been primed by this Herschel-Maxwell derivation,
[22:21] where the defining property for a Gaussian is this coincidence of having a distribution that's both radially symmetric and also independent along each axis, then the very first step of our proof, which seemed so strange bumping the problem
[22:35] up one dimension, was really just a way of opening the door to let that defining And if you think back, the essence of the proof came down to using that radial symmetry on the one hand, and then also using the ability to factor the function on the other.
[22:51] trick that happened to work, and more like an inevitable necessity. this is still not entirely satisfying.
[23:06] Using the Herschel-Maxwell derivation, saying this property of a multi-dimensional distribution is what defines a Gaussian, well that presumes that we're already in some kind of multi-dimensional situation in the first place.
[23:18] arises in practice doesn't feel spatial or geometric at all. about adding together many different independent variables.
[23:30] So to bring it all home here, what we need to do is explain why the function that's characterized by this Herschel-Maxwell derivation should be the same thing as the function that sits at the heart of the central limit theorem.
[23:42] And at this point, those of you following along are probably going to make fun of me, I think it makes sense to pull this last step out as its own video. After making a Patreon post about this particular project, one patron,
[23:55] who's a mathematician named Kevin Ega, shared something completely delightful that I had never seen before, which is that if you apply this integration trick in higher dimensions, it lets you derive the formulas for volumes of higher dimensional
[24:08] It's a very fun exercise, I'm leaving the details up on the screen Thank you very much to Kevin for sharing that one, and thanks to all patrons, and also for all the feedback you offer on the early drafts of videos.
[24:34] Thank you.
⚡ Saved you 0h 24m reading this? Transcribe any YouTube video for free — no signup needed.