Random Variables and Expectation: Measurable Functions and the Lebesgue Integral
Prerequisite:Probability Spaces and Kolmogorov's Axioms: Probability as a Measure of Total Mass One
0. Key points
Section titled “0. Key points”- A random variable is a measurable function from a probability space to the real line. Measurability is demanded so that quantities such as make sense, that is, so that is an event to which a probability has been assigned.
- A random variable induces a distribution (its image measure) on the real line. The distinction between “discrete” and “continuous” is nothing more than the question of whether is concentrated on a countable set or has a density with respect to Lebesgue measure.
- The expectation is the Lebesgue integral . It is built in four stages: indicator functions, then nonnegative simple functions, then nonnegative measurable functions, then integrable functions.
- The transfer formula translates an integral over an abstract into a familiar sum or an integral over the real line. Every actual computation passes through this theorem.
- Variance and covariance are precisely the inner-product structure of . The correlation coefficient is the cosine of the angle between two centred random variables, and is nothing but the Cauchy–Schwarz inequality.
- Chebyshev’s inequality bounds tail probabilities using the variance alone, with no assumption whatever on the shape of the distribution. The weak law of large numbers is a two-line consequence of it.
1. Motivation
Section titled “1. Motivation”1.1. “A variable determined at random” is not a definition
Section titled “1.1. “A variable determined at random” is not a definition”In high-school texts and in elementary statistics, a random variable is described as “a variable whose value is determined by the outcome of a trial”. The description is intuitively correct, but as a mathematical object it defines nothing: neither “variable” nor “determined at random” has been given a meaning.
The resolution Kolmogorov gave in his 1933 Grundbegriffe der Wahrscheinlichkeitsrechnung is almost disappointingly simple. Push all the randomness into the probability measure , and make the random variable itself a deterministic function. A point of the sample space records which outcome occurred, and is the number attached to that outcome. Only the selection of is random; the assignment rule contains no randomness at all.
For instance, for the experiment of rolling two dice we may take , and the sum of the faces is the completely deterministic function . The “probability that the sum is ” is the -measure of the set .
1.2. Why measurability is required
Section titled “1.2. Why measurability is required”Once this point of view is adopted, the condition to impose on a random variable becomes visible by itself. What we wish to compute are quantities such as and . But the probability measure is not defined on all subsets of ; it is defined only on the sets (events) belonging to the -algebra . Hence unless
holds, the symbol is meaningless. Requiring this for every Borel subset of is exactly measurability, and it becomes the definition of a random variable.
This condition is by no means automatic. Let , let be the Lebesgue measurable sets, let be Lebesgue measure, and let be a Vitali non-measurable set. Then is a real-valued function on , but , so cannot be defined. Measurability is “the minimal etiquette that excludes pathological functions”. For details see Probability spaces and Kolmogorov’s axioms, and in particular the definition of a σ-algebra and a measurable space(Definition 3.1)[Probability Spaces and Kolmogorov's Axioms].
1.3. Why the expectation is defined as a Lebesgue integral
Section titled “1.3. Why the expectation is defined as a Lebesgue integral”Elementary probability defines the expectation in two separate ways: in the discrete case and in the continuous case. This carries three disadvantages.
First, random variables of neither kind arise constantly in practice. The payout of an insurance policy with a deductible, for example, is a mixture: an atom of positive probability at with a continuous distribution sitting above it. Second, every theorem has to be proved twice, once for each case. Third, and most importantly, the law of large numbers and the central limit theorem deal with limits of sequences of random variables, and the theorems guaranteeing that limits may be interchanged with integrals (monotone convergence, dominated convergence) are available only within the Lebesgue framework.
Defining the expectation as the single Lebesgue integral resolves all three at once. Below we follow that construction, starting from approximation by simple functions.
2. Preliminaries
Section titled “2. Preliminaries”2.1. Notation and standing assumptions
Section titled “2.1. Notation and standing assumptions”Throughout, is a probability space (Definition 4.1[Probability Spaces and Kolmogorov's Axioms]): is a nonempty set, is a -algebra on , and is a countably additive measure with . The -algebra generated by the open subsets of is written and called the Borel -algebra; its members are Borel sets. We set .
The indicator function of a set is written : it equals when and otherwise. Following the usual convention, sets describing events are abbreviated as .
2.2. Facts taken as known
Section titled “2.2. Facts taken as known”We use the following theorem, established in The Lebesgue integral and the convergence theorems, as known; the proof is left to that article.
Theorem 2.1(Monotone convergence theorem)
Let be a measure space and let be a sequence of measurable functions. If and for every , then is measurable and
holds (as an equality, including the case where both sides are ).
The proof is the monotone convergence theorem (Beppo Levi)(Theorem 5.2)[ルベーグ積分の定義と収束定理] in The Lebesgue integral and the convergence theorems. Wherever below we write “interchange the limit and the integral as ”, the justification is, unless stated otherwise, an application of Theorem 2.1.
2.3. Independence
Section titled “2.3. Independence”To handle additivity of the variance in Section 5, we record independence in the language of random variables.
Definition 2.3(Independence of random variables)
Random variables are independent if for all Borel sets
holds. An infinite family of random variables is independent if every finite subfamily is independent.
3. Random variables and their distributions
Section titled “3. Random variables and their distributions”3.1. Definition and criteria for measurability
Section titled “3.1. Definition and criteria for measurability”Definition 3.1(Random variable)
Let be a probability space. A map is a random variable (equivalently, an -measurable function) if
holds. If the range is enlarged to , one uses the Borel -algebra of in place of and speaks of an extended-real-valued random variable.
Verifying the definition literally would require checking for every member of , which is not feasible: the Borel sets include a vast supply of sets with no explicit description. The next proposition guarantees that it suffices to check on a generating family, and thereby turns the verification of measurability into a practical task.
Proposition 3.2(Measurability criterion via a generating family)
Let be a family of sets with . If a map satisfies
then is a random variable.
Proof(Proposition 3.2)
Put ; we show that is a -algebra on . The key is the elementary fact that preimages commute with all set operations.
(i) , so .
(ii) Let . The statement is equivalent to , that is, to ; hence . Since is closed under complements, , so .
(iii) Let . The statement is equivalent to ” for some ”, so . Since is closed under countable unions, this lies in , whence .
Therefore is a -algebra. By hypothesis , and is the smallest -algebra containing , so . This says that for every Borel set , which is the condition in Definition 3.1.
Corollary 3.3(Criterion via half-lines)
A map is a random variable if and only if
Proof(Corollary 3.3)
Necessity follows from Definition 3.1, since is a Borel set. For sufficiency, put ; we claim . Indeed, every open interval can be produced from members of by countably many set operations:
(where is a natural number large enough that ), and every open subset of is a countable union of open intervals. Hence contains all open sets, so . The reverse inclusion is clear because is closed. It now suffices to apply Proposition 3.2 to .
We must also check that operations on random variables produce random variables again. Without this we could not even write “the expectation of ”.
Proposition 3.4(Closure of random variables under operations)
Let be random variables on a probability space and let . Then , , and are all random variables. Moreover, for a sequence of random variables, , , and are extended-real-valued random variables.
Proof(Proposition 3.4)
Everything is reduced to Corollary 3.3.
Sum. For every ,
Indeed, if then , so by density of the rationals there is with , placing in the right-hand side; the reverse inclusion is clear. Since and is countable, the right-hand side lies in . Consequently .
Scalar multiple. If then ; if then ; if then . In each case the set lies in .
Product. First consider : if then , and if then . So is a random variable, and by the identity together with the closure under sums and scalar multiples already proved, is a random variable too.
Maximum. .
Supremum and infimum. is a countable intersection, hence lies in ; the infimum follows from . Finally and , so applying the supremum and infimum results twice gives the claim.
3.2. Distribution (image measure) and distribution function
Section titled “3.2. Distribution (image measure) and distribution function”The “probabilistic information” carried by a random variable does not depend on what kind of set is. It is determined solely by how much probability mass is placed where on the real line. The object that extracts this is the distribution.
Definition 3.5(Distribution and distribution function)
Let be a random variable on a probability space . The set function on defined by
is called the distribution (or image measure, or law) of . Furthermore
is called the distribution function of .
Proposition 3.6(The distribution is a probability measure, determined by its distribution function)
The set function of Definition 3.5 is a probability measure on . Moreover its distribution function has the following three properties.
- Nondecreasing: .
- Right-continuous: for every .
- Limits at the ends: and .
Proof(Proposition 3.6)
It is a measure. by nonnegativity of , and because . For countable additivity, let be pairwise disjoint; then their preimages are pairwise disjoint as well, since would give , a contradiction. Hence, using the identity from the proof of Proposition 3.2 together with countable additivity of ,
Property 1. If then , so monotonicity of measures applies.
Property 2. Let be any decreasing sequence; then . Since is a finite measure, continuity from above applies and . As is monotone, validity along every decreasing sequence gives the limit as .
Property 3. From and continuity from above, ; from and continuity from below, . Monotonicity makes this sufficient.
Conversely, given any with properties 1–3, there is exactly one probability measure having as its distribution function (the Lebesgue–Stieltjes construction). Thus “distributions of random variables” and “functions with properties 1–3” correspond one to one. The construction uses the same outer-measure machinery as in Measurable sets and Lebesgue measure.
3.3. Discrete and continuous types
Section titled “3.3. Discrete and continuous types”Definition 3.8(Discrete and continuous types)
A random variable is discrete if there is a countable set with . In this case is called the probability mass function.
is continuous (absolutely continuous) if there is a Borel measurable with
where the right-hand side is a Lebesgue integral. Such an is called a probability density function.
These two do not exhaust the possibilities. There are random variables whose distribution function is continuous but which have no density (the Cantor distribution), and there are mixtures of a discrete and a continuous part. Defining the expectation by the Lebesgue integral is stronger than the elementary dichotomy precisely because it treats all of these without distinction.
Example 3.9(Standard discrete distributions)
Bernoulli distribution (with ): , . It records whether a single trial succeeded.
Binomial distribution : for . That this is a probability mass function follows from the binomial theorem:
It is the distribution of the number of successes in independent trials with success probability .
Poisson distribution (with ): for . The total mass is by the Taylor expansion of the exponential:
It arises as the limit of the binomial distribution when , and (the law of small numbers), and it is used to model the number of rare events per unit time.
Example 3.10(Standard continuous distributions)
Exponential distribution (with ): density . The total integral is
Its distribution function is for , and it is memoryless: .
Normal distribution (with and ): density
That the total integral is follows by substituting and then using the Gaussian integral . By the central limit theorem it appears universally as the limit of sums of many small independent effects. For details see the central limit theorem (Lindeberg–Lévy)(Theorem 5.6)[The Law of Large Numbers and the Central Limit Theorem] in The law of large numbers and the central limit theorem.
4. Defining the expectation as a Lebesgue integral
Section titled “4. Defining the expectation as a Lebesgue integral”4.1. The four stages of the construction
Section titled “4.1. The four stages of the construction”The expectation is defined in the following four stages. At each stage one must check that the definition extends the previous one consistently.
flowchart LR A["indicator 1_A<br/>E = P(A)"] --> B["nonnegative simple<br/>E = Σ a_i P(A_i)"] B --> C["nonnegative measurable<br/>E = sup{integrals of simple functions}"] C --> D["integrable<br/>E = E[X⁺] − E[X⁻]"]
Definition 4.1(Expectation)
Let be a probability space.
(Stages 1 and 2.) If form a partition of and , a random variable of the form is called a nonnegative simple function, and its expectation is defined by
(Stage 3.) For measurable set
(Stage 4.) For a general random variable put and , so that and . If we call integrable and define
The set of all integrable random variables is written .
The definition in stages 1 and 2 requires a check: since the same simple function admits several representations, we must show that the expectation does not depend on the representation chosen.
Lemma 4.2(The expectation of a simple function is well defined)
If a nonnegative simple function has two representations , where and are both partitions of into members of and , then
Proof(Lemma 4.2)
Consider the common refinement . The family is a partition of with and , so finite additivity of gives
If , then for the value equals in the first representation and in the second, so . Therefore
(the terms with have and hence do no harm even when ).
One must also check that the definition at stage 3 is consistent with stage 2, that is, that the two agree when is itself a nonnegative simple function. Since itself is among the simple functions , we have ; conversely, for any simple with , passing to a common refinement shows by finite additivity. Hence .
4.2. The standard approximating sequence
Section titled “4.2. The standard approximating sequence”The definition by a supremum at stage 3 is not usable for computation as it stands. What one actually uses is the following construction, which approximates an arbitrary nonnegative measurable function from below by an increasing sequence of simple functions. (The same statement for a general measure space is the simple approximation theorem(Theorem 3.7)[ルベーグ積分の定義と収束定理].)
Proposition 4.4(Standard approximating sequence)
Let be measurable and for set
Then each is a nonnegative simple function, and for every we have and . Moreover .
Proof(Proposition 4.4)
They are simple functions. Since is measurable, each of the sets and lies in by Definition 3.1. They form a finite partition of , and the coefficients are nonnegative.
Monotonicity. Passing from stage to stage halves the width of the intervals from to . If , then the value at stage is either or , so . On the part where , the value at stage is at least .
Pointwise convergence. If , then for all beyond some we have . If , then .
Convergence of the integrals. The sequence consists of nonnegative measurable functions increasing pointwise to , so Theorem 2.1 gives .
4.3. The transfer formula
Section titled “4.3. The transfer formula”The expectation was defined as an integral over , but is an abstract set and nothing can be computed there directly. The theorem that makes every actual computation possible is the following. In statistics texts it is sometimes called the law of the unconscious statistician.
Theorem 4.5(Transfer formula)
Let be a random variable on a probability space and let be Borel measurable.
- If , then as an identity in ,
- For general , integrability of with respect to is equivalent to integrability of with respect to , and in that case the identity above holds.
Proof(Theorem 4.5)
We argue by the standard machine (indicator functions, then simple functions, then nonnegative functions, then the general case). Note first that is a random variable: for any Borel set we have , and measurability of gives while measurability of gives .
Stage 1 (). Let . We have exactly when , that is, when ; hence . So stage 1 of Definition 4.1 gives
The third equality is Definition 3.5, and the fourth is stage 1 of the definition of the integral with respect to .
Stage 2 (nonnegative simple functions). Let with and the Borel sets partitioning . Then , and is a partition of . Hence
Here the outer equalities are independent of the chosen representation by Lemma 4.2.
Stage 3 (). Apply the standard approximating sequence of Proposition 4.4 to , obtaining an increasing sequence of nonnegative simple functions . Composing, for every , so is an increasing sequence of nonnegative simple functions as well. Applying Theorem 2.1 on both sides (with measure on the left and on the right),
where the middle equality is the result of stage 2.
Stage 4 (general case). Decompose . Since and , applying stage 3 to gives
so one side is finite exactly when the other is; this is the asserted equivalence of integrability. When is integrable, apply stage 3 separately to and and subtract; stage 4 of Definition 4.1 then yields the stated identity.
Corollary 4.6(Computational formulas in the discrete and continuous cases)
Let be Borel measurable.
- If is discrete with for and , then whenever the variable is integrable and
- If is continuous with density , then whenever the variable is integrable and
Proof(Corollary 4.6)
By Theorem 4.5, both reduce to computing .
1. Put , so that and . Suppose first that and set ; then is simple and . By the definition at stage 2, , and letting with Theorem 2.1 gives . Since , we have . For general split into , run the same argument, and subtract under the hypothesis of absolute convergence.
2. The claim has the form “if then ”, and this too is proved by the standard machine: for it is exactly the definition of a density in Definition 3.8; simple functions follow by linearity; nonnegative measurable functions follow from Theorem 2.1 (if then ); and the general case follows by splitting into positive and negative parts.
4.4. Worked computations
Section titled “4.4. Worked computations”Example 4.7(Means and second moments of standard distributions)
Poisson distribution : by part 1 of Corollary 4.6,
Here we used that the term vanishes and that . Similarly, for ,
Hence .
Exponential distribution : by part 2 of Corollary 4.6, integrating by parts,
In the last equality we used the value just computed.
Normal distribution : substituting ,
(the function is odd and absolutely integrable, so its integral is ). Next, , and integrating by parts with and (so ),
so .
Example 4.8(A random variable with no expectation (the Cauchy distribution))
Consider a random variable with density . This really is a density, since .
The density is symmetric about the origin, so one is tempted to say that the mean “ought to be ”; but the integrability condition at stage 4 of Definition 4.1 fails. Indeed
Thus , and cannot be defined. The correct statement is that does not exist.
This is not a pathological example. The Cauchy distribution arises naturally as the distribution of a ratio of independent standard normal variables, and moreover the sample mean of independent Cauchy samples follows the same Cauchy distribution for every . In other words, increasing the sample size does not improve the precision. Here one sees what it means that the law of large numbers assumes the existence of the expectation.
5. Variance, covariance and correlation
Section titled “5. Variance, covariance and correlation”5.1. Definitions and basic properties
Section titled “5.1. Definitions and basic properties”The expectation is a single number describing the “location” of a distribution. The standard quantity measuring its “spread” is the variance.
Definition 5.1(Variance, covariance and correlation coefficient)
Write for the set of random variables with . For , put and and define
called the variance of and the covariance of and . The quantity is the standard deviation. If moreover and , then
is the correlation coefficient. When we say and are uncorrelated.
For this definition to make sense, and must be integrable whenever .
Proposition 5.2(On a probability space, L² ⊂ L¹)
On a probability space we have , and if then .
Proof(Proposition 5.2)
Expanding for real gives . Taking yields , so
(we used ; that this is finite is exactly what a probability measure buys us). Taking and in the same inequality gives , whose right-hand side is integrable, so . Consequently is integrable as well.
Proposition 5.3(Basic properties of variance and covariance)
Let and .
- and .
- . In particular the variance is invariant under translation by a constant.
- is a symmetric bilinear form, and .
- .
- holds if and only if .
Proof(Proposition 5.3)
1. By linearity of the expectation (a basic property following from Definition 4.1),
Setting gives the formula for the variance. By Proposition 5.2 every term is finite, so the expansion is legitimate.
2. Since , we get , whence .
3. Symmetry holds because the defining expression is symmetric in and . Bilinearity follows by expanding using linearity of the expectation.
4. Put . Then , so by the bilinearity of part 3,
The diagonal terms give , and by symmetry the off-diagonal terms amount to twice the terms with .
5. Put . If , then for every the simple function gives , that is, . Since is a countable union, , i.e. . Conversely, if then with probability and hence .
Example 5.4(Computing variances, and additivity)
From the results of Example 4.7 and part 1 of Proposition 5.3 we immediately obtain the following.
- : . That the mean and the variance coincide is characteristic of the Poisson distribution.
- : . The standard deviation equals the mean .
- : . The parameter is the variance itself.
For the binomial distribution it is quicker to decompose into an independent sum than to compute directly. Let be independent with distribution , so that . Since ,
By independence and Proposition 7.5 (Appendix) we have for , so part 4 of Proposition 5.3 gives
No sum involving binomial coefficients has to be evaluated directly.
5.2. Uncorrelated is not the same as independent
Section titled “5.2. Uncorrelated is not the same as independent”5.3. The Cauchy–Schwarz inequality and the meaning of the correlation coefficient
Section titled “5.3. The Cauchy–Schwarz inequality and the meaning of the correlation coefficient”Theorem 5.6(Cauchy–Schwarz inequality for the covariance)
Let . Then
Moreover, when and , equality holds if and only if there exist real numbers and with .
Proof(Theorem 5.6)
Put and , so that , , and . All of these are finite by Proposition 5.2.
Case : by part 5 of Proposition 5.3 we have , so with probability and ; the inequality holds with equality.
Now suppose . For every we have , so by monotonicity of the expectation
(the expansion uses linearity of the expectation). This is a quadratic in with leading coefficient . Nonnegativity for all is equivalent to the discriminant being at most , so
that is, .
Equality. Equality holds exactly when the discriminant vanishes, i.e. when there is a (double) root with . From and the argument inside the proof of part 5 of Proposition 5.3 (a nonnegative random variable with expectation vanishes with probability ), we get . Rewriting,
so it suffices to set and . We do need , i.e. : if then , contradicting the assumption , so indeed . (In the degenerate case the variable is constant, and although equality does hold, , one has .) Conversely, if with , then parts 2 and 3 of Proposition 5.3 give and , so and equality holds.
Corollary 5.7(Range of the correlation coefficient)
If and then . Moreover holds if and only if there exist with , the sign of agreeing with the sign of .
Proof(Corollary 5.7)
Dividing both sides in Theorem 5.6 by gives , and the equality condition follows from the same theorem. As for the sign, when we have .
The statement that the correlation coefficient measures “the strength of a linear relationship” acquires a precise meaning through the following computation.
Example 5.8(The correlation coefficient determines the error of the best linear prediction)
Suppose we predict from by an expression of the form ; we find the minimising the mean squared error, and the minimum value. Assume and .
First fix and minimise over . Put . With we have (apply part 1 of Proposition 5.3 to ). So is optimal, and the resulting value is
(by parts 4, 2 and 3 of Proposition 5.3). This is a quadratic in ; completing the square gives
so the minimum is attained at
Thus is “the fraction of the variance of explained by linear prediction from ”. If the error is , that is, is an affine function of , exactly as in Corollary 5.7. If , linear prediction cannot beat the constant . The coefficient of determination in regression analysis is a generalisation of this . Handling more complicated dependences requires dropping the linearity constraint, which leads to the conditional expectation (see Conditional expectation, and in particular the conditional expectation as an orthogonal projection(Theorem 5.1)[条件付き期待値]).
6. The Markov and Chebyshev inequalities
Section titled “6. The Markov and Chebyshev inequalities”Even with no knowledge whatever of the shape of a distribution, tail probabilities can be estimated from the mean and the variance alone. This is the content of the next two inequalities. Both proofs take two lines, yet they are among the most heavily used tools in probability.
Theorem 6.1(Markov's inequality)
Let be a nonnegative random variable, i.e. . Then for every ,
Proof(Theorem 6.1)
Put (measurability of ). For every ,
Indeed, if the left-hand side is and by definition of we have ; if the left-hand side is and the right-hand side is nonnegative since .
The expectation is monotone ( implies , which follows at once from the supremum in stage 3 of Definition 4.1), so taking expectations of both sides,
Dividing by gives the claim.
Corollary 6.2(Chebyshev's inequality)
Let , and . Then for every ,
In particular, when , taking with gives
Proof(Corollary 6.2)
Put ; by Proposition 5.2 we have . The events satisfy
(since is strictly increasing on , the conditions and are equivalent). Applying Theorem 6.1 to and ,
Taking gives .
Example 6.3(How loose is the Chebyshev bound?)
For , Corollary 6.2 gives . Let us compare with the true values for concrete distributions.
- : .
- (so ): .
Both are about one fifth of the bound . If the distribution is known, far sharper estimates are available.
Chebyshev’s inequality is nonetheless important because it assumes nothing about the distribution. It applies where no distributional form can be posited: unknown populations, the outputs of complicated machine-learning models. Moreover the bound cannot be improved if one demands validity across all distributions, because there are distributions attaining equality (Exercise 7.3).
Example 6.4(Application to the weak law of large numbers)
Let be independent and identically distributed with and , and put . Linearity of the expectation gives . Using parts 2 and 4 of Proposition 5.3 together with the vanishing of the covariance terms under independence (Proposition 7.5),
Applying Corollary 6.2 to , for every ,
This is the weak law of large numbers (the law of large numbers in the sense of convergence in probability). Versions dispensing with the variance assumption, and the strong law asserting almost sure convergence (the strong law of large numbers (Kolmogorov)(Theorem 4.5)[The Law of Large Numbers and the Central Limit Theorem]), are treated in The law of large numbers and the central limit theorem.
This estimate can be used directly to design a sample size. With and , to keep the error probability at most it suffices to take . The central limit theorem shows that is already enough, but that is an approximation valid for large , whereas the Chebyshev estimate is an inequality that is exactly correct for every .
7. Exercises
Section titled “7. Exercises”Exercise 7.1Standard
Suppose satisfies for every rational . Show that is a random variable.
Solution
Apply Proposition 3.2 with . We must show .
First, because each is open.
For the converse, let be any real number. By density of the rationals there is a rational sequence , and
(if then for all large ). Hence .
Next, also lies in , since
(if then for every ; conversely if for every then letting gives ), and each set on the right lies in by the previous step. Therefore, for any open interval,
Every open subset of is a countable union of open intervals (take intervals with rational endpoints around each point), so all open sets lie in , and hence .
This proves , and Proposition 3.2 shows that is a random variable. The point is that countably many conditions already imply measurability.
Exercise 7.2Standard
Let and , and let have the geometric distribution, that is,
Compute and .
Solution
First we check that this is a probability mass function: , the geometric series converging because .
For , differentiating termwise gives , and differentiating once more gives (a power series may be differentiated termwise inside its radius of convergence).
By part 1 of Corollary 4.6,
Next, for ,
Hence , and part 1 of Proposition 5.3 gives
where we used . For example, with (the number of rolls until a die first shows a ) we get and .
Exercise 7.3Standard
Fix and . Construct a random variable with , and
thereby showing that the bound in Corollary 6.2 cannot be improved.
Solution
Take the following three-point distribution.
Since we have , and the probabilities sum to .
By part 1 of Corollary 4.6,
Hence part 1 of Proposition 5.3 gives .
On the other hand holds exactly when , so
and Chebyshev’s inequality holds with equality. Therefore, among bounds of the form valid for all random variables, one cannot take smaller than . Improvements are possible under extra assumptions on the distribution (symmetry, unimodality, boundedness, and so on).
Exercise 7.4Hard
Let be a random variable with . Show that
holds as an identity in , the right-hand side being a Lebesgue integral.
Solution
On the product space consider the function
Measurability. The map is the composition of the projection onto the first coordinate with , hence measurable with respect to the product -algebra; the same holds for . By the result on sums (differences) in Proposition 3.4, the map is measurable, and
is the preimage of under that measurable function, hence a measurable set. Thus is nonnegative and measurable.
Applying Tonelli’s theorem. Since is a probability measure and Lebesgue measure on is -finite, Tonelli’s theorem for nonnegative measurable functions (The Lebesgue integral and the convergence theorems) says that the two iterated integrals agree.
Integrating in first, with fixed,
(since , the integrand in is the indicator of the interval ). So this iterated integral equals .
Integrating in first, with fixed,
so this iterated integral is .
By Tonelli’s theorem the two are equal in , which proves the claim.
Check. For we have , so , agreeing with the result of Example 4.7. The formula is especially useful when the density is unknown and only the survival probability is available, as in reliability engineering and survival analysis.
References
Section titled “References”- A. N. Kolmogorov, Grundbegriffe der Wahrscheinlichkeitsrechnung, Springer, 1933 (English translation: Foundations of the Theory of Probability, Chelsea, 1956) — the original source defining random variables as measurable functions. Chapter III.
- Kiyosi Ito, Kakuritsuron (Probability Theory), Iwanami Shoten (Iwanami Foundations of Mathematics), 1991 (in Japanese) — Chapters 1 and 2. A standard Japanese reference on measure-theoretic probability.
- Naohisa Funaki, Kakuritsuron (Probability Theory), Asakura Shoten (Ways of Mathematical Thinking 20), 2004 (in Japanese) — Chapter 2 organises random variables, expectation and the various inequalities.
- P. Billingsley, Probability and Measure, 3rd ed., Wiley, 1995 — Sections 13, 15, 21. A careful treatment of the construction of the expectation and of the transfer formula.
- R. Durrett, Probability: Theory and Examples, 5th ed., Cambridge University Press, 2019 — Chapter 1. Concise, with a wealth of exercises.
- D. Williams, Probability with Martingales, Cambridge University Press, 1991 — Chapters 3–6. Makes the “standard machine” (from indicators to the general case) explicit as a method.
Appendix: independence and products of expectations
Section titled “Appendix: independence and products of expectations”Here we prove the fact used in Example 5.4 and Example 6.4 in the main text, namely that independence implies uncorrelatedness.
Proposition 7.5(Expectation of a product of independent integrable random variables)
Let be independent random variables, both integrable. Then is integrable and
In particular, if are independent then .
Proof(Proposition 7.5)
Write for the joint distribution of on . Definition 2.3 says that
for all Borel sets . The rectangles form a multiplicative family (closed under finite intersections) generating , and both sides are probability measures on , so Dynkin’s - theorem gives (the product measure) everywhere.
Suppose first that . Apply the two-dimensional version of Theorem 4.5 (with , regarding as an -valued measurable map; the proof is the same standard machine as in the main text), and then Tonelli’s theorem for the product measure :
and by Theorem 4.5 again the right-hand side is .
In the general case decompose and . Since is a Borel measurable function of and is a Borel measurable function of , and Borel measurable functions of independent random variables are again independent (because the events have the form ), the nonnegative case applies to all four products . Integrability follows from , and the identity from
Finally, part 1 of Proposition 5.3 gives .
The converse fails: Example 5.5 is a counterexample.
The space is a Hilbert space under the inner product ; the variance is the squared norm of the centred variable, the covariance is the corresponding inner product, and the correlation coefficient is the cosine of the angle between them. This is why Theorem 5.6 is nothing but the Cauchy–Schwarz inequality in a Hilbert space. The viewpoint becomes essential when the conditional expectation is understood as an “orthogonal projection onto a subspace”. For as a Hilbert space see L² is a Hilbert space(Theorem 6.2)[L^p 空間と関数解析への導入] in L^p spaces and an introduction to functional analysis.
Report an error in this article ・Operated by: Mugen Giken LLC ・Pricing ・Terms ・Legal notice
© 2026 夢現技研合同会社 ・Feeding the text to an LLM is welcome. Code samples are MIT licensed.