课程导航
课程首页:Probability Theory
前置内容:Distribution Functions, Mass Functions, and Density Functions
后续内容:Transformation Techniques and Probability Integral Transform
课程资料:02 Transformations and Expectations, pp.1–17;Casella & Berger, Statistical Inference, Ch2
Transformations and Expected Values / 随机变量的变换与期望
1. Distributions of Functions of a Random Variable / 随机变量函数的分布
设随机变量 的分布已知,并定义
目标是由 的分布求出 的分布。首先要确定 在 的 support 上是否 one-to-one,再选择离散求和、CDF method 或 change of variables。
1.1 Discrete Transformations / 离散变换
Theorem
If is discrete and , then
同一个 可能对应多个 。因此, 的概率等于所有满足 的原像点概率之和。
Example
设 在集合 上均匀分布,令 。则
不是 one-to-one,所以除 外,每个 的可能取值都要合并两个原像点的概率。
1.2 CDF Method / 分布函数法
Definition
For , begin with
Rewrite the event in terms of , then differentiate when a density exists.
CDF method 直接处理事件 ,不要求 单调。它特别适用于非单调变换、最大值或最小值,以及变换后 support 随 改变的情形。
求解时需要先把不等式改写为关于 的事件:
若最后得到的 可微,则
1.3 One-to-One Change of Variables / 一一变量变换
Theorem
Suppose and is strictly monotone and differentiable. If , then
for in the transformed support.
绝对导数
是 one-dimensional Jacobian。它修正了变量变换对区间长度的伸缩,使变换前后的概率保持一致。
Example
设 ,并令
由 可知 的 support 为 ,且
因此
故
也可用 CDF method 验证:对 ,
1.4 Non-One-to-One Transformations / 非一一变换
Theorem
If has inverse branches , then
当一个 对应多个原像点时,每个 inverse branch 都贡献一项,最后把这些密度相加。
Example
令 。对 ,两个 inverse branches 为
相应的 Jacobian 绝对值均为 ,所以
1.5 Normal and Chi-Square Distributions / 正态分布与卡方分布
Theorem
If , then
令 。标准正态密度关于原点对称,因此对 ,
这正是自由度为 的 chi-square density。
进行变量变换时,应先写出 的 support,再求 的 support,并列出全部 inverse branches。所得 PMF 或 PDF 还应满足非负性和总质量为 。
2. Expected Values / 期望
2.1 Law of the Unconscious Statistician (LOTUS) / 无意识统计学家法则
Theorem
For a function for which the expectation exists,
for discrete , and
for continuous .
该结论称为 Law of the Unconscious Statistician (LOTUS)。计算 时,可以直接对 的分布求和或积分,不必先推导 的分布。
例如,若 连续且二阶矩存在,则
2.2 Linearity of Expectation / 期望的线性性质
Theorem
Whenever the expectations exist,
Independence is not required.
更一般地,
期望的线性性只要求相关期望存在,不要求随机变量之间独立。独立性会在乘积期望或方差加法中发挥作用,但不是线性性的前提。
2.3 Indicator Variables / 指示变量
Definition
For an event , let equal when occurs and otherwise. Then
示性变量把事件是否发生转化为取值为 或 的随机变量。若
表示落入集合 的观测个数,则
若 identically distributed,且 ,则
这里仍不需要独立性;只要每一项的边缘概率为 即可。
2.4 Variance / 方差
Definition
The variance of is
方差度量随机变量围绕均值的平方离散程度。由定义可得
平移常数 不改变离散程度,而缩放 会使方差乘以 。
2.5 The Mean Minimizes Squared Distance / 均值最小化平方距离
Theorem
If , then the constant minimizing
is .
Proof
令 ,则
第一项与 无关,第二项在且仅在 时为 ,所以均值是平方损失下的最优常数预测。