
確率論と統計学において、分散はばらつきの尺度であり、つまり、一連の数値が平均値からどれだけばらついているかを示す尺度です。これは、確率変数の平均からの偏差の二乗の期待値として定義されます。標準偏差は分散の平方根です。厳密には、これは分布の2次中心モーメントであり、確率変数とそれ自身との共分散であり、多くの場合、で表されます。、 、 、 、または . [ 1 ]
分散をばらつきの尺度として用いる利点は、期待絶対偏差などの他のばらつきの尺度よりも代数的な操作が容易であることです。例えば、無相関の確率変数の和の分散は、それぞれの分散の和に等しくなります。実用的な応用において分散が不利な点は、標準偏差とは異なり、その単位が確率変数の単位と異なることです。そのため、計算が終わった後は、ばらつきの尺度として標準偏差が報告されるのが一般的です。また、多くの分布において分散が有限ではないことも不利な点です。
「分散」と呼ばれる概念には、大きく分けて2種類あります。1つは、前述のように、理論的な確率分布の一部であり、方程式で定義されます。もう1つは、観測値の集合の特性です。観測値から分散を計算する場合、それらの観測値は通常、実際のシステムから測定されます。システムのすべての観測値が存在する場合、計算された分散は母集団分散と呼ばれます。しかし、通常は一部の観測値しか得られず、そこから計算された分散は標本分散と呼ばれます。標本から計算された分散は、母集団全体の分散の推定値とみなされます。標本分散に基づいて母集団分散を推定する方法は複数あり、これについては後述します。
2種類の分散は密接に関連しています。その関係性を理解するために、理論的な確率分布を仮想的な観測値の生成器として使用できることを考えてみましょう。ある分布を用いて無限個の観測値を生成すると、その無限集合から計算された標本分散は、その分布の分散の式を用いて計算された値と一致します。分散は統計学において中心的な役割を果たしており、記述統計、統計的推論、仮説検定、適合度、モンテカルロサンプリングなど、分散を用いる概念がいくつかあります。

確率変数の分散は、平均からの二乗偏差の期待値です。、 : この定義は、離散的、連続的、どちらでもない、または混合的な プロセスによって生成される確率変数を包含します。分散は、確率変数とそれ自身との 共分散と考えることもできます。
分散は、確率分布の2番目のキュムラントにも相当し、。 分散は通常次のように指定されます。、または時々または、または象徴的にまたは単に(「シグマ二乗」と発音します)。分散の式は次のように展開できます。
言い換えれば、 の分散ははの二乗の平均に等しい平均値の二乗を引いた値この式は、浮動小数点演算を用いた計算には使用しないでください。式の2つの要素の大きさが似ている場合、致命的な打ち消しが発生するためです。数値的に安定した他の代替方法については、「分散を計算するためのアルゴリズム」を参照してください。
乱数発生器確率質量関数を持つ離散的である、 それから どこは期待値です。つまり、 (このような離散的な加重分散が、合計が 1でない重みによって指定されている場合は、重みの合計で割る。)
コレクションの分散等確率の値は次のように表すことができます。 どこは平均値です。つまり、
一連の分散等確率の値は、平均を直接参照することなく、点同士のすべてのペアワイズ二乗距離の二乗偏差で同等に表現できます。[ 2 ]
ランダム変数確率密度関数 を持つ、そしては対応する累積分布関数であり、 または同等に、 どこは期待値ですによって与えられた
これらの式では、そしてそれぞれ、ルベーグ積分とルベーグ・スティルチェス積分である。
パラメータを持つ指数分布は連続分布であり、その確率密度関数は次のように与えられる。 区間[ 0, ∞)上。その平均は次のように表される。
部分積分法を用い、既に計算済みの期待値を利用すると、次のようになります。
したがって、の分散はは次のように与えられる
公平な6面サイコロは離散確率変数としてモデル化できる。、結果1から6までがあり、それぞれ確率は1/6です。 の期待値ははしたがって、の分散はは
結果の分散の一般式は、、の面ダイスは
以下の表は、一般的に用いられるいくつかの確率分布の分散を示しています。
分散は非負である。なぜなら、二乗は正またはゼロだからである。
定数の分散はゼロである。
逆に、確率変数の分散が0であれば、それはほぼ確実に定数である。つまり、常に同じ値をとる。
コーシー分布のように、分布の期待値が有限でない場合、分散も有限にはなり得ません。ただし、期待値が有限であっても、分散が有限でない分布もあります。例えば、指数が 1/2であるパレート分布などが挙げられます。満たす
分散分解の一般式、または全分散の法則は次のとおりです。そしては 2 つの確率変数であり、分散は存在するならば、
条件付き期待値の与えられた、条件付き分散次のように理解できます。確率変数Yの任意の特定の値yが与えられた場合、条件付き期待値があります。 事象Y = yが与えられた場合、この量は特定の値yに依存します。これは関数です。 同じ関数を確率変数Yで評価したものが条件付き期待値です。 .
特に、は離散確率変数であり、取りうる値をとる。対応する確率とともにすると、全分散の式において、右辺の最初の項は次のようになる。 どこ同様に、右辺の第2項は次のようになる。 どこでそしてしたがって、総分散は次のように表されます 。
同様の式は分散分析にも適用され、対応する式は次のようになる。 ここは二乗平均を表します。線形回帰分析における対応する式は次のとおりです。
これは分散の加法性からも導き出せる。なぜなら、合計(観測)スコアは予測スコアと誤差スコアの合計であり、後者2つは無相関だからである。
二乗偏差の合計 (二乗、):
非負の確率変数の母分散は、累積分布関数Fを用いて 表現できます。
この式は、累積分布関数は簡単に表現できるが、密度関数は簡単に表現できないような状況で、分散を計算するために使用できます。
確率変数の2次モーメントは、確率変数の1次モーメント(つまり平均)付近で最小値をとる。. Conversely, if a continuous function satisfies for all random variables X, then it is necessarily of the form , where a > 0. This also holds in the multidimensional case.[3]
Unlike the expected absolute deviation, the variance of a variable has units that are the square of the units of the variable itself. For example, a variable measured in meters will have a variance measured in meters squared. For this reason, describing data sets via their standard deviation or root mean square deviation is often preferred over using the variance. In the dice example the standard deviation is √2.9 ≈ 1.7, slightly larger than the expected absolute deviation of 1.5.
The standard deviation and the expected absolute deviation can both be used as an indicator of the "spread" of a distribution. The standard deviation is more amenable to algebraic manipulation than the expected absolute deviation, and, together with variance and its generalization covariance, is used frequently in theoretical statistics; however the expected absolute deviation tends to be more robust as it is less sensitive to outliers arising from measurement anomalies or an unduly heavy-tailed distribution.
Variance is invariant with respect to changes in a location parameter. That is, if a constant is added to all values of the variable, the variance is unchanged:
If all values are scaled by a constant, the variance is scaled by the square of that constant:
The variance of a sum of two random variables is given by where is the covariance.
In general, for the sum of random variables , the variance becomes: see also general Bienaymé's identity.
These results lead to the variance of a linear combination as:
If the random variables are such that then they are said to be uncorrelated. It follows immediately from the expression given earlier that if the random variables are uncorrelated, then the variance of their sum is equal to the sum of their variances, or, expressed symbolically:
Since independent random variables are always uncorrelated (see Covariance § Uncorrelatedness and independence), the equation above holds in particular when the random variables are independent. Thus, independence is sufficient but not necessary for the variance of the sum to equal the sum of the variances.
Define as a column vector of random variables , and as a column vector of scalars . Therefore, is a linear combination of these random variables, where denotes the transpose of . Also let be the covariance matrix of . The variance of is then given by:[4]
This implies that the variance of the mean can be written as (with a column vector of ones)
One reason for the use of the variance in preference to other measures of dispersion is that the variance of the sum (or the difference) of uncorrelated random variables is the sum of their variances:
This statement is called the Bienaymé formula[5] and was discovered in 1853.[6][7] It is often made with the stronger condition that the variables are independent, but being uncorrelated suffices. So if all the variables have the same variance σ2, then, since division by n is a linear transformation, this formula immediately implies that the variance of their mean is
That is, the variance of the mean decreases when n increases. This formula for the variance of the mean is used in the definition of the standard error of the sample mean, which is used in the central limit theorem.
To prove the initial statement, it suffices to show that
The general result then follows by induction. Starting with the definition,
Using the linearity of the expectation operator and the assumption of independence (or uncorrelatedness) of X and Y, this further simplifies as follows:
In general, the variance of the sum of n variables is the sum of their covariances: (注:2番目の等式は、Cov( X i , X i ) = Var( X i ) という事実から導かれる。)
ここ、は共分散であり、独立な確率変数の場合(存在する場合)はゼロになります。この式は、和の分散が、構成要素の共分散行列のすべての要素の合計に等しいことを示しています。次の式は、和の分散が共分散行列の対角要素の合計と、その上三角要素(または下三角要素)の合計の 2 倍の合計に等しいことを同等に示しています。これは、共分散行列が対称であることを強調しています。この式は、古典的テスト理論におけるクロンバックのアルファの理論で使用されています。
したがって、変数の分散が等しいσ²であり、異なる変数の平均相関がρである場合、それらの平均の分散は
これは、平均の分散が相関の平均とともに増加することを意味します。言い換えれば、相関のある観測値を追加しても、独立した観測値を追加しても、平均の不確実性を減らす効果はそれほど高くありません。さらに、変数の分散が1である場合(例えば、標準化されている場合)、これは次のように単純化されます。
この式は、古典的テスト理論のスピアマン・ブラウン予測式で使用されます。平均相関が一定であるか収束する場合、nが無限大に近づくとρに収束します。したがって、相関が等しいか平均相関が収束する標準化変数の平均の分散については、次のようになります。
したがって、多数の標準化変数の平均の分散は、それらの平均相関にほぼ等しくなります。このことから、相関のある変数の標本平均は、独立変数の標本平均が母平均に収束するという大数の法則があるにもかかわらず、一般的には母平均に収束しないことが明らかになります。
事前に、ある基準に従って許容される観測値がいくつになるか分からないままサンプルが採取される場合がある。このような場合、サンプルサイズNは、その変動がXの変動に加算される確率変数であり、[ 8 ]これは全分散の法則 から導かれる。
Nがポアソン分布に従う場合、推定量n = Nの場合、推定量はになる, giving (see Standard error § Standard error of the sample mean).
The scaling property and the Bienaymé formula, along with the property of the covarianceCov(aX, bY) = ab Cov(X, Y) jointly imply that
This implies that in a weighted sum of variables, the variable with the largest weight will have a disproportionally large weight in the variance of the total. For example, if X and Y are uncorrelated and the weight of X is two times the weight of Y, then the weight of the variance of X will be four times the weight of the variance of Y.
The expression above can be extended to a weighted sum of multiple variables:
If two variables X and Y are independent, the variance of their product is given by[9]
Equivalently, using the basic properties of expectation, it is given by
In general, if two variables are statistically dependent, then the variance of their product is given by:
The delta method uses second-order Taylor expansions to approximate the variance of a function of one or more random variables (see Taylor expansions for the moments of functions of random variables). For example, the approximate variance of a function of one variable is given by provided that f is twice differentiable and that the mean and variance of X are finite.
Real-world observations such as the measurements of yesterday's rain throughout the day typically cannot be complete sets of all possible observations that could be made. As such, the variance calculated from the finite set will in general not match the variance that would have been calculated from the full population of possible observations. This means that one estimates the mean and variance from a limited set of observations by using an estimator equation. The estimator is a function of the sample of nobservations drawn without observational bias from the whole population of potential observations. In this example, the sample would be the set of actual measurements of yesterday's rainfall from available rain gauges within the geography of interest.
The simplest estimators for population mean and population variance are simply the mean and variance of the sample, the sample mean and (uncorrected) sample variance – these are consistent estimators (they converge to the value of the whole population as the number of samples increases) but can be improved. Most simply, the sample variance is computed as the sum of squared deviations about the (sample) mean, divided by n as the number of samples. However, using values other than n improves the estimator in various ways. Four common values for the denominator are n, n − 1, n + 1, and n − 1.5: n is the simplest (the variance of the sample), n − 1 eliminates bias,[10]n + 1 minimizes mean squared error for the normal distribution,[11] and n − 1.5 mostly eliminates bias in unbiased estimation of standard deviation for the normal distribution.[12]
Firstly, if the true population mean is unknown, then the sample variance (which uses the sample mean in place of the true mean) is a biased estimator: it underestimates the variance by a factor of (n − 1) / n; correcting this factor, resulting in the sum of squared deviations about the sample mean divided by n − 1 instead of n, is called Bessel's correction.[10] The resulting estimator is unbiased and is called the (corrected) sample variance or unbiased sample variance. If the mean is determined in some other way than from the same samples used to estimate the variance, then this bias does not arise, and the variance can safely be estimated as that of the samples about the (independently known) mean.
第二に、標本分散は一般的に標本分散と母集団分散の間の平均二乗誤差を最小化しません。バイアスを補正すると、多くの場合、状況が悪化します。補正された標本分散よりも優れた性能を発揮するスケールファクターを常に選択できますが、最適なスケールファクターは母集団の過剰尖度に依存し(平均二乗誤差§ 分散を参照)、バイアスを導入します。これは常に不偏推定量を縮小(n − 1より大きい数で割る)することによって行われ、縮小推定量の簡単な例です。不偏推定量をゼロに向かって「縮小」します。正規分布の場合、n + 1 ( n − 1またはnの代わりに)で割ると、平均二乗誤差が最小化されます。[ 11 ]ただし、結果として得られる推定量はバイアスがあり、バイアスのある標本変動として知られています。
一般に、有限サイズの母集団の母分散は値x iは次のように与えられる 母平均はそして、そこでは期待値演算子です。
母集団分散は[ 13 ]を用いて計算することもできる。
(右辺には合計に重複する項があるが、中央の辺には合計する項が重複しない。)これは、
母集団の分散は、生成確率分布の分散と一致する。この意味で、母集団の概念は、無限母集団を持つ連続確率変数にも拡張できる。
多くの実際的な状況では、母集団の真の分散は事前にわかっておらず、何らかの方法で計算する必要があります。非常に大きな母集団を扱う場合、母集団内のすべてのオブジェクトを数えることは不可能なので、母集団のサンプルに対して計算を実行する必要があります。 [ 14 ]これは一般にサンプル分散または経験的分散と呼ばれます。サンプル分散は、連続分布のサンプルからその分布の分散を推定するためにも適用できます。
復元抽出でサンプルを採取します。サイズ の母集団からの値Y 1 、 ...、Y n 、ここでn < N、このサンプルに基づいて分散を推定します。 [ 15 ]サンプルデータの分散を直接取ると、二乗偏差の平均が得られます。 [ 16 ] (この式の導出については、§ 母集団分散の項を参照してください。)ここで、標本平均を表します。
Y iはランダムに選択されるため、そしては確率変数です。期待値は、サイズ のすべての可能なサンプル{ Y i }のアンサンブルで平均することによって評価できます。人口から。これにより、以下が得られます。
こここのセクションで導出されるのは母集団分散であり、独立性によりそして .
したがって母集団の分散の推定値を示すそれは、期待値が母集団分散(真の分散)よりもその係数だけ小さい。このため、これは偏りのある標本分散と呼ばれます。
このバイアスを補正すると、不偏標本分散が得られ、それは と表記される。 :
どちらの推定量も、文脈からどちらのバージョンであるかが判断できる場合は、単に標本分散と呼ぶことができる。同様の証明は、連続確率分布から抽出された標本にも適用できる。
n − 1という項の使用はベッセル補正と呼ばれ、標本共分散や標本標準偏差(分散の平方根)にも用いられます。平方根は凹関数であるため、分布に依存する負のバイアス(イェンセンの不等式による)が生じ、ベッセル補正を用いた補正済み標本標準偏差はバイアスを持ちます。標準偏差の不偏推定は技術的に複雑な問題ですが、正規分布の場合、n − 1.5という項を用いることでほぼ不偏推定値が得られます。
The unbiased sample variance is a U-statistic for the function f(y1, y2) = (y1 − y2)2/2, meaning that it is obtained by averaging a 2-sample statistic over 2-element subsets of the population.
For a set of numbers {10, 15, 30, 45, 57, 52, 63, 72, 81, 93, 102, 105}, if this set is the whole data population for some measurement, then variance is the population variance 932.743 as the sum of the squared deviations about the mean of this set, divided by 12 as the number of the set members. If the set is a sample from the whole population, then the unbiased sample variance can be calculated as 1017.538 that is the sum of the squared deviations about the mean of the sample, divided by 11 instead of 12. A function VAR.S in Microsoft Excel gives the unbiased sample variance while VAR.P is for population variance.
Being a function of random variables, the sample variance is itself a random variable, and it is natural to study its distribution. In the case that Yi are independent observations from a normal distribution, Cochran's theorem shows that the unbiased sample varianceS2 follows a scaled chi-squared distribution (see also: asymptotic properties and an elementary proof):[17] where σ2 is the population variance. As a direct consequence, it follows that and[18]
If Yi are independent and identically distributed, but not necessarily normally distributed, then[19] where κ is the kurtosis of the distribution and μ4 is the fourth central moment.
二乗観測値に対して大数の法則の条件が満たされる場合、 S 2はσ 2の一致推定量となります。実際、推定量の分散は漸近的にゼロに近づくことがわかります。漸近的に等価な式は、Kenney と Keeping (1951:164)、Rose と Smith (2002:264)、および Weisstein (nd) で与えられています。[ 20 ] [ 21 ] [ 22 ]
サムエルソンの不等式は、標本平均と(偏りのある)分散が計算されている場合、標本内の個々の観測値が取り得る値の上限を示す結果である。[ 23 ]値は、以下の範囲内に収まらなければならない。 .
1つの新しい観測セットに追加されます平均値を持つ観測値および分散新しい差異再帰的な更新式を用いて表現することができる。Chan ら (1983) [ 24 ]によって提供された二乗和の恒等式に基づくと、次のようになる。
この関係から、新しい観測値が分散に与える影響は、現在の平均からの距離によって決まります。、その場合、分散は変化しません。したがって、新しい観測値が平均に近い場合() の場合、分散は減少します。また、平均から遠ざかるほど、) すると分散が増加します。
標本分散の導出
平方和の更新式を使用して():
標本分散の関係を代入すると():
設定:
解決する収量:
母集団分散の導出
母集団の分散については()更新式は次のとおりです。
設定:
解決する収量:
正の実数の サンプル{ y i }に対して、 ここで、y maxはサンプルの最大値であり、は算術平均です。はサンプルの調和平均であり、これは標本の(偏りのある)分散です。
この境界は改善され、分散は以下によって制限されることが知られています。 ここで、y minはサンプルの最小値である。[ 26 ]
標本が正規分布に従う場合、分散の等質性を検定するF検定とカイ二乗検定は適切である。正規分布に従わない場合、2つ以上の分散の等質性を検定することはより困難になる。
ノンパラメトリック検定がいくつか提案されている。これには、バートン・デイビッド・アンサリ・フロイント・シーゲル・テューキー検定、カポン検定、ムード検定、クロッツ検定、スカトメ検定などがある。スカトメ検定は2つの分散に適用され、両方のメジアンが既知でゼロに等しい必要がある。ムード検定、クロッツ検定、カポン検定、バートン・デイビッド・アンサリ・フロイント・シーゲル・テューキー検定も2つの分散に適用され、メジアンが未知であってもよいが、2つのメジアンが等しい必要がある。
レーマン検定は、2つの分散を比較するパラメトリック検定です。この検定にはいくつかのバリエーションが知られています。分散の等質性を検定する他の方法としては、ボックス検定、ボックス・アンダーソン検定、モーゼス検定などがあります。
確率分布の分散は、 古典力学における、重心周りの回転に関する、線に沿った対応する質量分布の慣性モーメントに類似している。 [ 27 ]この類似性から、分散のようなものは確率分布のモーメント と呼ばれる。[ 27 ]共分散行列は、多変量分布の慣性モーメントテンソルと関連している。共分散行列を持つn個の点のクラウドの慣性モーメントは、は
物理学と統計学における慣性モーメントの違いは、直線上に集まった点の場合に明らかです。多くの点がx軸の近くにあり、x軸に沿って分布していると仮定します。共分散行列は次のようになります 。
つまり、 x方向の変動が最も大きい。物理学者はこれをx軸周りのモーメントが小さいとみなすので、慣性モーメントテンソルは
半分散は分散と同じ方法で計算されますが、計算には平均値より低い観測値のみが含まれます。 また、さまざまな応用分野において、特定の尺度として説明されています。歪んだ分布の場合、半分散は分散では得られない追加情報を提供できます。[ 28 ]
半変性に関連する不等式については、チェビシェフの不等式 § 半変性 を参照してください。
分散という用語は、ロナルド・フィッシャーが1918年の論文「メンデル遺伝の仮定に基づく親族間の相関」で初めて導入した。[ 29 ]
入手可能な膨大な統計データは、人間の測定値の平均値からの偏差が正規誤差法則に非常に近いことを示し、したがって、変動性は平均二乗誤差の平方根に対応する標準偏差によって一様に測定できることを示しています。均一な母集団分布に標準偏差を持つ2つの独立した変動原因が存在する場合、そして両方の原因が同時に作用する場合、分布の標準偏差はしたがって、変動の原因を分析する際には、変動の尺度として標準偏差の二乗を用いることが望ましい。この量を分散と呼ぶことにする。
もしは、 の値をとるようなスカラー複素数値確率変数です。、その分散は、そこではの複素共役ですこの分散は実数スカラーです。
もしは、値の範囲が であるベクトル値確率変数です。列ベクトルとして考えると、分散の自然な一般化は次のようになる。どこそしてはの転置です、そして行ベクトルです。結果として得られるのは正の半正定値正方行列で、一般に分散共分散行列(または単に共分散行列)と呼ばれます。
もしは、値の範囲が であるベクトル値かつ複素数値の確率変数です。、すると共分散行列は、そこではの共役転置ですこの行列は正半定値かつ正方行列です。
ベクトル値確率変数の分散に関する別の一般化行列ではなくスカラー値となるのは、一般化分散である。共分散行列の行列式。一般化分散は、点の平均値を中心とした多次元散布と関連していることが示される。 [ 30 ]
スカラー分散の式を考慮すると、別の一般化が得られる。そして再解釈する確率変数とその平均値との間のユークリッド距離の二乗、または単にベクトルのスカラー積としてそれ自体と。その結果、これは共分散行列のトレースです。