統計用語
統計学 、 特に 回帰分析 において、てこ 比とは、ある 観測値 の 独立変数の 値が他の観測値の 独立変数の 値からどれだけ離れているかを示す尺度である。てこ 比の高い点 があれば、それは独立変数 に対する 外れ値 である。つまり、てこ比の高い点には 空間内に隣接する点がなく、 は 回帰モデル内の独立変数の数である。このため、適合モデルはてこ比の高い観測値に近づく可能性が高くなる。 [1]したがって、てこ比の高い点は、削除された場合、 すなわち影響力のある点 となった場合に、パラメータ推定値に大きな変化をもたらす可能性がある。影響力のある点は一般にてこ比が高いが、てこ比の高い点が必ずしも影響力のある点であるわけではない。てこ比は、一般に ハット行列 の対角要素として定義される 。 [2]
R
p
{\displaystyle \mathbb {R} ^{p}}
p
{\displaystyle {p}}
定義と解釈
線形回帰 モデル 、を 考えてみましょう 。つまり、 、ここで、 は行が観測値に対応し、列が独立変数または説明変数に対応する 設計行列 です。 独立観測値 の てこ比スコアは 次のように与えられます。
ええ
私
=
x
私
⊤
β
+
ε
私
{\displaystyle {y}_{i}={\boldsymbol {x}}_{i}^{\top }{\boldsymbol {\beta }}+{\varepsilon }_{i}}
私
=
1
、
2
、
…
、
ん
{\displaystyle i=1,\,2,\ldots ,\,n}
ええ
=
バツ
β
+
ε
{\displaystyle {\boldsymbol {y}}=\mathbf {X} {\boldsymbol {\beta }}+{\boldsymbol {\varepsilon }}}
バツ
{\displaystyle \mathbf {X} }
ん
×
p
{\displaystyle n\times p}
私
t
h
{\displaystyle {i}^{th}}
x
私
{\displaystyle {\boldsymbol {x}}_{i}}
h
私
私
=
[
H
]
私
私
=
x
私
⊤
(
バツ
⊤
バツ
)
−
1
x
私
{\displaystyle h_{ii}=\left[\mathbf {H} \right]_{ii}={\boldsymbol {x}}_{i}^{\top }\left(\mathbf {X} ^{\top }\mathbf {X} \right)^{-1}{\boldsymbol {x}}_{i}}
、 正射影行列 ( 別名 ハット行列) の対角要素 。
私
t
h
{\displaystyle {i}^{th}}
H
=
バツ
(
バツ
⊤
バツ
)
−
1
バツ
⊤
{\displaystyle \mathbf {H} =\mathbf {X} \left(\mathbf {X} ^{\top }\mathbf {X} \right)^{-1}\mathbf {X} ^{\top }}
したがって、レバレッジスコアは 、 の平均に対する の 間の「加重」距離として見ることができます (マハラノビス距離との関係を参照)。また、 測定された(従属)値(つまり、 )が 適合された(予測された)値(つまり、 ) に影響を与える 度合いとして解釈することもできます。 数学的には、
i
t
h
{\displaystyle {i}^{th}}
x
i
{\displaystyle {\boldsymbol {x}}_{i}}
x
i
{\displaystyle {\boldsymbol {x}}_{i}}
i
t
h
{\displaystyle {i}^{th}}
y
i
{\displaystyle y_{i}}
i
t
h
{\displaystyle {i}^{th}}
y
^
i
{\displaystyle {\widehat {y\,}}_{i}}
h
i
i
=
∂
y
^
i
∂
y
i
{\displaystyle h_{ii}={\frac {\partial {\widehat {y\,}}_{i}}{\partial y_{i}}}}
。
したがって、レバレッジスコアは、観測自己感度または自己影響とも呼ばれます。 [3] 上記の式で (つまり、予測は の範囲空間への の正射影である) という事実を使用すると、 が得られます 。このレバレッジは、すべての観測の説明変数の値に依存します が、従属変数の値には依存しないことに注意してください 。
y
^
=
H
y
{\displaystyle {\boldsymbol {\widehat {y}}}={\mathbf {H} }{\boldsymbol {y}}}
y
^
{\displaystyle {\boldsymbol {\widehat {y}}}}
y
{\displaystyle {\boldsymbol {y}}}
X
{\displaystyle \mathbf {X} }
h
i
i
=
[
H
]
i
i
{\displaystyle h_{ii}=\left[\mathbf {H} \right]_{ii}}
(
X
)
{\displaystyle (\mathbf {X} )}
(
y
i
)
{\displaystyle (y_{i})}
プロパティ
レバレッジは 0 から 1 の間の数値です。 証明: はべき等行列 ( ) かつ対称行列 ( )である ことに注意してください 。したがって、 という事実を使用することで 、 が得られます 。 であることがわかっているので 、 が得られます 。
h
i
i
{\displaystyle h_{ii}}
0
≤
h
i
i
≤
1.
{\displaystyle 0\leq h_{ii}\leq 1.}
H
{\displaystyle \mathbf {H} }
H
2
=
H
{\displaystyle \mathbf {H} ^{2}=\mathbf {H} }
h
i
j
=
h
j
i
{\displaystyle h_{ij}=h_{ji}}
[
H
2
]
i
i
=
[
H
]
i
i
{\displaystyle \left[\mathbf {H} ^{2}\right]_{ii}=\left[\mathbf {H} \right]_{ii}}
h
i
i
=
h
i
i
2
+
∑
j
≠
i
h
i
j
2
{\displaystyle h_{ii}=h_{ii}^{2}+\sum _{j\neq i}h_{ij}^{2}}
∑
j
≠
i
h
i
j
2
≥
0
{\displaystyle \sum _{j\neq i}h_{ij}^{2}\geq 0}
h
i
i
≥
h
i
i
2
⟹
0
≤
h
i
i
≤
1
{\displaystyle h_{ii}\geq h_{ii}^{2}\implies 0\leq h_{ii}\leq 1}
レバレッジの合計は、(切片を含む) の パラメータの数に等しくなります。 証明: 。
(
p
)
{\displaystyle (p)}
β
{\displaystyle {\boldsymbol {\beta }}}
∑
i
=
1
n
h
i
i
=
Tr
(
H
)
=
Tr
(
X
(
X
⊤
X
)
−
1
X
⊤
)
=
Tr
(
X
⊤
X
(
X
⊤
X
)
−
1
)
=
Tr
(
I
p
)
=
p
{\displaystyle \sum _{i=1}^{n}h_{ii}=\operatorname {Tr} (\mathbf {H} )=\operatorname {Tr} \left(\mathbf {X} \left(\mathbf {X} ^{\top }\mathbf {X} \right)^{-1}\mathbf {X} ^{\top }\right)=\operatorname {Tr} \left(\mathbf {X} ^{\top }\mathbf {X} \left(\mathbf {X} ^{\top }\mathbf {X} \right)^{-1}\right)=\operatorname {Tr} (\mathbf {I} _{p})=p}
レバレッジを用いたXの外れ値の判定
大きなレバレッジは、 極端な に対応します。一般的なルールは、 レバレッジ値が 平均レバレッジの 2 倍以上であるを特定することです (上記の特性 2 を参照)。つまり、 の場合 、 は外れ値と見なされます。統計学者の中 には、 ではなく のしきい値を好む人もいます 。
h
i
i
{\displaystyle {h_{ii}}}
x
i
{\displaystyle {{\boldsymbol {x}}_{i}}}
x
i
{\displaystyle {{\boldsymbol {x}}_{i}}}
h
i
i
{\displaystyle {h}_{ii}}
h
¯
=
1
n
∑
i
=
1
n
h
i
i
=
p
n
{\displaystyle {\bar {h}}={\dfrac {1}{n}}\sum _{i=1}^{n}h_{ii}={\dfrac {p}{n}}}
h
i
i
>
2
p
n
{\displaystyle h_{ii}>2{\dfrac {p}{n}}}
x
i
{\displaystyle {{\boldsymbol {x}}_{i}}}
3
p
/
n
{\displaystyle 3p/{n}}
2
p
/
n
{\displaystyle 2p/{n}}
マハラノビス距離との関係
レバレッジはマハラノビス距離と密接な関係があります(証明 [4] )。具体的には、ある 行列に対して、 長さ の 平均のベクトルからの ( の行 ) の マハラノビス距離の2乗 は で 、 は の 推定 共 分散行列 です。これは、1の列ベクトルを追加した後のハット行列の レバレッジと関係があります 。2つの関係は次のとおりです。
n
×
p
{\displaystyle n\times p}
X
{\displaystyle \mathbf {X} }
x
i
{\displaystyle {{\boldsymbol {x}}_{i}}}
x
i
⊤
{\displaystyle {\boldsymbol {x}}_{i}^{\top }}
i
t
h
{\displaystyle {i}^{th}}
X
{\displaystyle \mathbf {X} }
μ
^
=
∑
i
=
1
n
x
i
{\displaystyle {\widehat {\boldsymbol {\mu }}}=\sum _{i=1}^{n}{\boldsymbol {x}}_{i}}
p
{\displaystyle p}
D
2
(
x
i
)
=
(
x
i
−
μ
^
)
⊤
S
−
1
(
x
i
−
μ
^
)
{\displaystyle D^{2}({\boldsymbol {x}}_{i})=({\boldsymbol {x}}_{i}-{\widehat {\boldsymbol {\mu }}})^{\top }\mathbf {S} ^{-1}({\boldsymbol {x}}_{i}-{\widehat {\boldsymbol {\mu }}})}
S
=
X
⊤
X
{\displaystyle \mathbf {S} =\mathbf {X} ^{\top }\mathbf {X} }
x
i
{\displaystyle {{\boldsymbol {x}}_{i}}}
h
i
i
{\displaystyle h_{ii}}
X
{\displaystyle \mathbf {X} }
D
2
(
x
i
)
=
(
n
−
1
)
(
h
i
i
−
1
n
)
{\displaystyle D^{2}({\boldsymbol {x}}_{i})=(n-1)(h_{ii}-{\tfrac {1}{n}})}
この関係により、レバレッジを意味のある要素に分解することができ、高いレバレッジの原因を分析的に調査できるようになります。 [5]
影響関数との関係
回帰分析では、てこ関数と 影響関数 を組み合わせて、1つのデータポイントを削除した場合に推定係数がどの程度変化するかを計算します。回帰残差を と表すと、 式 [6] [7] を使用して推定係数を1つ除外した推定係数と 比較することができます。
e
^
i
=
y
i
−
x
i
⊤
β
^
{\displaystyle {\widehat {e}}_{i}=y_{i}-{\boldsymbol {x}}_{i}^{\top }{\widehat {\boldsymbol {\beta }}}}
β
^
{\displaystyle {\widehat {\boldsymbol {\beta }}}}
β
^
(
−
i
)
{\displaystyle {\widehat {\boldsymbol {\beta }}}^{(-i)}}
β
^
−
β
^
(
−
i
)
=
(
X
⊤
X
)
−
1
x
i
e
^
i
1
−
h
i
i
{\displaystyle {\widehat {\boldsymbol {\beta }}}-{\widehat {\boldsymbol {\beta }}}^{(-i)}={\frac {(\mathbf {X} ^{\top }\mathbf {X} )^{-1}{\boldsymbol {x}}_{i}{\widehat {e}}_{i}}{1-h_{ii}}}}
Young (2019) は、残差制御後のこの式のバージョンを使用しています。 [8] この式を直感的に理解するには、が 観測値が回帰パラメータに影響を及ぼす可能性を捉えていること、したがって、 回帰パラメータに対するその観測値の適合値からの偏差の実際の影響を捉えていることに注目してください。次に、式は で割って、 観測値を調整するのではなく削除するという事実を考慮に入れています。これは、除去によって共変量の分布がより変化するという事実を反映しており、レバレッジの高い観測値(つまり、外れ値の共変量値)に適用した場合、共変量の分布がより変化するという事実を反映しています。回帰コンテキストで統計的影響関数の一般的な式を適用すると、同様の式が生まれます。 [9] [10]
∂
β
^
∂
y
i
=
(
X
⊤
X
)
−
1
x
i
{\displaystyle {\frac {\partial {\hat {\beta }}}{\partial y_{i}}}=(\mathbf {X} ^{\top }\mathbf {X} )^{-1}{\boldsymbol {x}}_{i}}
(
X
⊤
X
)
−
1
x
i
e
^
i
{\displaystyle (\mathbf {X} ^{\top }\mathbf {X} )^{-1}{\boldsymbol {x}}_{i}{\widehat {e}}_{i}}
(
1
−
h
i
i
)
{\displaystyle (1-h_{ii})}
残差分散への影響
固定された 等 分散回帰誤差 を持つ 通常の最小二乗 設定の場合 、 回帰残差 は 分散を持つ。
X
{\displaystyle \mathbf {X} }
ε
i
,
{\displaystyle \varepsilon _{i},}
y
=
X
β
+
ε
;
Var
(
ε
)
=
σ
2
I
{\displaystyle {\boldsymbol {y}}=\mathbf {X} {\boldsymbol {\beta }}+{\boldsymbol {\varepsilon }};\ \ \operatorname {Var} ({\boldsymbol {\varepsilon }})=\sigma ^{2}\mathbf {I} }
i
t
h
{\displaystyle {i}^{th}}
e
i
=
y
i
−
y
^
i
{\displaystyle e_{i}=y_{i}-{\widehat {y}}_{i}}
Var
(
e
i
)
=
(
1
−
h
i
i
)
σ
2
{\displaystyle \operatorname {Var} (e_{i})=(1-h_{ii})\sigma ^{2}}
。
言い換えれば、観測値のてこ比スコアは、その観測値に対するモデルの予測誤りのノイズの度合いを決定し、てこ比が高いほどノイズは少なくなります。これは、 がべき等かつ対称であり、 したがって であるという事実から導かれます 。
I
−
H
{\displaystyle \mathbf {I} -\mathbf {H} }
y
^
=
H
y
{\displaystyle {\widehat {\boldsymbol {y}}}=\mathbf {H} {\boldsymbol {y}}}
Var
(
e
)
=
Var
(
(
I
−
H
)
y
)
=
(
I
−
H
)
Var
(
y
)
(
I
−
H
)
⊤
=
σ
2
(
I
−
H
)
2
=
σ
2
(
I
−
H
)
{\displaystyle \operatorname {Var} ({\boldsymbol {e}})=\operatorname {Var} ((\mathbf {I} -\mathbf {H} ){\boldsymbol {y}})=(\mathbf {I} -\mathbf {H} )\operatorname {Var} ({\boldsymbol {y}})(\mathbf {I} -\mathbf {H} )^{\top }=\sigma ^{2}(\mathbf {I} -\mathbf {H} )^{2}=\sigma ^{2}(\mathbf {I} -\mathbf {H} )}
対応する スチューデント化残差 (観測値固有の推定残差分散を調整した残差)は、
t
i
=
e
i
σ
^
1
−
h
i
i
{\displaystyle t_{i}={e_{i} \over {\widehat {\sigma }}{\sqrt {1-h_{ii}\ }}}}
ここで は の適切な推定値です 。
σ
^
{\displaystyle {\widehat {\sigma }}}
σ
{\displaystyle \sigma }
部分的レバレッジ
部分レバレッジ ( PL ) は、各観測値の総レバレッジに対する個々の 独立変数 の寄与度を測る尺度です。つまり、PL は、 変数としての変更が回帰モデルにどの程度追加されるかの尺度です。次のように計算されます。
h
i
i
{\displaystyle h_{ii}}
(
P
L
j
)
i
=
(
X
j
∙
[
j
]
)
i
2
∑
k
=
1
n
(
X
j
∙
[
j
]
)
k
2
{\displaystyle \left(\mathrm {PL} _{j}\right)_{i}={\frac {\left(\mathbf {X} _{j\bullet [j]}\right)_{i}^{2}}{\sum _{k=1}^{n}\left(\mathbf {X} _{j\bullet [j]}\right)_{k}^{2}}}}
ここで、 は独立変数のインデックス、 は観測のインデックス、は 残りの独立変数に対する 回帰からの 残差 です。部分的てこ比は、変数 の 部分回帰プロット 内の点のてこ比であることに注意してください 。独立変数の部分的てこ比が大きいデータ ポイントは、自動回帰モデル構築手順でその変数の選択に過度の影響を及ぼす可能性があります。
j
{\displaystyle j}
i
{\displaystyle i}
X
j
∙
[
j
]
{\displaystyle \mathbf {X} _{j\bullet [j]}}
X
j
{\displaystyle \mathbf {X} _{j}}
i
t
h
{\displaystyle {i}^{th}}
j
t
h
{\displaystyle {j}^{th}}
ソフトウェア実装
R 、 Python などの多くのプログラムや統計パッケージには 、Leverage の実装が含まれています。
参照
参考文献
^ Everitt, BS (2002). ケンブリッジ統計辞典 . ケンブリッジ大学出版局. ISBN 0-521-81099-X 。
^ ジェームズ、ガレス; ウィッテン、ダニエラ; ハスティー、トレバー; ティブシラニ、ロバート (2021)。統計学習入門: R での応用 (第 2 版)。ニューヨーク、NY: シュプリンガー。p. 112。ISBN 978-1-0716-1418-1 . 2024年 10月29日 閲覧 。
^ Cardinali, C. (2013 年 6 月). 「データ同化: データ同化システムの観測影響診断」 (PDF) 。
^ マハラノビス距離とレバレッジの関係を証明しますか?
^ Kim, MG (2004). 「線形回帰モデルにおける高レバレッジの源 (Journal of Applied Mathematics and Computing, Vol 16, 509–513)」. arXiv : 2006.04024 [math.ST].
^ ミラー 、ルパート ・ G. (1974年9月)。「アンバランス・ジャックナイフ」 。Annals of Statistics。2 (5): 880–891。doi : 10.1214 /aos/1176342811。ISSN 0090-5364 。
^ ヒヤシ・フミオ(2000年) 『計量経済学 』プリンストン大学出版局、21頁。
^ Young, Alwyn (2019). 「チャネリング・フィッシャー:ランダム化テストと 一見 有意な実験結果の統計的無意味さ」 季刊経済学ジャーナル 。134 (2): 567. doi : 10.1093/qje/qjy029 。
^ Chatterjee, Samprit; Hadi, Ali S. (1986年8月). 「線形回帰における影響力のある観測、高いてこポイント、外れ値」. 統計科学 . 1 (3): 379–393. doi : 10.1214/ss/1177013622 . ISSN 0883-4237.
^ 「回帰 - 影響関数とOLS」。Cross Validated 。 2020年12月6日 閲覧。