A likelihood function (often simply called the likelihood) measures how well a statistical model explains observed data by calculating the probability of seeing that data under different parameter values of the model. It is constructed from the joint probability distribution of the random variable that (presumably) generated the observations.[1][2][3] When evaluated on the actual data points, it becomes a function solely of the model parameters.
In maximum likelihood estimation, the model parameter(s) or argument that maximizes the likelihood function serves as a point estimate for the unknown parameter, while the Fisher information (often approximated by the likelihood's Hessian matrix at the maximum) gives an indication of the estimate's precision.
In contrast, in Bayesian statistics, the estimate of interest is the converse of the likelihood, the so-called posterior probability of the parameter given the observed data, which is calculated via Bayes' rule.[4]
The likelihood function, parameterized by a (possibly multivariate) parameter , is usually defined differently for discrete and continuousprobability distributions (a more general definition is discussed below). Given a probability density or mass function
where is a realization of the random variable , the likelihood function is often written
In other words, when is viewed as a function of with fixed, it is a probability density function, and when viewed as a function of with fixed, it is a likelihood function. In the frequentist paradigm, the notation is often avoided and instead or are used to indicate that is regarded as a fixed unknown quantity rather than as a random variable being conditioned on.
The likelihood function does not specify the probability that is the truth, given the observed sample . Such an interpretation is a common error, with potentially disastrous consequences (see prosecutor's fallacy).
Let be a discrete random variable with probability mass function depending on a parameter . Then the function
considered as a function of , a possible value of the deterministic but unknown parameter , is the likelihood function, given the outcome of the random variable . Sometimes the probability of "the value of for the parameter value " is written as P(X = x | θ) or P(X = x; θ). The likelihood is the probability that a particular outcome is observed when the true value of the parameter is , equivalent to the probability mass on ; it is not a probability density over the parameter . The likelihood, , should not be confused with , which is the posterior probability of given the data .


Consider a simple statistical model of a coin flip: a single parameter that expresses the "fairness" of the coin. The parameter is the probability that a coin lands heads up ("H") when tossed. can take on any value within the range 0.0 to 1.0. For a perfectly fair coin, .
Imagine flipping a fair coin twice, and observing two heads in two tosses ("HH"). Assuming that each successive coin flip is i.i.d., then the probability of observing HH is
Equivalently, the likelihood of observing "HH" assuming is
This is not the same as saying that , a conclusion which could only be reached via Bayes' theorem given knowledge about the marginal probabilities and .
Now suppose that the coin is not a fair coin, but instead that . Then the probability of two heads on two flips is
Hence
More generally, for each value of , we can calculate the corresponding likelihood. The result of such calculations is displayed in Figure 1. The integral of over [0, 1] is 1/3; likelihoods need not integrate or sum to one over the parameter space.
Let be a random variable following an absolutely continuous probability distribution with density function (a function of ) which depends on a parameter . Then the function
considered as a function of , is the likelihood function (of , given the outcome). Again, is not a probability density or mass function over , despite being a function of given the observation .
The use of the probability density in specifying the likelihood function above is justified as follows. Given an observation , the likelihood for the interval , where is a constant, is given by . Observe that since is positive and constant. Because
where is the probability density function, it follows that
The first fundamental theorem of calculus provides that
Then
Therefore, and so maximizing the probability density at amounts to maximizing the likelihood of the specific observation .
In measure-theoretic probability theory, the density function is defined as the Radon–Nikodym derivative of the probability distribution relative to a common dominating measure.[5] The likelihood function is this density interpreted as a function of the parameter, rather than the random variable.[6] Thus, we can construct a likelihood function for any distribution, whether discrete, continuous, a mixture, or otherwise. (Likelihoods are comparable, e.g. for parameter estimation, only if they are Radon–Nikodym derivatives with respect to the same dominating measure.)
The above discussion of the likelihood for discrete random variables uses the counting measure, under which the probability density at any outcome equals the probability of that outcome.
The above can be extended in a simple way to allow consideration of distributions which contain both discrete and continuous components. Suppose that the distribution consists of a number of discrete probability masses and a density , where the sum of all the 's added to the integral of is always one. Assuming that it is possible to distinguish an observation corresponding to one of the discrete probability masses from one which corresponds to the density component, the likelihood function for an observation from the continuous component can be dealt with in the manner shown above. For an observation from the discrete component, the likelihood function for an observation from the discrete component is simply where is the index of the discrete probability mass corresponding to observation , because maximizing the probability mass (or probability) at amounts to maximizing the likelihood of the specific observation.
The fact that the likelihood function can be defined in a way that includes contributions that are not commensurate (the density and the probability mass) arises from the way in which the likelihood function is defined up to a constant of proportionality, where this "constant" can change with the observation , but not with the parameter .
In the context of parameter estimation, the likelihood function is usually assumed to obey certain conditions, known as regularity conditions. These conditions are assumed in various proofs involving likelihood functions, and need to be verified in each particular application. For maximum likelihood estimation, the existence of a global maximum of the likelihood function is of the utmost importance. By the extreme value theorem, it suffices that the likelihood function is continuous on a compact parameter space for the maximum likelihood estimator to exist.[7] While the continuity assumption is usually met, the compactness assumption about the parameter space is often not, as the bounds of the true parameter values might be unknown. In that case, concavity of the likelihood function plays a key role.
More specifically, if the likelihood function is twice continuously differentiable on the k-dimensional parameter space assumed to be an openconnected subset of there exists a unique maximum if the matrix of second partials is negative definite for every at which the gradient vanishes, and if the likelihood function approaches a constant on the boundary of the parameter space, i.e., which may include the points at infinity if is unbounded. Mäkeläinen and co-authors prove this result using Morse theory while informally appealing to a mountain pass property.[8] Mascarenhas restates their proof using the mountain pass theorem.[9]
In the proofs of consistency and asymptotic normality of the maximum likelihood estimator, additional assumptions are made about the probability densities that form the basis of a particular likelihood function. These conditions were first established by Chanda.[10] In particular, for almost all, and for all exist for all in order to ensure the existence of a Taylor expansion. Second, for almost all and for every it must be that where is such that This boundedness of the derivatives is needed to allow for differentiation under the integral sign. And lastly, it is assumed that the information matrix, is positive definite and is finite. This ensures that the score has a finite variance.[11]
上記の条件は十分条件ではあるが、必要条件ではない。つまり、これらの正則性条件を満たさないモデルは、上述の特性の最尤推定量を持つ場合もあれば、持たない場合もある。さらに、観測値が独立または同一分布に従わない場合には、追加の特性を仮定する必要があるかもしれない。
ベイズ統計学では、事後確率の漸近正規性を証明するため、尤度関数にほぼ同一の正則性条件が課せられ、[ 12 ] [ 13 ]したがって、大標本における事後確率のラプラス近似を正当化する。[ 14 ]
尤度比とは、指定された2つの尤度の比であり、多くの場合次のように表記されます。
尤度比は尤度統計学の中心となる概念である。尤度の法則によれば、データ(証拠とみなされる)が一方のパラメータ値を他方のパラメータ値よりもどの程度支持しているかは、尤度比によって測定される。
頻度論的推論では、尤度比は検定統計量の基礎であり、いわゆる尤度比検定である。ネイマン・ピアソン補題によれば、これは与えられた有意水準で2つの単純な仮説を比較するための最も強力な検定である。他の多くの検定は、尤度比検定またはその近似と見なすことができる。[ 15 ]検定統計量として考えられる対数尤度比の漸近分布は、ウィルクスの定理によって与えられる。
尤度比はベイズ推論においても中心的な重要性を持ち、ベイズ因子として知られ、ベイズの定理で使用されます。オッズの観点から述べると、ベイズの定理は、2 つの選択肢の事後オッズが、そして、イベントが与えられた場合 は、事前オッズに尤度比を掛けたものです。数式で表すと次のようになります。
尤度比はAICに基づく統計手法では直接使用されません。代わりに、モデルの相対尤度が使用されます(下記参照)。
根拠に基づいた医療では、診断検査を実施する価値を評価するために、診断検査において尤度比が用いられる。
尤度関数の実際の値はサンプルに依存するため、標準化された尺度を用いると便利な場合が多い。パラメータθの最尤推定値が次のようになると仮定する。他のθ値の相対的な妥当性は、それらの他の値の尤度と、θの相対尤度は次のように定義されます[ 16 ] [ 17 ] [ 18 ] [ 19 ] [ 20 ] したがって、相対尤度は、固定分母を持つ尤度比(上記で説明した)である。これは、最大確率が1になるように標準化することに相当します。
尤度領域とは、相対尤度が与えられた閾値以上であるθのすべての値の集合です。パーセンテージで表すと、 θのp %尤度領域は次のように定義されます[ 16 ] [ 18 ] [ 21 ] 。
θが単一の実数パラメータである場合、 p %尤度領域は通常、実数値の区間で構成されます。領域が区間で構成されている場合、それは尤度区間と呼ばれます。[ 16 ] [ 18 ] [ 22 ]
尤度区間、より一般的には尤度領域は、尤度統計学における区間推定に用いられます。これらは、頻度統計学における信頼区間やベイズ統計学における信用区間に類似しています。尤度区間は、包含確率(頻度主義)や事後確率(ベイズ主義)ではなく、相対尤度の観点から直接解釈されます。
モデルが与えられた場合、尤度区間は信頼区間と比較できます。θ が単一の実数パラメータである場合、特定の条件下では、θの 14.65% 尤度区間 (約 1:7 の尤度)は 95% 信頼区間 (19/20 のカバレッジ確率) と同じになります。[ 16 ] [ 21 ]対数尤度の使用に適した少し異なる定式化 (ウィルクスの定理を参照) では、検定統計量は対数尤度の差の 2 倍であり、検定統計量の確率分布は、2 つのモデル間の自由度の差に等しい自由度 (df) を持つカイ二乗分布に近似します (したがって、 e − 2尤度区間は 0.954 信頼区間と同じです。df の差が 1 であると仮定します)。[ 21 ] [ 22 ]
多くの場合、尤度は複数のパラメータの関数ですが、関心は 1 つまたはせいぜい数個のパラメータの推定に集中し、他のパラメータは不要パラメータとみなされます。このような不要パラメータを排除して尤度を関心のあるパラメータ (または複数のパラメータ) の関数として記述できるようにするために、いくつかの代替アプローチが開発されています。主なアプローチは、プロファイル尤度、条件付き尤度、周辺尤度です。[ 23 ] [ 24 ]これらのアプローチは、グラフを可能にするために、高次元の尤度曲面を 1 つまたは 2 つの関心のあるパラメータに縮小する必要がある場合にも役立ちます。
パラメータのサブセットに対する尤度関数を集中させることで次元を削減することが可能であり、そのためには、不要なパラメータを関心のあるパラメータの関数として表現し、それらを尤度関数に置き換える必要がある。[ 25 ] [ 26 ]一般に、パラメータベクトルに依存する尤度関数については、分割できる、そして対応関係がある場合明示的に決定することができ、集中により元の最大化問題の計算負荷が軽減される。 [ 27 ]
例えば、誤差が正規分布する線形回帰では、係数ベクトルは以下のように分割できます。(そして結果として設計マトリックスも))に関して最大化する最適な値関数が得られるこの結果を用いると、すると、次のように導き出せる。 どこ is the projection matrix of . This result is known as the Frisch–Waugh–Lovell theorem.
Since graphically the procedure of concentration is equivalent to slicing the likelihood surface along the ridge of values of the nuisance parameter that maximizes the likelihood function, creating an isometricprofile of the likelihood function for a given , the result of this procedure is also known as profile likelihood.[28][29] In addition to being graphed, the profile likelihood can also be used to compute confidence intervals that often have better small-sample properties than those based on asymptotic standard errors calculated from the full likelihood.[30][31]
Sometimes it is possible to find a sufficient statistic for the nuisance parameters, and conditioning on this statistic results in a likelihood which does not depend on the nuisance parameters.[32]
One example occurs in 2×2 tables, where conditioning on all four marginal totals leads to a conditional likelihood based on the non-central hypergeometric distribution. This form of conditioning is also the basis for Fisher's exact test.
Sometimes we can remove the nuisance parameters by considering a likelihood based on only part of the information in the data, for example by using the set of ranks rather than the numerical values. Another example occurs in linear mixed models, where considering a likelihood for the residuals only after fitting the fixed effects leads to residual maximum likelihood estimation of the variance components.
A partial likelihood is an adaption of the full likelihood such that only a part of the parameters (the parameters of interest) occur in it.[33] It is a key component of the proportional hazards model: using a restriction on the hazard function, the likelihood does not contain the shape of the hazard over time.
The likelihood, given two or more independentevents, is the product of the likelihoods of each of the individual events: This follows from the definition of independence in probability: the probabilities of two independent events happening, given a model, is the product of the probabilities.
これは、事象が独立同分布の確率変数から生じる場合、例えば独立観測や復元抽出によるサンプリングの場合に特に重要です。このような状況では、尤度関数は個々の尤度関数の積に分解されます。
空の積の値は1であり、これは事象が発生しない場合の尤度が1であることに対応します。つまり、データが与えられる前は、尤度は常に1です。これはベイズ統計における一様事前分布に似ていますが、尤度統計では尤度は積分されないため、これは不適切な事前分布ではありません。
対数尤度関数は尤度関数の対数であり、小文字のlまたはで表されることが多い。大文字のLと対比して、尤度については、対数は厳密に増加する関数であるため、尤度を最大化することは対数尤度を最大化することと同等です。しかし、実際的な目的においては、最尤推定では対数尤度関数を使用する方が便利です。特に、最も一般的な確率分布(特に指数族)は対数的に凹であるだけであり、[ 34 ] [ 35 ]目的関数の凹性が最大化において重要な役割を果たします。
各事象が独立していることを前提とすると、交差の全体的な対数尤度は、個々の事象の対数尤度の合計に等しくなります。これは、全体的な対数確率が個々の事象の対数確率の合計であるという事実と類似しています。この数学的な利便性に加えて、対数尤度の加算プロセスには、データからの「支持」としてよく表現されるように、直感的な解釈があります。最尤推定のために対数尤度を使用してパラメータを推定する場合、各データポイントは全体の対数尤度に加算されます。データは推定されたパラメータを支持する証拠と見なすことができるため、このプロセスは「独立した証拠からの支持が加算される」と解釈でき、対数尤度は「証拠の重み」となります。負の対数確率を情報量または驚き度として解釈すると、ある事象が与えられた場合のモデルの支持度(対数尤度)は、そのモデルが与えられた場合の事象の驚き度の負の値になります。つまり、その事象がモデルが与えられた場合に驚き度が低いほど、モデルはその事象によって支持されると言えます。
尤度比の対数は、対数尤度の差に等しい。
事象が発生しない場合の尤度が1であるのと同様に、事象が発生しない場合の対数尤度は0であり、これは空の合計値に対応します。つまり、データがなければ、どのモデルも支持されません。
対数尤度のグラフは、(単変量の場合)サポート曲線と呼ばれます。[ 36 ]多変量の場合、 この概念はパラメータ空間上のサポート曲面に一般化されます。これは分布のサポートと関連がありますが、それとは異なります。
この用語は、統計的仮説検定の文脈でAWF エドワーズ[ 36 ]によって造語されました。つまり、データがテストされている仮説 (またはパラメータ値) のうちの 1 つを他の仮説よりも「支持」しているかどうかということです。
グラフにプロットされている対数尤度関数は、スコア(対数尤度の勾配)とフィッシャー情報量(対数尤度の曲率)の計算に使用されます。したがって、このグラフは最尤推定法と尤度比検定の文脈において直接的な解釈が可能です。
対数尤度関数が滑らかである場合、パラメータに関するその勾配はスコアとして知られ、次のように表されます。が存在し、微分計算の適用を可能にします。微分可能な関数を最大化する基本的な方法は、停留点(導関数がゼロになる点)を見つけることです。和の導関数は導関数の和に等しいですが、積の導関数には積の法則が必要なので、独立事象の尤度よりも独立事象の対数尤度の停留点を計算する方が簡単です。
スコア関数の停留点によって定義される方程式は、最尤推定量の 推定方程式として機能する。 その意味で、最尤推定量は暗黙のうちに次の値によって定義される。逆関数の、 どこはd次元ユークリッド空間であり、はパラメータ空間である。逆関数定理を用いると、次のことが示される。オープンな近隣地域で明確に定義されています確率は1に近づき、は、結果として、数列が存在するそのため漸近的にほぼ確実に、そして[ 37 ]ロルの定理を用いて同様の結果を確立することができる。[ 38 ] [ 39 ]
2階微分を評価したところ、フィッシャー情報として知られるこの値は、尤度曲面の曲率を決定し、[ 40 ]推定値の精度を示します。 [ 41 ]
対数尤度は、多くの一般的なパラメトリック確率分布を含む指数分布族にも特に有用です。指数分布族の確率分布関数(したがって尤度関数)は、指数を含む因子の積を含んでいます。このような関数の対数は積の和であり、元の関数よりも微分が容易です。
指数族とは、確率密度関数が次の形式である分布族のことである(一部の関数では、次のように書くことができる)。内積の場合):
これらの各項には解釈がありますが、[ a ]確率から尤度に切り替えて対数を取るだけで、次の合計が得られます。
のそしてそれぞれが座標変換に対応しているため、これらの座標では、指数型分布族の対数尤度は次の単純な式で与えられます。
言い換えれば、指数型分布族の対数尤度は自然パラメータの内積である。そして十分統計量正規化係数(対数分割関数)を差し引いた値したがって、例えば、十分統計量Tと対数分割関数Aの導関数を取ることによって、最尤推定値を計算できます。
ガンマ分布は2つのパラメータを持つ指数族である。そして尤度関数は
最尤推定値を見つける単一の観測値に対してかなり難しそうに見える。対数の方がずっと扱いやすい。
対数尤度を最大化するために、まずに関して偏微分をとります。:
複数の独立した観測がある場合すると、結合対数尤度は個々の対数尤度の合計となり、この合計の導関数は各個々の対数尤度の導関数の合計となる。
結合対数尤度の最大化手順を完了するために、方程式をゼロに設定して解きます。:
ここは最尤推定値を表し、これは観測値の標本平均です。
「尤度」という用語は、少なくとも中期英語後期から英語で使われてきました。[ 42 ]数学統計学における特定の関数を指す正式な用語として、ロナルド・フィッシャー[ 43 ]が1921年[ 44 ]と1922年[ 45 ]に発表した2つの研究論文で提唱しました。1921年の論文では、今日「尤度区間」と呼ばれるものが導入され、1922年の論文では「最尤法」という用語が導入されました。フィッシャーの言葉を引用すると、次のようになります。
1922年に私は「尤度」という用語を提案しました。これは、[パラメータ]に関して、尤度は確率ではなく、確率の法則に従わない一方で、[パラメータ]の可能な値の中から合理的に選択するという問題に対して、確率が偶然のゲームにおける事象の予測という問題に対して持つ関係に似た関係を持つという事実に基づいています。…しかし、心理的判断に関して、尤度は確率とある程度似ていますが、この2つの概念は全く異なります。…」[ 46 ]
ロナルド・フィッシャー卿が述べたように、可能性の概念は確率と混同してはならない。
私がこの点を強調するのは、確率と尤度の違いを常に強調してきたにもかかわらず、尤度を一種の確率であるかのように扱う傾向が依然としてあるためです。したがって、最初の結果は、異なるケースに適した2つの異なる合理的信念の尺度が存在するということです。母集団を知っている場合、標本に関する不完全な知識または期待を確率で表現できます。標本を知っている場合、母集団に関する不完全な知識を尤度で表現できます。[ 47 ]
フィッシャーによる統計的尤度の発明は、逆確率と呼ばれる以前の推論形式に対する反動であった。[ 48 ]彼が「尤度」という用語を使用したことで、数学統計学におけるその用語の意味が固定された。
AWF エドワーズ(1972) は、対数尤度比をある仮説に対する別の仮説の相対的な支持の尺度として使用するための公理的基礎を確立しました。支持関数は、尤度関数の自然対数です。これらの用語はどちらも系統発生学で使用されていますが、統計的証拠のトピックの一般的な扱いには採用されていません。[ 49 ]
統計学者の間では、統計の基礎が何であるべきかについて合意はありません。基礎として提案されている主なパラダイムは、頻度主義、ベイズ主義、尤度主義、AICベースという4つです。[ 50 ]提案されているそれぞれの基礎において、尤度の解釈は異なります。4つの解釈については、以下のサブセクションで説明します。
ベイズ推論では、ある命題または確率変数が別の確率変数を与えられた場合の尤度について話すことができます。たとえば、特定のデータまたはその他の証拠が与えられた場合のパラメータ値または統計モデルの尤度(周辺尤度を参照) などです。[ 51 ] [ 52 ] [ 53 ] [ 54 ]尤度関数は同じ実体のままですが、(i)パラメータが与えられた場合のデータの条件付き密度(パラメータは確率変数であるため) および (ii) パラメータ値またはモデルに関するデータによってもたらされる情報の尺度または量という追加の解釈があります。 [ 51 ] [ 52 ] [ 53 ] [ 54 ] [ 55 ]パラメータ空間またはモデルの集合に確率構造が導入されたことにより、パラメータ値または統計モデルが、与えられたデータに対して大きな尤度値を持つにもかかわらず、低い確率を持つ、またはその逆の可能性があります。[ 53 ] [ 55 ]これは医療の文脈でよく見られます。[ 56 ]ベイズの定理に従うと、条件付き密度として見た場合の尤度は、パラメータの事前確率密度を乗じて正規化することで、事後確率密度が得られます。[ 51 ] [ 52 ] [ 53 ] [ 54 ] [ 55 ]より一般的には、未知の量の尤度は別の未知の量が与えられた場合確率に比例する与えられた[ 51 ] [ 52 ] [ 53 ] [ 54 ] [ 55 ]
頻度統計学では、尤度関数自体が母集団から抽出した単一の標本を要約する統計量であり、その計算値は、いくつかのパラメータθ 1 ... θ pの選択に依存します。ここで、pは既に選択された統計モデルにおけるパラメータの数です。尤度の値は、パラメータの選択に対する評価指標として機能し、利用可能なデータに基づいて、尤度が最大となるパラメータセットが最良の選択となります。
尤度の具体的な計算は、選択されたモデルといくつかのパラメータθの値が、観測されたサンプルが抽出された母集団の頻度分布を正確に近似していると仮定した場合に、観測されたサンプルが割り当てられる確率です。経験的に、実際に観測されたサンプルが起こった事後確率を最大にするようなパラメータを選択するのが良い選択であるというのは理にかなっています。ウィルクスの定理は、推定値のパラメータ値によって生成される尤度の対数と、母集団の「真の」(ただし未知の)パラメータ値によって生成される尤度の対数の差が漸近的にχ 2分布に従うことを示すことで、この経験則を定量化しています。
各独立標本の最尤推定値は、標本抽出された母集団を記述する「真の」パラメータセットの個別の推定値です。多数の独立標本から得られる連続的な推定値は、母集団の「真の」パラメータ値セットがそれらのどこかに隠れている状態で、互いにクラスターを形成します。最尤推定値と隣接するパラメータセットの尤度の対数の差は、座標がパラメータθ 1 ... θ pであるプロット上に信頼領域を描くために使用できます。この領域は最尤推定値を囲み、その領域内のすべての点(パラメータセット)は、対数尤度においてある固定値のみ異なります。ウィルクスの定理によって与えられるχ 2分布は、この領域の対数尤度の差を、母集団の「真の」パラメータセットがその領域内に存在する「信頼度」に変換します。固定対数尤度の差を選択する際のコツは、信頼度を許容できるほど高くしつつ、領域を許容できるほど小さく(推定値の範囲を狭く)することです。
観測データが増えるにつれて、それらを個別の推定値を算出するために用いるのではなく、以前のサンプルと組み合わせて単一の結合サンプルを作成し、その大きなサンプルを用いて新たな最尤推定値を算出することができます。結合サンプルのサイズが大きくなるにつれて、同じ信頼度を持つ尤度領域のサイズは縮小します。最終的には、信頼領域のサイズがほぼ一点になるか、母集団全体がサンプリングされるかのいずれかになります。どちらの場合も、推定されたパラメータセットは母集団のパラメータセットと本質的に同じになります。