正在学习
6.11.1 Schwarz Information Criterion
6.11.1 Schwarz Information Criterion
The Schwarz Bayesian information criterion (BIC) ranks models by their posterior probabilities. To make this idea operational, we require model priors, , that sum to 1 across all models: . We also need priors over the parameters of each model, . With these in place, the posterior probability of each model, , is given by
where is the likelihood for model and is the marginal density for the data, Z. For notational convenience we use to denote the entire data set used to construct the forecasting model, . To obtain the model posterior given the data, must be integrated out. To do this we employ an approximation that allows the normal prior to be used as an approximate conjugate for any distribution.
Using a two-term expansion for the likelihood around the maximum likelihood estimate , we have
where is the average Fisher information evaluated at the sample estimate scaled by the sample size T, which is an term. Exponentiating (6.46) and integrating over we get
This approximate result allows the normal distribution to become a conjugate. Setting the prior over the coefficients to and evaluating the integral in (6.47) yields
Using this result in the expression for the posterior probability in (6.45), we have
Taking logarithms in (6.48),
The logarithm of the marginal distribution of Z is the same for each model and so can be ignored in model rankings. Maximizing (6.49) is therefore the same as minimizing
As the sample size, T , gets large, the first and last terms in (6.50) are bounded and so remain “small” and can be ignored, while the second and third terms get larger.
This suggests choosing the model that minimizes the following criterion:
or, as is more common in practice, minimizing (6.51) divided by the sample size,
Choosing the model that minimizes the expression in (6.51) or (6.52) is thus equivalent, in large samples, to selecting the model with the highest posterior probability.
练习题
What is the primary purpose of the Schwarz Bayesian Information Criterion (BIC)?
What condition must the model priors, , satisfy?
What does the expression represent?
What is the role of in the posterior probability formula?
In the two-term expansion for the likelihood, what does represent?
The average Fisher information evaluated at the sample estimate scaled by the sample size is an term.
The integral approximation for the normal prior involves integrating over without exponentiating the likelihood.
The integral result with a normal prior is approximately , where is the ___.
The posterior probability expression involves the term , where is the ___.
Explain the significance of taking logarithms in the posterior probability expression.
Which of the following are key components in the BIC formula?
Which of the following statements are true about the BIC and posterior probability equivalence?
Select the correct statements regarding the large sample approximation for model selection.
Describe the relationship between the BIC and the AIC in the context of model selection.
How does the BIC help in addressing the issue of overfitting in model selection?
Which of the following is a key requirement for the model priors in the Schwarz Bayesian Information Criterion (BIC) to be operational?
Which of the following statements are true regarding the integral approximation for the normal prior in the BIC?
The BIC formula is derived by minimizing the expression and then dividing by the sample size .
In the context of the BIC, the posterior probability expression is approximated by , which simplifies to by factoring out ___.
登录后解锁笔记、知识点解析、AI 问答
立即登录