正在学习

6.11.1 Schwarz Information Criterion

6.11.1 Schwarz Information Criterion

The Schwarz Bayesian information criterion (BIC) ranks models by their posterior probabilities. To make this idea operational, we require model priors, , that sum to 1 across all models: . We also need priors over the parameters of each model, . With these in place, the posterior probability of each model, , is given by

where is the likelihood for model and is the marginal density for the data, Z. For notational convenience we use to denote the entire data set used to construct the forecasting model, . To obtain the model posterior given the data, must be integrated out. To do this we employ an approximation that allows the normal prior to be used as an approximate conjugate for any distribution.

Using a two-term expansion for the likelihood around the maximum likelihood estimate , we have

where is the average Fisher information evaluated at the sample estimate scaled by the sample size T, which is an term. Exponentiating (6.46) and integrating over we get

This approximate result allows the normal distribution to become a conjugate. Setting the prior over the coefficients to and evaluating the integral in (6.47) yields

Using this result in the expression for the posterior probability in (6.45), we have

Taking logarithms in (6.48),

The logarithm of the marginal distribution of Z is the same for each model and so can be ignored in model rankings. Maximizing (6.49) is therefore the same as minimizing

As the sample size, T , gets large, the first and last terms in (6.50) are bounded and so remain “small” and can be ignored, while the second and third terms get larger.

This suggests choosing the model that minimizes the following criterion:

or, as is more common in practice, minimizing (6.51) divided by the sample size,

Choosing the model that minimizes the expression in (6.51) or (6.52) is thus equivalent, in large samples, to selecting the model with the highest posterior probability.

练习题

What is the primary purpose of the Schwarz Bayesian Information Criterion (BIC)?

A. To maximize the likelihood function
B. To rank models by their posterior probabilities
C. To minimize the sum of squared residuals
D. To calculate the marginal density of the data

What condition must the model priors, , satisfy?

A. They must be greater than 1
B. They must be equal to the likelihood
C. They must sum to 1 across all models
D. They must be less than 0

What does the expression represent?

A. The sum of the likelihoods of all models
B. The sum of the posterior probabilities of all models
C. The normalization condition for model priors
D. The sum of the marginal densities of the data

What is the role of in the posterior probability formula?

A. It is the likelihood for model
B. It is the prior over the parameters of each model
C. It is the marginal density for the data,
D. It is the maximum likelihood estimate of

In the two-term expansion for the likelihood, what does represent?

A. The average Fisher information
B. The marginal density of the data
C. The maximum likelihood estimate of
D. The prior over the coefficients

The average Fisher information evaluated at the sample estimate scaled by the sample size is an term.

The integral approximation for the normal prior involves integrating over without exponentiating the likelihood.

The integral result with a normal prior is approximately , where is the ___.

The posterior probability expression involves the term , where is the ___.

Explain the significance of taking logarithms in the posterior probability expression.

Which of the following are key components in the BIC formula?

A. Likelihood function
B. Number of parameters,
C. Sample size,
D. Prior over the coefficients
E. Marginal density of the data

Which of the following statements are true about the BIC and posterior probability equivalence?

A. BIC minimization is equivalent to maximizing posterior probability in large samples
B. BIC is always the best criterion for model selection
C. The equivalence holds when the sample size, , is large
D. BIC and posterior probability are unrelated
E. Small sample sizes do not affect the equivalence

Select the correct statements regarding the large sample approximation for model selection.

A. The first and last terms in the criterion remain significant
B. The second and third terms grow larger with sample size
C. The first and last terms are bounded and can be ignored
D. The criterion suggests minimizing
E. Sample size has no effect on the terms

Describe the relationship between the BIC and the AIC in the context of model selection.

How does the BIC help in addressing the issue of overfitting in model selection?

Which of the following is a key requirement for the model priors in the Schwarz Bayesian Information Criterion (BIC) to be operational?

A. The priors over the parameters of each model, , must be zero.
B. The model priors, , must sum to 1 across all models.
C. The likelihood for each model, , must be zero.
D. The marginal density for the data, , must be infinite.

Which of the following statements are true regarding the integral approximation for the normal prior in the BIC?

A. It involves exponentiating the two-term expansion of the likelihood.
B. It requires integrating over .
C. It assumes the prior over the coefficients, , is zero.
D. It results in an expression involving the maximum likelihood estimate .
E. It is not necessary for calculating the BIC.

The BIC formula is derived by minimizing the expression and then dividing by the sample size .

In the context of the BIC, the posterior probability expression is approximated by , which simplifies to by factoring out ___.

登录后解锁笔记、知识点解析、AI 问答

立即登录