正在学习
17.7 IDENTIFYING SUPERIOR MODELS
17.7 IDENTIFYING SUPERIOR MODELS
The approach taken by White (2000) and Hansen (2005) asks whether a superior model exists. By using the max-statistic, both White and Hansen identify models that are significantly better than the benchmark, i.e., models whose individual t-statistics exceed the appropriate critical value. Their approach does not go beyond this first step in an attempt to identify which particular model is better than the benchmark, nor does it address how many benchmark-beating models exist. This issue is addressed by Romano and Wolf (2005) who develop a stepwise approach that iterates on White’s bootstrap in a way that controls the so-called familywise error rate, i.e., the event of wrongly identifying at least one forecasting model as superior, and can be used to identify additional significant alternatives by tightening the critical value in subsequent steps. Like White’s approach, their method achieves power gains relative to the Bonferroni bound by accounting for correlation among the individual models’ test statistics.
To see how the Basic StepM method of Romano and Wolf (2005) works, without loss of generality suppose that we have ordered the models from largest to smallest by their out-of-sample forecasting performance, . The first step in the procedure estimates the critical value, , associated with the null hypothesis that not even the best model is superior using the one-step White bootstrap conducted at the α% critical level. If this step does not result in a rejection of the null, there is no need to continue because none of the models beats the benchmark. Conversely, if the null is rejected, all models with performance are removed.12 Next, the previous step is repeated including only those models that remain in the candidate set to establish a new cutoff, for these models. The procedure continues until no additional models are removed. Romano and Wolf propose algorithms that asymptotically identify all models that are better than the benchmark, while removing inferior models and maintaining control over the familywise error rate.
Hansen, Lunde, and Nason (2011) avoid the need for having a random or fixed benchmark against which a multitude of models are compared. Their model confidence set identifies the subset of models that includes the best, but usually unknown, model with a certain level of confidence. The model confidence set is constructed given an equivalence test and an elimination rule. The equivalence test is a test for equal predictive performance of the models contained in some set of (superior) models. If the equivalence test rejects, the models in the maintained set are not equally good and the elimination rule identifies which of the models to eliminate.
Specifically, let denote the initial set of models under consideration. Each model, i , generates a loss in period t of , so the differential loss when comparing pairs of models i and at time t is
Hansen, Lunde, and Nason (2011) propose a simple algorithm for constructing the model confidence set. The first step sets , the full set of models under consideration. The next step uses the equivalence test to test for all at a critical level α. If is accepted, the estimated model confidence set is . If the null is rejected, the elimination rule is used to reject one model from and the procedure is repeated on the reduced set of models. The procedure continues until the equivalence test does not reject and no further model needs to be eliminated.
This description still leaves the issue of which test to use to eliminate forecast models and construct the model confidence set. A natural approach is to consider the models’ out-of-sample forecasting performance. To this end, Hansen, Lunde, and Nason (2011) define two relative sample loss measures, which captures the relative sample loss of models i and and which captures the relative sample loss of model i against the average loss across the m models in . Studentized test statistics can be used in the tests:
where and are sample estimates of the corresponding population objects. The first of the test statistics in (17.45) is similar to the t-test proposed by Diebold and Mariano (1995) for comparing the predictive accuracy of models i and j and testing the null that . To convert the many t-tests in (17.45) into simpler objects that allow inspection of the null that all models have equivalent forecast performance, Hansen, Lunde, and Nason (2011) propose the following test statistics for
These test statistics have a nonstandard asymptotic distribution whose critical values can be bootstrapped along the lines described in White (2000) and Hansen (2005). The natural elimination rule associated with denoted , is to set , so that in the case of rejection of , the model with the largest t-statistic against the average loss (and hence the worst performance) gets eliminated. For the test statistic, Hansen, Lunde, and Nason (2011) propose the elimination rule , thus eliminating the model whose performance looks worst relative to some other model.
练习题
Which approach uses the max-statistic to identify models significantly better than the benchmark?
What is the main goal of Romano and Wolf's (2005) stepwise approach?
Which of the following are true about Hansen, Lunde, and Nason's (2011) model confidence set approach? (Select all that apply)
The Romano and Wolf procedure continues until no additional models are removed.
The model confidence set constructed by Hansen, Lunde, and Nason (2011) includes models that are not equally good.
The differential loss when comparing pairs of models and at time is given by , where represents the loss of model in period given information up to period . The differential loss is used in the __________ test.
The __________ bound is a conservative approach that guards against worst-case scenarios by adjusting the p-value based on the number of models under consideration.
Explain the purpose of the elimination rule in Hansen, Lunde, and Nason's (2011) model confidence set approach.
What is the significance of the critical value in the Romano and Wolf procedure?
Which of the following statements are true about the impact of searching across many prediction models? (Select all that apply)
登录后解锁笔记、知识点解析、AI 问答
立即登录