正在学习
18.3 TESTS BASED ON THE PROBABILITY INTEGRAL TRANSFORM
18.3 TESTS BASED ON THE PROBABILITY INTEGRAL TRANSFORM
The probability integral transform2 (PIT) of a continuous cumulative density function evaluated at some outcome, , is defined as . Given a density forecast and an outcome , we can compute as the probability of observing a value less than or equal to . If the density forecast comes from a parametric model, this number can be calculated directly. If the density estimate is provided as a set of simulated pairs , for a range of y-values, as we discuss in chapter 13, the PIT value can
be estimated as
where is the difference in the values of .
The PIT has the useful property that if is truly distributed as , then U is uniformly distributed on [0, 1]. To see this, notice that and is the uniform density. Since has support on [0,1], it follows that U ∼ Uniform[0, 1]. The key assumption here is that the density used to compute the PIT is identical to that of the actual outcome.
Moreover, provided that the sequence of conditional densities is correctly specified each period, a sequence of PIT values, , generated from a sample of density forecasts, , and outcomes, will also be Uniform [0, 1] and mutually independent: For this to hold requires that the conditional density is correctly specified at each point in time; misspecified densities lead to violations of this property. For example, if the densities do not condition correctly on prior information, values of can be serially correlated.
Example 18.3.1 (PIT score for GARCH(1, 1) process). Suppose that the true density model is a GARCH(1, 1),
If a forecaster uses a homoskedastic density model of the form
then the PIT becomes
where is the standard Gaussian c.d.f. Specifically, at times when is higher than the average volatility, there is a higher than predicted chance of observing large forecast errors of either sign, and so this misspecification will show up in the form of higher probabilities of very large or very small PIT values at such times.
These results suggest a simple way of testing whether the density forecasts are correctly specified. Given a sequence of PIT values, , we can examine whether they are drawn from a uniform distribution.
As an empirical illustration, we use our daily data on S&P500 stock returns and generate one-step-ahead out-of-sample density forecasts from a GARCH(1,1) model with standard normal innovations and computed PIT values. A plot of these values for the last year of the sample preceding 03/31/2015, along with their squares, are shown in figure 18.1. There is no noticeable pattern in these.
PIT values
Squared PIT values
Figure 18.1: Time-series plot of probability integral transform (PIT) values in 2010 from a GARCH(1, 1) model fitted to daily US stock market returns. The top window plots the PIT values, while the bottom window shows the squared PIT values from one-step-ahead density forecasts.
Figure 18.2 shows a histogram of the PIT values using 20 bins, each of which should cover 5%. There is a slight overrepresentation of values near 0.5, suggesting that the GARCH(1,1) model misses the many days with very small return outcomes. Conversely, the model seems to perform reasonably well in the tails of the return distribution, perhaps with a slight tendency to overestimate the probability of large positive outcomes. Of course, the histogram shown here does not reveal what happens in the extreme tails.
To more formally test the hypothesis of correct specification, we can use a Kolmogorov–Smirnov or Cramer–von Mises test. For the Kolmogorov–Smirnov (KS) test, we construct the cumulative density function for U as as a function of u. The cumulative density function of a uniform variable is simply u on [0, 1], so the KS test for this problem is
When the predictive distribution is known, the limit distribution for the KS test statistic in (18.15) is , where is a standard

Figure 18.2: Histogram of one-step-ahead probability integral transform (PIT) values from a GARCH(1, 1) model with recursively estimated parameter values.
Brownian motion process. When parameters of the forecast distribution are estimated, the asymptotic distribution of the KS statistic must account for parameter estimation errors, leading to different critical values; see Durbin (1973). Bai (2003) shows that when the density estimate comes from a distribution known up to a finite set of parameters that must be estimated, the limiting distribution of the Kolmogorov test statistic depends on the true parameters and the true underlying distribution which are of course unknown. This makes it difficult to approximate percentiles of the limit distribution. Bai’s analysis is in-sample, whereas Corradi and Swanson (2006c) address the effect of parameter estimation error when such tests are applied to out-of-sample forecasts.
For the Cramer–von Mises test statistic we compute
The integral in (18.16) can be approximated numerically by using a fine mesh for . Alternatively, given a sample of density forecasts, the expression
TABLE 18.2:
Tests of the i.i.d. and zero-mean property of standardized forecast errors from a GARCH(1,1) model with Gaussian shocks estimated on daily US stock market returns.
| Equation, null hypothesis | Test | p-value | ||
| 0.0350 | 1 | 1.3238 | 0.1858 | |
| 0.0165 | 0.0210 | 1.6934 | 0.1934 | |
| 0.9494 | -0.0252 | -1.1104 | 0.2670 |
can be employed. The limit distribution for the CvM statistic is , which is the integral of a Brownian bridge. The critical value for a test with a size of 5% is approximately 0.461, with larger values leading to rejections.
Berkowitz (2001) suggests transforming the PIT values through the inverse of the standard Gaussian . Defining
and noting that . Uniform(0, 1), it follows that ind N(0, 1). This suggests simple likelihood ratio tests for correct specification of the distribution forecast. For example, one can run a simple unconstrained regression of on a constant and its lagged value,
and use conventional likelihood ratio tests to see whether the mean and variance of are 0 and 1, respectively:
where denotes the likelihood function in (18.19); should also be independent of its lagged values which suggests using the test statistic
Even if the density model passes these tests, it is possible that the distribution forecasts ignore time-varying volatility. This can be detected by considering persistence in the squared values of :
where, under the null of a correctly specified density, and . These regression tests are convenient and easy to conduct, but do not, of course, change the size or power of the original tests for uniformity.
As an empirical illustration of these tests, table 18.2 applies simple regression tests to GARCH(1,1) forecasts of daily US stock returns. We use 2000–2009 as an initial estimation period and 2010–2015 as the forecast evaluation period and assume a constant daily mean return. Tests of zero mean and no serial correlation in are not rejected, nor does there appear to be evidence of serial dependence in the squared value \widehat { u } _ { t + 1 } ^ { 2 } . ^
Constructing a correctly specified density forecast may be too tall an order; after all, correctly specifying the conditional mean is a hard problem. We might therefore expect that tests based on the assumption of a correct density forecast would mostly reject. However, tests such as (18.15) and (18.16) are notorious for having low power unless the sample size is very large, so a failure to reject the null hypothesis may not be very informative. Another problem, detailed in examples by Hamill (2001) and Gneiting and Raftery (2007) is that it is entirely possible that forecast distributions that are not particularly informative about the outcome can look reasonable in tests based on the PIT. We illustrate this by returning to the example from Gneiting and Raftery (2007).
Example 18.3.2 (Comparing Gaussian density forecasts, continued). Consider the two density forecasts from Example 18.2.1, and Both the density forecasts and yield probability integral transforms that are uniformly distributed, despite the second forecast clearly being a better one. Unconditionally, , so the first density forecast, is correctly specified and results in a uniform distribution for . However, conditional on observing μt, , so is also correctly specified and is also uniformly distributed.
These examples show that methods based on the PIT may not be able to distinguish between density forecasts that are correctly specified, but use different conditioning variables. Passing the test of uniformity does not imply that conditioning information has been utilized well, or even at all. For this reason, Hamill (2001) regards the PIT uniformity property as a necessary but not sufficient condition for a good forecast density.
练习题
What is the key property of the probability integral transform (PIT) when the density used to compute it is identical to that of the actual outcome?
In the context of the probability integral transform, what does it mean for a sequence of PIT values to be mutually independent?
Which of the following are conditions for a sequence of PIT values to be i.i.d. Uniform[0, 1]?
What are the implications of using a misspecified density model in the probability integral transform, as shown in the GARCH(1, 1) example?
If the density forecast comes from a parametric model, the PIT value can always be calculated directly.
The probability integral transform is only applicable to continuous cumulative density functions.
The formula for estimating the PIT value when the density estimate is provided as a set of simulated pairs is , where is the difference in the values of .
In the GARCH(1, 1) example, if a forecaster uses a homoskedastic density model, the PIT becomes , where is the standard Gaussian .
Explain how a misspecified density model can lead to serial correlation in PIT values.
What is the significance of the PIT values being uniformly distributed on [0, 1] for testing density forecasts?
Which of the following is a consequence of using a homoskedastic density model for a GARCH(1, 1) process in PIT calculation?
Combining the concepts of probability integral transform and calibration (from prior knowledge), which of the following statements are correct?
登录后解锁笔记、知识点解析、AI 问答
立即登录