正在学习
21.3 REAL-TIME DATA
21.3 REAL-TIME DATA
Real-time forecasting refers to the computation of forecasts by means of data restricted to be available at the point in time when the forecast is constructed. This may seem like an obvious requirement to impose and is automatically satisfied in the case of survey forecasts. Often, however, forecasts are constructed “after the fact,” using simulated or pseudo out-of-sample methods—or backtests as they are called in finance—and it is for these exercises that the data availability restriction becomes relevant. Suppose, for example, that we find that we could have predicted some event, say the sharp decline in stock markets during the fall of 2008, by means of data that were not available at the time such as detailed information on banks’ balance sheets or counterparty risk exposures, and the delinquency rates on mortgages. Clearly, we would not conclude from this finding that, historically, markets were not efficient. At most we could conclude that, had such data been available and used to construct a forecasting model, it might have helped investors and other decision makers predict the ensuing events.
Close attention therefore has to be paid to which data are available to forecasters in real time. Many types of macroeconomic data undergo substantial revisions. This leads to the concept of data vintages, i.e., snapshots of the data series available at a given point in time, v. For example, at the monthly frequency the 2012:12 data vintage might comprise all data available at the end of December 2012. This includes data on earlier observations known at that point in time. Data vintages can be represented as an expanding triangular shape, as shown in table 21.1.
Data revisions come in different shapes. First, there are the regularly scheduled revisions for variables such as GDP. For example, in the US, an advance estimate of GDP in a given quarter is published by the Bureau of Economic Analysis near the end of the month following the end of the quarter. The following year, data from income tax records and economic census data allow this estimate to be updated to include more precise data. Second, benchmark revisions may reflect changes in how a variable is being measured. A prime example here is the change from measuring GDP using fixed weighting as opposed to chain weighting, which was introduced in 1996 for US GDP.
Let t be the date of the forecast and let refer to date-t observations from vintage v. Only the current data vintage—data available at time t—can be used to generate forecasts. Moreover, let be the parameter of the forecasting model associated with vintage v. Taking explicit account of data vintages, we can write the forecast for period t + h given information at time t on the latest vintage v as
Here v refers to the vintage of data used for the model, is the data vintage available at time t, and are the parameters associated with this model.
As pointed out by Croushore (2006) and Stark and Croushore (2002), data revisions can affect the forecasting model, and the conditional forecast, , in three ways. First, data revisions directly affect the forecast through the conditioning information, . To illustrate this point, suppose we compare two forecasts from identical linear models, except that one forecast uses vintage , while the other uses vintage . Then the two forecasts would be and , respectively. Second, the use of different data will almost certainly lead to different parameter estimates, i.e., , and so . Finally, the forecast model could itself change with the data vintage—a point that applies both to the choice of lag order for linear models and the inclusion of nonlinear terms for nonlinear models.
These considerations appear to matter empirically. In a study of predictability of money demand, Amato and Swanson (2001) find that the latest available data on the aggregate stock of money, as measured by M1 and M2, appear to predict output growth, while real-time measures of the money stock do not have such predictive power. This suggests that measures of growth in money should perhaps be omitted from real-time forecasting models. Swanson and White (1997) examine the effect of real-time data revisions on model selection.
ABLE 21.1 Data vintages for quarterly US real GNP.
| DATE | ROUTPUT | ROUTPUT | ROUTPUT | ROUTPUT | ROUTPUT | ROUTPUT | ROUTPUT | ROUTPUT |
| 12Q2 | 12Q3 | 12Q4 | 13Q1 | 13Q2 | 13Q3 | 13Q4 | 14Q1 | |
| 2010:Q1 | 12937.7 | 12947.6 | 12947.6 | 12947.6 | 12947.6 | 14597.7 | 14597.7 | 14597.7 |
| 2010:Q2 | 13058.5 | 13019.6 | 13019.6 | 13019.6 | 13019.6 | 14738 | 14738 | 14738 |
| 2010:Q3 | 13139.6 | 13103.5 | 13103.5 | 13103.5 | 13103.5 | 14839.3 | 14839.3 | 14839.3 |
| 2010:Q4 | 13216.1 | 13181.2 | 13181.2 | 13181.2 | 13181.2 | 14942.4 | 14942.4 | 14942.4 |
| 2011:Q1 | 13227.9 | 13183.8 | 13183.8 | 13183.8 | 13183.8 | 14894 | 14894 | 14894 |
| 2011:Q2 | 13271.8 | 13264.7 | 13264.7 | 13264.7 | 13264.7 | 15011.3 | 15011.3 | 15011.3 |
| 2011:Q3 | 13331.6 | 13306.9 | 13306.9 | 13306.9 | 13306.9 | 15062.1 | 15062.1 | 15062.1 |
| 2011:Q4 | 13429 | 13441 | 13441 | 13441 | 13441 | 15242.1 | 15242.1 | 15242.1 |
| 2012:Q1 | 13502.4 | 13506.4 | 13506.4 | 13506.4 | 13506.4 | 15381.6 | 15381.6 | 15381.6 |
| 2012:Q2 | -999 | 13558 | 13548.5 | 13548.5 | 13548.5 | 15427.7 | 15427.7 | 15427.7 |
| 2012:Q3 | -999 | -999 | 13616.2 | 13652.5 | 13652.5 | 15534 | 15534 | 15534 |
| 2012:Q4 | -999 | -999 | -999 | 13647.6 | 13665.4 | 15539.6 | 15539.6 | 15539.6 |
| 2013:Q1 | -999 | -999 | -999 | -999 | 13750.1 | 15583.9 | 15583.9 | 15583.9 |
| 2013:Q2 | -999 | -999 | -999 | -999 | -999 | 15648.7 | 15679.7 | 15679.7 |
| 2013:Q3 | -999 | -999 | -999 | -999 | -999 | -999 | 15790.1 | 15839.3 |
| 2013:Q4 | -999 | -999 | -999 | -999 | -999 | -999 | -999 | 15965.6 |
Dependent on what is being modeled, data revisions can also lead to changes in the predicted variable. If the predicted variable is subject to revisions, what is considered the “actual” or outcome is no longer so clear-cut. Are forecasters trying to predict the first measurement of a variable because this is what will generate most publicity, or are they trying to predict some underlying, unobserved, “true” measure of the outcome which, over time, is measured with increasing precision? This important issue is often difficult to address.
Stark and Croushore (2002) consider using either the latest available data vintage, the most recent vintage prior to a benchmark revision, or the vintage obtained one year after the observation date. If information is improving over time, using the latest available vintage would seem to make most sense. In this case, data on older observation dates have had more time to “settle” and so could introduce heterogeneity in the quality of the data, with more recent data being subject to the greatest measurement errors. For GDP growth, Stark and Croushore (2002) find that forecasts are not notably improved if based on the latest available data. Conversely, for inflation forecasts, models based on the latest available data seem to lead to better forecasts than models based on the real-time data vintages.
练习题
What is the main characteristic of real - time forecasting?
What is the concept of data vintages related to?
In the US, when is an advance estimate of GDP in a given quarter published?
What are the types of data revisions? (Select all that apply)
Which of the following are true about the parameters in the real - time forecasting formula f _ { t + h | t v } = f _ { v } ( z _ { t v }, eta _ { t v } )? (Select all that apply)
Only the current data vintage can be used to generate forecasts in real - time forecasting.
Benchmark revisions are always due to changes in the measurement of a variable.
Data vintages can be represented as an expanding ___ shape.
In the US, data from income tax records and economic census data allow the advance estimate of GDP to be updated to include more precise data in the following ___.
Explain why we cannot conclude that markets were not efficient if we could have predicted an event using data that was not available at the time.
What is the role of in the formula f _ { t + h | t v } = f _ { v } ( z _ { t v}, eta _ { t v } )?
Which of the following are examples of data that may undergo revisions? (Select all that apply)
Regularly scheduled revisions for GDP are made immediately after the quarter ends.
The 2012:12 data vintage comprises all data available at the end of ___.
Describe the difference between regularly scheduled revisions and benchmark revisions of data.
If we want to use the real - time forecasting formula f _ { t + h | t v } = f _ { v } ( z _ { t v}, eta _ { t v } ) to make a forecast for period given information at time , what does represent?
Which of the following are true about the data used in real - time forecasting? (Select all that apply)
The change from fixed - weighting to chain - weighting for GDP measurement is an example of a regularly scheduled revision.
The formula f _ { t + h | t v } = f _ { v } ( z _ { t v}, eta _ { t v } ) shows the forecast for period given information at time on the latest vintage ___.
Explain the significance of data vintages in real - time forecasting.
登录后解锁笔记、知识点解析、AI 问答
立即登录