正在学习

The Parametric Forecasting Problem

The Parametric Forecasting Problem

The forecaster’s objective is to use data—outcomes of the random variable to predict the value of the random variable Y. Let be the date when the forecast is computed and define as the sigma algebra generated by so is a filtration . Following conventional terminology in the forecasting literature, we refer to as the information set at time T. We define as when the timing of the forecasting problem is clear and as when we need to be clearer about the time subscript, and write objects as conditional on or conditional on as meaning the same thing. By we mean or . The random variable typically includes past values of the predicted variable, as well as other possibly useful variables, so often . In some cases (such as ARIMA and VAR forecasts, which we discuss in more detail in chapters 7 and 9), the information set might include only past values of the predicted variable(s). The outcome, Y, may be a vector or could be univariate.

We assume throughout the analysis that a joint distribution for exists, which can be written as , where the vector comprises all the parameters of the joint distribution. Through a small abuse of notation we refer to the true parameters as θ regardless of whether we are referring to the conditional distribution for Y given or to the joint distribution. We refer to this joint distribution as the data-generating process or true model.

This chapter focuses on the construction of optimal point forecasts, i.e., finding the best possible for a general forecasting problem. There are many reasons for considering the optimality of a forecast. First, clearly we want the best possible forecast. Second, questions related to forecast optimality can be complicated and there will not be a single best forecasting approach even in the simplest of problems. Selecting a good forecasting approach often requires deeper thinking about the problem than merely stating a loss function and it need not be straightforward to determine a reasonable set of models for a realistic forecast problem. This point motivates the wide range of approaches considered in the empirical forecasting literature (and in part II of this book), with different approaches being tailored to different data-generating processes and forecast objectives. Claims of forecast optimality are typically limited to restrictive classes of models and so can be very weak—optimality often does not extend too far in the sense that there may be better forecast models available, given the same data.

We use “risk” as our measure of the performance of a forecasting method, as well as the best measure for distinguishing between methods. Using jargon from the statistics literature, the risk of a forecasting method is simply the expected loss when integrating over both Y and Z. This focus on risk is unlikely to be controversial. Many forecasting papers motivate their methods through Monte Carlo experiments that examine risk at a few, carefully chosen points in the space of the parameters and data-generating processes. Such risk evaluation exercises require us to think carefully about the model spaces and to avoid assuming that methods that work for some Monte Carlo studies will necessarily work more generally, especially for quite different data-generating processes. It also highlights the role of the loss function; results favoring a particular forecasting model for one loss function need not generalize to other loss functions.

Examining in a single chapter the forecasting problem in the context of decision theory requires us to cut many corners. Hence, this chapter is really an overview of some key points from decision theory and estimation theory adapted to the forecasting problem and using language from forecasting. We demonstrate results not by proofs with technical conditions but through relevant examples.

Forecast models may be parametric, semiparametric or nonparametric. Parametric models for the forecast are fully specified up to a finite-dimensional unknown vector, β. Nonparametric models typically have an infinite-dimensional set of unknown parameters. Semiparametric models fall in the middle. Throughout this chapter we assume that the true joint distribution of the data, or data-generating process, has a parametric representation with a joint density for the outcome variable and the conditioning information used by the forecaster. This is the setting that the title of the chapter refers to.

An alternative to examining point forecasts is to provide a predictive density. The construction of a predictive density arises naturally under the Bayesian approach. From the classical perspective, the estimation problem will be different from that of point forecasting. We will see below that point forecasts are features of the predictive distribution in the sense that knowing the predictive distribution allows the construction of the optimal point forecast. Hence, generating a predictive distribution is in a sense a harder problem since the estimation problem is to construct something more informative than the point estimate. For this reason it is often suggested that provision of a predictive distribution (known as density forecasting in the forecasting literature) is better than provision of a point forecast. However, when the predictive distribution must be estimated, additional concerns arise.

The chapter restricts our focus to forecasts with loss functions of the form L ( f, Y), i.e., we drop the dependence of L on Z, except through f . This reduces the notation and, as we saw in the previous chapter, most commonly used loss functions simplify in this manner.

The chapter proceeds as follows. Section 3.1 examines optimality of point forecasts under known model parameters. Section 3.2 considers classical (frequentist) approaches when parameters are estimated, again focusing on optimality issues. This is followed in section 3.3 by an examination of Bayesian methods for constructing point forecasts. Section 3.4 relates the two approaches and discusses density forecasts.

Section 3.5 compares the methods in the context of an empirical portfolio decision problem. Section 3.6 concludes.

练习题

What is the forecaster's primary objective?

A. To analyze the sigma algebra
B. To use data from the random variable to predict the value of
C. To define the filtration
D. To construct the joint distribution

What is defined as?

A. The sigma algebra generated by
B. The sigma algebra generated by
C. The sigma algebra generated by
D. The sigma algebra generated by

What does the notation represent? (Select all that apply)

A.
B.
C.
D.

The random variable includes only past values of the predicted variable .

In ARIMA and VAR forecasts, the information set might include only past values of the predicted variable(s).

The outcome may be a ___ or could be univariate.

We assume a joint distribution for exists, which can be written as , where comprises all the parameters of the ___.

What is referred to as the data-generating process or true model?

Which of the following are reasons for considering forecast optimality? (Select all that apply)

A. We want the best possible forecast.
B. There will be a single best forecasting approach even in the simplest of problems.
C. Selecting a good forecasting approach requires deeper thinking about the problem.
D. It is straightforward to determine a reasonable set of models for a realistic forecast problem.

What is a limitation of claims of forecast optimality?

A. They are always valid for any data-generating process.
B. They are typically limited to restrictive classes of models and may not extend too far.
C. They guarantee the best forecast for any given data set.
D. They are not influenced by the choice of loss function.

Which of the following are types of forecast models? (Select all that apply)

A. Parametric
B. Semiparametric
C. Nonparametric
D. Deterministic

What is the relationship between point forecasts and predictive distribution?

What is often suggested as better than provision of a point forecast?

A. Provision of a parametric model
B. Provision of a predictive distribution
C. Provision of a nonparametric model
D. Provision of a loss function

Which of the following are true about the construction of a predictive density? (Select all that apply)

A. It arises naturally under the Bayesian approach.
B. It is the same as the estimation problem of point forecasting from the classical perspective.
C. From the classical perspective, the estimation problem is different from that of point forecasting.
D. It is not influenced by the choice of the loss function.

What is the primary measure of the performance of a forecasting method?

A. Accuracy
B. Precision
C. Risk
D. Bias

Which of the following are considered when evaluating risk in forecasting papers? (Select all that apply)

A. Monte Carlo experiments
B. The model spaces
C. The choice of loss function
D. The forecasting method's computational efficiency

What is the main focus of this chapter in the context of decision theory?

Which of the following are assumptions made in this chapter about the data-generating process? (Select all that apply)

A. It has a parametric representation.
B. It has a joint density for the outcome variable and conditioning information.
C. It is nonparametric.
D. It is semiparametric.

What is an alternative to examining point forecasts?

A. Examining the loss function
B. Providing a predictive density
C. Using Monte Carlo experiments
D. Constructing a parametric model

Which of the following combinations correctly describe the relationship between point forecasts and predictive distribution? (Select all that apply)

A. Point forecasts are independent of the predictive distribution.
B. Knowing the predictive distribution allows the construction of the optimal point forecast.
C. Point forecasts are features of the predictive distribution.
D. The predictive distribution is derived from point forecasts.

In the context of forecasting excess returns on stocks, if the data-generating process is given by , where , and , which of the following represents the information set at time ?

A.
B.
C.
D.

Which of the following statements are true regarding the construction of optimal point forecasts and the data-generating process? Select all that apply.

A. Optimal point forecasts aim to find the best possible for a general forecasting problem.
B. The data-generating process for excess returns on stocks assumes that .
C. The joint distribution assumption is necessary for constructing optimal point forecasts as it provides the basis for calculating conditional expectations.
D. In the context of forecasting excess returns, the optimal stock holding for an informed investor depends on the data-generating process parameters.

The risk of a forecasting method, which is used as a performance measure, is simply the expected loss when integrating over both and . In the context of forecasting excess returns on stocks, if an investor uses a forecasting method with a higher expected loss, it implies that the investor's expected utility at optimal stock holdings will be lower.

In the parametric forecasting problem, the random variable typically includes past values of the predicted variable and other possibly useful variables . In the data-generating process for excess returns on stocks , the variable can be considered as a type of ___ variable.

登录后解锁笔记、知识点解析、AI 问答

立即登录