正在学习

4.2.2 Estimation Based on Functions Other Than the Loss Function

4.2.2 Estimation Based on Functions Other Than the Loss Function

An alternative to loss-based estimation is to use plug-in estimators derived from a variety of popular estimation methods. The leading case is perhaps to use maximum likelihood methods to estimate the model parameters and then plug these into the forecast model. Alternatively, often estimates obtained from least squares regression are used in forecast problems involving very different loss functions.

Maximum likelihood estimates for or are commonly used in forecasting. Maximum likelihood methods typically require the construction of a likelihood for the data, although limited information maximum likelihood methods might be employed when models are complicated or it is thought that some of the predictor variables can be treated as exogenous. We can then construct forecasts from the model , using that is typically easily derived from and often is just a subset of the latter.

Example 4.2.2 (Forecast of a binary outcome variable). Suppose Y is a binary variable that takes values of either 0 or 1 so the optimal forecast is a function of the “success” probability, . To generate a forecast of Y we might use a linear index model, for some , F . Popular choices are probit and logit models, the maximum likelihood estimates for which can be employed in a decision rule such as , which forecasts if the predicted probability falls above a certain threshold, c, and otherwise forecasts

Using maximum likelihood estimates to construct a forecast model requires some knowledge of the form of and . When the likelihood is correctly specified so the entire joint density of and posited by the forecaster is the true data-generating process, under fairly general conditions the maximum likelihood estimators for the parameters will be consistent for the true parameters of the data-generating process, θ. Assuming that the random variables from which the data are drawn have absolutely continuous distributions, and that the expected likelihood exists and is continuous in the consistency result follows provided that (i) the model is identified (so the expected likelihood has a unique maximum at the true θ for the data-generating process) and (ii) the suitably scaled log-likelihood

satisfies a uniform law of large numbers. This result follows from White (1996, Theorem 4.2).

In practice, the assumption that the likelihood is correctly specified is unlikely to hold and we can instead view the estimator as a quasi-maximum likelihood estimator (QMLE). QMLEs may still prove to have good properties, as discussed in White (1996), although their properties must be established in relation to the forecaster’s loss function. When the kernel of the likelihood function agrees with the forecaster’s loss function, the results from above apply.

Example 4.2.3 (Quasi-maximum likelihood estimation with a normal distribution). The pseudo MLE from a model that assumes normality can be equivalent to a forecast that minimizes mean squared errors. To see this, let , so

and

Here the estimates of the forecasting model and so forecasts based on the MLE values may be reasonable even when the density is misspecified.

Example 4.2.4 (QMLE for tick-exponential family). Example 4.1.2 showed that the M-estimator for the relevant quantile provided an estimator for the forecast model under lin-lin loss. Komunjer (2005) examines QMLEs for the tick-exponential family whose pseudo densities take the form

where η is the model for the conditional quantile. When the tick-exponential is an asymmetric Laplace (double exponential) density, the tick-exponential QMLEs coincide with the conventional quantile regression estimator. Setting and yields a likelihood that is an order-preserving transform of the M-estimator based on the log function. It follows from Komunjer’s results that the QMLE provides the best forecast among linear models.

Equivalence results such as those in the previous examples do not hold more generally. Rather, the forecasts based on MLEs will differ from those that minimize the loss function. Under reasonable conditions forecasts based on the correct loss function have the property that—at least in large enough samples—they yield losses that are approximately the smallest possible ones given the loss function and the forecasting model. In general we therefore expect forecasts based on the correct loss function to be superior to those based on MLEs. In small samples, efficiency may be more of an issue and it is possible that utility- or loss-based estimators are more sensitive to outliers than MLE-based estimators and so could lead to worse finitesample performance. This issue has to be evaluated in each individual case.

Another case where the parameters are estimated under objective functions other than the forecaster’s loss arises when standard econometric methods such as least squares estimation are used even though the forecaster has a different loss function such as mean absolute error loss or perhaps is interested in predicting turning points of the data. Indeed, the large literature on turning point prediction often examines forecasts in the context of logit or probit models; see, e.g., Chauvet and Potter (2005). It follows directly from the fact that differs for different loss functions (see the examples in chapter 3) that this approach can easily lead to forecast models that do not achieve average losses that converge to the average expected loss given . For further discussion and examples, see Weiss (1996).

Example 4.2.5 (Least squares estimates under absolute error loss). Suppose least squares estimates are employed to generate forecasts under an absolute error loss function

where . Equality in (4.9) holds when the conditional mean and conditional median are equivalent, , when the conditional distribution of Y given X is symmetric. Equality of the expected loss need not hold, however, for skewed or asymmetrical distributions.

练习题

Which of the following is a common alternative to loss-based estimation in forecasting?

A. Using only historical averages
B. Using plug-in estimators derived from estimation methods like maximum likelihood
C. Using only judgmental forecasts
D. Using only exponential smoothing methods

In forecasting, what are maximum likelihood estimates for or typically used for?

A. To construct the loss function directly
B. To estimate the model parameters and then plug these into the forecast model
C. To minimize the forecast errors directly
D. To determine the optimal sample size for forecasting

What are the conditions for the consistency of maximum likelihood estimators? (Select all that apply)

A. The model is identified
B. The expected likelihood exists and is continuous in
C. The random variables have discrete distributions
D. The suitably scaled log-likelihood satisfies a uniform law of large numbers

Quasi-maximum likelihood estimators (QMLEs) always require the likelihood to be correctly specified.

The pseudo MLE from a model that assumes normality can be equivalent to a forecast that minimizes mean squared errors.

The decision rule for forecasting a binary outcome variable using maximum likelihood estimates often involves comparing to a certain ___.

When the kernel of the likelihood function agrees with the forecaster’s loss function, the results for the consistency of maximum likelihood estimators ___.

Explain why quasi-maximum likelihood estimators (QMLEs) may still prove to have good properties even when the likelihood is misspecified.

What is the main concern with using plug-in estimators in forecasting?

Which of the following are true about the pseudo MLE from a model that assumes normality? (Select all that apply)

A. It minimizes the sum of absolute errors
B. It can be equivalent to a forecast that minimizes mean squared errors
C. It assumes the data follows a Poisson distribution
D. It may still provide reasonable forecasts even when the density is misspecified

When using maximum likelihood estimates to construct a forecast model, which of the following is NOT a required condition for the consistency of the maximum likelihood estimators?

A. The model is identified so the expected likelihood has a unique maximum at the true for the data-generating process.
B. The suitably scaled log-likelihood satisfies a uniform law of large numbers.
C. The random variables from which the data are drawn have discrete distributions.
D. The expected likelihood exists and is continuous in .

Which of the following statements are true regarding plug-in estimators in forecasting? Select all that apply.

A. Plug-in estimators can be based on maximum likelihood estimates.
B. The loss function used to estimate model parameters always aligns with the forecaster's loss function when using plug-in estimators.
C. Quasi-maximum likelihood estimators (QMLEs) may still have good properties even when the likelihood is misspecified.
D. Plug-in estimators are only used when the form of the joint density of and is known.

When the kernel of the likelihood function agrees with the forecaster's loss function, the results regarding the properties of ___ estimators apply.

登录后解锁笔记、知识点解析、AI 问答

立即登录