正在学习
2.1.4 Loss Functions Not Based on Expected Loss
2.1.4 Loss Functions Not Based on Expected Loss
So far we have characterized the loss function for a univariate outcome, and defined its properties with reference to a “one-shot” problem. This makes sense when forecasting is placed in a decision-theoretic or utility-maximization context. This approach to forecasting is internally consistent, from initially setting up the problem to defining the expected loss and conducting model estimation and forecast evaluation.
Some loss functions that have been used in practice are based directly on sample statistics without relating the sample loss to a population loss function. In cases where such a population loss function exists and satisfies reasonable properties, this does not cause any problems. Basing the loss function directly on a sample of losses can, however, sometimes yield a loss function that does not make sense in population or for fully specified decision problems. Loss functions that do not map back to decision problems often have poor and unintended properties. We consider one such example below.
Example 2.1.2 (Kuipers score for binary outcome). Let be a forecast of the binary variable and let be the number of observations for which the forecast equals j and the outcome equals k. The Kuipers score is given by
This is the positive hit rate, i.e., the proportion of times where is correctly predicted less the “false positive rate,” i.e., the proportion of times where is wrongly predicted. This can equivalently be thought of as
which is the hit rate for plus the hit rate for minus a centering constant of 1. The Kuipers score is positive if the sum of the positive and negative hit rates exceeds 1. For a sample with a single observation, this definition makes no sense, as one of the denominators in (2.10) is 0: either
For a single observation, this sample statistic does not follow from any obvious loss function. The first term in is the sample analog of and the second is the sample analog of . However, they do not combine to a loss function with this sample analog. This failure to embed the loss function into the expected loss framework results in odd properties for the objective. For example, the definition of KuS in (2.9) implies that the marginal value of an extra , a correct call, depends on the sample proportion of hits. To see this, consider the improvement in KuS from adding a single successfully predicted observation
The resulting improvement in the hit rate is
Thus the marginal value of a correct call depends on the total number of observations and the proportion of missed hits prior to the new observation. The Kuipers score’s poor properties arise from the lack of justification of its setup for a population problem.
2.2 SPECIFIC LOSS FUNCTIONS
We next review various families of loss functions that have been suggested in the forecasting literature. The vast majority of empirical work on forecasting assumes that the loss function depends only on the forecast error, , i.e., the difference between the outcome and the forecast. In this case we can write In general, loss functions can be more complicated functions of the outcome and forecast and take the form or
2.2.1 Loss That Depends Only on Forecast Errors
The most commonly used loss functions, including squared error loss and absolute error loss, depend only on the forecast error. For such loss functions, L (e), so the loss function takes a particularly simple form.
2.2.1.1 Squared Error Loss
By far the most popular loss function in empirical studies is squared error loss, also known as quadratic or mean squared error (MSE) loss:
This loss function clearly satisfies the three Granger properties listed in ( 2.3). When viewed as a family of loss functions—corresponding to different values of the scalar a—squared error loss forms a homogeneous class.7 It is symmetric, bowl shaped, and differentiable everywhere and penalizes large forecast errors at an increasing rate due to its convexity in |e|. The loss function is not bounded from above. Large forecast errors or “outliers” are thus very costly under this loss function.
练习题
Which of the following best describes the Kuipers score for a binary outcome?
What is the main issue with using the Kuipers score for a sample with a single observation?
Which of the following are properties of the squared error loss function? (Select all that apply)
The Kuipers score can be thought of as the hit rate for plus the hit rate for minus a centering constant of 1.
The marginal value of a correct call in the Kuipers score depends only on the number of correct calls made so far.
The squared error loss function is given by , where . The loss function is __________ in .
The Kuipers score is positive if the sum of the positive and negative hit rates exceeds ___.
Explain why the Kuipers score can have poor properties for a single observation.
What is the relationship between the loss function and the utility function as described by Granger and Machina (2006)?
Which of the following are requirements for loss functions that depend only on the forecast error ? (Select all that apply)
Which property of the Kuipers score demonstrates its poor design for population problems?
Which conditions are satisfied by squared error loss but NOT by the Kuipers score? (Select all that apply)
The Kuipers score can be derived from a utility function using the relationship
The improvement in Kuipers score from adding a correct prediction is , where the blank represents ___.
登录后解锁笔记、知识点解析、AI 问答
立即登录