Forecasting - Causal Models
Forecasting - Causal Models
Time-series approaches are limited by their simplistic assumption that the future will be the same as the past. Time-series approaches use only historical data on the variable to be forecasted. Causal models refer to the introduction into the analysis of factors that are hypothesized to cause or influence, either directly or indirectly, the object of the forecast. Thus, if forecasting transit usage is the goal, such causal factors as population growth, the price of gasoline, and frequency of service might be introduced as factors that might cause or influence transit usage.
There are at least three reasons for moving to a causal analysis. First, time-series analyses usually cannot predict turning points and thus are incapable of forecasting accurately at times when an accurate forecast is most needed. If causal factors can be identified and linked to the forecast, then they can be used to help predict turning points. Actually, some of the qualitative techniques used, like the jury of executive opinion, consider causal factors but not in a specific formal way. Second, sometimes there is too much variation in a time-series to enable the detection of a trend. If a factor can be identified that is causally linked to some of that variation, the potential exists to remove it from the time-series, just as seasonal components were removed. Third, an understanding of causal relationships can be interesting and useful. For example, if the effect of schedules on transit usage could be predicted, policy concerning schedules might be affected.
Leading Indicators
Perhaps the simplest form of causal analysis is to attempt to identify leading indicators of the object to be forecasted. These indicators then could be monitored. In particular, we would hope that sharp changes in the growth rate of the indicators would predict turning points. Examples of leading indicators include the following:
1. The sales of color televisions are a leading indicator of the demand for components of color television sets. The sales of a product class are always a leading indicator of the sales of a supplier.
2. An analysis of demographic data frequently can provide useful indicators. The number of births is a leading indicator of the demand for education, and the number of people reaching 65 is a leading indicator of the demand for retirement facilities.
3. In technological forecasting, a leading indicator of the speed of commercial aircraft is the speed of military aircraft.
4. Economic data, such as disposable income, often are used as leading indicators of demand for such durables as automobiles, compactors, and the like.
Regression Models
A more formal causal model would be in the regression context. The dependent variable would be the object of the forecast. The independent variables would be the causal factors that are thought to cause or influence the forecast. The task, then, is to identify the causal factors and to obtain data representing them.
It is possible to include time as an independent variable, just as in time-series analyses. The causal factors then will be additional independent variables. The sales of the Simmons Company, makers of Beautyrest Mattresses, were hypothesized to have a linear trend and, in addition, to be caused by marriages during the year, housing starts, and annual disposable personal income.6 The resulting model was:
S = 49.85 - 0.07M + 0.04H + 1.22D - 19.54T r2 = .92 (1.2) (3.1) (2.0) (8.4) (7.3)
where
S = annual sales
M = annual number of marriages
H = annual housing starts
D = disposable income
T = time in years (the first year is 1, the second 2, etc.)
The numbers in parentheses, which are the t-values, indicate that the trend is highly significant, as is the disposable-income variable.
In an effort to improve the model's validity and performance, three changes were explored. Since marriages were thought to be related positively to sales, a negative coefficient was unexpected and cast doubt on the use of this variable, and the variable was therefore deleted. Also, the housing-start variable was thought to cause sales activity after about a year. To represent this lag between cause and effect, the number of housing starts in the previous year, Ht_l, was used instead of the number of current housing starts. Finally, last year's sales, St_ x, is introduced to reflect current sales momentum. If last year's sales were exceptionally good, perhaps conditions that contributed to their success still will hold. Thus, the inclusion of last year's sales might provide a better model. The new model has the following form:
S = -33.51 + 0.37St_! + 0.03^! + 0.67D - 11.03T r2 = 0.95
(2.1) (2.8) (2.4) (5.7) (5.1)
The revised model has an improved fit to the data, as reflected by the r2. Of course, the fit to the data is only an indication of the model's ability to forecast, in the future. In evaluating a causal model, an important consideration is how reasonable and defensible the causal variables and their coefficients are. In this case, the revised model was thought to represent more closely the realities of the modeling problem.
Forecasts also can be made by cross-sectional data, data that involve a fixed point in time. For example, one firm wanted to forecast the soft-drink consumption per capita by state.7 The hypothesis was that consumption was influenced by the mean temperature and by per-capita income. The resulting model was
C = -145.5 + 6.46X - 2.37Y r2 = 0.66
where
C = soft drink consumption per capita X = mean annual temperature Y = per-capita income
Thus, the consumption rate of a particular state can be predicted by knowing the values of the two causal variables. Such a prediction can be used to evaluate a special promotion in a particular state. The model can provide a basic prediction from which the results of the promotion can be compared.
Time-series approaches are limited by their simplistic assumption that the future will be the same as the past. Time-series approaches use only historical data on the variable to be forecasted. Causal models refer to the introduction into the analysis of factors that are hypothesized to cause or influence, either directly or indirectly, the object of the forecast. Thus, if forecasting transit usage is the goal, such causal factors as population growth, the price of gasoline, and frequency of service might be introduced as factors that might cause or influence transit usage.
There are at least three reasons for moving to a causal analysis. First, time-series analyses usually cannot predict turning points and thus are incapable of forecasting accurately at times when an accurate forecast is most needed. If causal factors can be identified and linked to the forecast, then they can be used to help predict turning points. Actually, some of the qualitative techniques used, like the jury of executive opinion, consider causal factors but not in a specific formal way. Second, sometimes there is too much variation in a time-series to enable the detection of a trend. If a factor can be identified that is causally linked to some of that variation, the potential exists to remove it from the time-series, just as seasonal components were removed. Third, an understanding of causal relationships can be interesting and useful. For example, if the effect of schedules on transit usage could be predicted, policy concerning schedules might be affected.
Leading Indicators
Perhaps the simplest form of causal analysis is to attempt to identify leading indicators of the object to be forecasted. These indicators then could be monitored. In particular, we would hope that sharp changes in the growth rate of the indicators would predict turning points. Examples of leading indicators include the following:
1. The sales of color televisions are a leading indicator of the demand for components of color television sets. The sales of a product class are always a leading indicator of the sales of a supplier.
2. An analysis of demographic data frequently can provide useful indicators. The number of births is a leading indicator of the demand for education, and the number of people reaching 65 is a leading indicator of the demand for retirement facilities.
3. In technological forecasting, a leading indicator of the speed of commercial aircraft is the speed of military aircraft.
4. Economic data, such as disposable income, often are used as leading indicators of demand for such durables as automobiles, compactors, and the like.
Regression Models
A more formal causal model would be in the regression context. The dependent variable would be the object of the forecast. The independent variables would be the causal factors that are thought to cause or influence the forecast. The task, then, is to identify the causal factors and to obtain data representing them.
It is possible to include time as an independent variable, just as in time-series analyses. The causal factors then will be additional independent variables. The sales of the Simmons Company, makers of Beautyrest Mattresses, were hypothesized to have a linear trend and, in addition, to be caused by marriages during the year, housing starts, and annual disposable personal income.6 The resulting model was:
S = 49.85 - 0.07M + 0.04H + 1.22D - 19.54T r2 = .92 (1.2) (3.1) (2.0) (8.4) (7.3)
where
S = annual sales
M = annual number of marriages
H = annual housing starts
D = disposable income
T = time in years (the first year is 1, the second 2, etc.)
The numbers in parentheses, which are the t-values, indicate that the trend is highly significant, as is the disposable-income variable.
In an effort to improve the model's validity and performance, three changes were explored. Since marriages were thought to be related positively to sales, a negative coefficient was unexpected and cast doubt on the use of this variable, and the variable was therefore deleted. Also, the housing-start variable was thought to cause sales activity after about a year. To represent this lag between cause and effect, the number of housing starts in the previous year, Ht_l, was used instead of the number of current housing starts. Finally, last year's sales, St_ x, is introduced to reflect current sales momentum. If last year's sales were exceptionally good, perhaps conditions that contributed to their success still will hold. Thus, the inclusion of last year's sales might provide a better model. The new model has the following form:
S = -33.51 + 0.37St_! + 0.03^! + 0.67D - 11.03T r2 = 0.95
(2.1) (2.8) (2.4) (5.7) (5.1)
The revised model has an improved fit to the data, as reflected by the r2. Of course, the fit to the data is only an indication of the model's ability to forecast, in the future. In evaluating a causal model, an important consideration is how reasonable and defensible the causal variables and their coefficients are. In this case, the revised model was thought to represent more closely the realities of the modeling problem.
Forecasts also can be made by cross-sectional data, data that involve a fixed point in time. For example, one firm wanted to forecast the soft-drink consumption per capita by state.7 The hypothesis was that consumption was influenced by the mean temperature and by per-capita income. The resulting model was
C = -145.5 + 6.46X - 2.37Y r2 = 0.66
where
C = soft drink consumption per capita X = mean annual temperature Y = per-capita income
Thus, the consumption rate of a particular state can be predicted by knowing the values of the two causal variables. Such a prediction can be used to evaluate a special promotion in a particular state. The model can provide a basic prediction from which the results of the promotion can be compared.
Comments
Post a Comment