Introducing a Third Variable - Spurious Association
Introducing a Third Variable - Spurious Association
Association measures, by themselves, do not demonstrate causation. This statement merits repeating because it is so easy to forget or suppress in the context of data analysis. Association measures do not demonstrate causation because they can be the result of extraneous variables. For example, the number of churches in a community is associated with the number liquor stores. Yet, few would maintain that churches tend to spawn liquor
stores (or the reverse). The fact is that this is a spurious association
because of the third variable, community size, which influences both the number of churches and the number of liquor stores. Another example is the fact that the amount of damage at a fire is associated with the number of fire trucks, only because a large fire both attracts fire trucks and causes substantial damage. If you add fire trucks, you don't increase the damage.
A major task of data analysis is to help the researcher identify such spurious associations, determine their nature and extent, and attempt to correct for them. Actually, the term "spurious association" is misleading, since the association is usually real enough. It is the causal interpretation that is spurious. However, the phrase has grown to mean an inappropriate causal interpretation of associations.
Association measures can be useful in prediction even if no causal analysis is involved. If a researcher needs to know the number of liquor stores in a community, a knowledge of the number of churches will be helpful, even though there is no causal link between the two. Often, however, there is a desire to gain a deeper understanding of the relationship. Even if prediction is the primary objective, a better understanding of the relationship may suggest the conditions under which the association can be used for predictions and the likely conditions under which the prediction might err. Sometimes, however, the desire is not only to predict the dependent variable but to influence it. Thus, an organization may want to influence the number of churches in a community—to increase their number, for example. Then there is a need to identify variables that influence the number of churches so that effective programs can be developed and Implemented.
Transit Study
The possibility that a third variable is a source of spurious association will be considered first in the context of an example. Table 15-1 shows an analysis of a study of advertising for a transit system. The system had not done any previous media advertising. To test the effectiveness of getting people to use the system more, a series of television advertisements were run with the appeal that money, time, and aggravation could be saved by riding the transit system. A survey of people in the area just after the test period generated 360 responses. Three variables were isolated for analysis:
I: Whether the respondent intended (J = 1) or did not intend (1 = 0) to use the system during the coming four-week period
A: Whether the respondent could (A = 1) or could not (A = 0) recall the transit advertising
U: Whether the respondent had (L7 = 1) or had not (U = 0) used the system during the four-week period preceding the test advertising
The objective was to determine if the advertising was successful in generating intentions to use the system. Thus, the analysis starts by looking at the relationship between advertising and intentions. The first cross-tabulation in Table 15-1 explores this relationship. Notice that the percentage is based on intention, which is assumed to be the variable to be influenced, the dependent variable. The advertising recall variable (A) is thus the causal or independent variable. Table 15-1 shows that the association between advertising (A) and intentions {I) is high. The association can be determined by comparing the percentage who intend to use the system. It is confirmed by the size and statistical significance of the chi-square statistic, here a measure of association. Thus an implication could be that advertising influences intentions via:
A CAUSAL RELATIONSHIP A-»I
The next step in the analysis was to introduce the usage variable. Of the 360 respondents, 160 were transit users and 200 were transit non-users. Of the 160 transit users recall of the advertising did not affect intentions. The percentage who intended to use was about the same for those who saw the advertising (85 percent) and those that did not (80 percent). Similarly, of the 200 transit nonusers intentions were virtually identical (20 percent vs. 18 percent) for the "recall advertising" group and the "not recall advertising" group. The reason that the association between advertising and intentions was found for all respondents is that transit users tended to recall the advertising much more than transit nonusers.
A SPURIOUS RELATIONSHIP
BETWEEN A AND I
A <- U-> I
Figure 15-1 shows the association between advertising recall and intentions in graphic form. It shows that the 63 percent of those with positive intentions who recalled the advertising represented an average of 50 non-users and 100 users. Further, the 36 percent of those not recalling the advertising but who intended to use represented an average of 150 non-users and 60 users. The reason that the association between advertising and intentions was explained by the usage variable was because of the difference between the composition of the "recalled-advertising" group and the "didn't-recall-advertising" group. The users tended to notice and recall the advertising whereas the nonusers were not as interested in the transit system or its advertising. If the composition had been identical, the association between advertising and intentions would not have been affected by the introduction of the usage variable.
When the association between two variables is explained by the identification of a third variable influencing both, the association is termed spurious. The extent to which an association provides evidence of a causal relationship is largely based on the confidence that the association is not spurious. Thus, an effort to discover measured or unmeasured extraneous variables is an important task of analysis.
Experimental Control
To understand this logic completely, it is necessary to recall the concept of experimental control introduced in Chapter 10. Suppose an experiment was designed to test the effect of advertising on intentions by selecting a group of 100 users and another group of 100 nonusers. Half of each of these groups would be selected randomly and exposed to the advertisements. Thus, a user subgroup of size 50 would be exposed to the advertisement and the remaining 50 users would not be exposed. Intentions would then be measured for all 200 subjects. The result would be a randomized block design:
Here, the usage variable is controlled by being introduced as a block. Using the notation of Chapter 10, the symbolic representation is:
[R X 01 n = 50
Nonusers \
^R Oa n = 50
(R X 03 n = 50
Users i
R 04 n = 50
An experiment also can control for variables without introducing them as a block in the design. Assume that the experiment is conducted on 200 randomly selected people, including users and nonusers. The advertisement-exposure treatment is given to half, who are selected randomly. The result would be a simple randomized design:
R 0 n = 100
R X 0 n = 100
Any resulting association between advertising and intentions then could be assumed to be devoid of any spuriousness. It might be that, by accident, the exposed group had an abnormally high representation of non-users, or users. Although that possibility might have to be considered (and could be considered), the fact that the exposed group was selected randomly acts to minimize that possibility, as well as the possibility that other variables besides usage were generating (spuriously) the association between advertising and intentions. This randomization is a way of controlling extraneous variables, such as usage and others that may be unknown.
In the "test" represented in Table 15-1, the exposed and unexposed groups were self-selected. People were assigned randomly to the two groups, and therein lies the possibility that other factors caused the self-selection to occur. One of these factors, as has been illustrated, was usage. There may be others.
Statistical Control
A second method to control for variables that are a source of spurious association is termed statistical control. An example is the process ::' introducing usage in Table 15-1 as a control variable. The association between A and I for the nonuser group then could be interpreted as association with usage held constant—all are nonusers. The association for tie user group has the same interpretation.
Another illustration comes from a chain of retail book stores. Suppose store sales were thought to be influenced by store advertising. Suppose however, that the association between advertising and sales was causer! spuriously by store size. The large stores tended to have higher sales and :: advertise more. A solution might be to control statistically for store size r determining the association between store sales and advertising for large stores and the association for small stores. In Chapter 20 the use of regression analysis to accomplish statistical control will be presented.
Statistical control provides a way to proceed when a true experiment is not conducted, and it often will control adequately for the identified variable. The problem is that there might be other variables that are unmeasured and even unknown. These will remain uncontrolled and will be a source of misinterpretation. The analyst should draw on existing theory and common sense to attempt to identify possible sources of spurious association and to determine the extent to which they may be affecting the analysis. The advantage of randomization in experimental control is that ■ will control for both known and unknown variables.
Combination of Causal Paths
There are times when the association between two variables is completely explained by the introduction of a third variable; however, the more common situation is where the association is reduced but not eliminated by the introduction of the control variable. In Table 15-1 the association was stil positive, although small, after the introduction of the control variable. If 1 were a bit larger, it might have been reasonable to conclude that the original association represented partly a causal link and partly a spurious association:
We now turn to the other three ways in which a third variable can be introduced into the analysis: as an intervening variable, as an additive cause, and as an interactive cause.
Association measures, by themselves, do not demonstrate causation. This statement merits repeating because it is so easy to forget or suppress in the context of data analysis. Association measures do not demonstrate causation because they can be the result of extraneous variables. For example, the number of churches in a community is associated with the number liquor stores. Yet, few would maintain that churches tend to spawn liquor
stores (or the reverse). The fact is that this is a spurious association
because of the third variable, community size, which influences both the number of churches and the number of liquor stores. Another example is the fact that the amount of damage at a fire is associated with the number of fire trucks, only because a large fire both attracts fire trucks and causes substantial damage. If you add fire trucks, you don't increase the damage.
A major task of data analysis is to help the researcher identify such spurious associations, determine their nature and extent, and attempt to correct for them. Actually, the term "spurious association" is misleading, since the association is usually real enough. It is the causal interpretation that is spurious. However, the phrase has grown to mean an inappropriate causal interpretation of associations.
Association measures can be useful in prediction even if no causal analysis is involved. If a researcher needs to know the number of liquor stores in a community, a knowledge of the number of churches will be helpful, even though there is no causal link between the two. Often, however, there is a desire to gain a deeper understanding of the relationship. Even if prediction is the primary objective, a better understanding of the relationship may suggest the conditions under which the association can be used for predictions and the likely conditions under which the prediction might err. Sometimes, however, the desire is not only to predict the dependent variable but to influence it. Thus, an organization may want to influence the number of churches in a community—to increase their number, for example. Then there is a need to identify variables that influence the number of churches so that effective programs can be developed and Implemented.
Transit Study
The possibility that a third variable is a source of spurious association will be considered first in the context of an example. Table 15-1 shows an analysis of a study of advertising for a transit system. The system had not done any previous media advertising. To test the effectiveness of getting people to use the system more, a series of television advertisements were run with the appeal that money, time, and aggravation could be saved by riding the transit system. A survey of people in the area just after the test period generated 360 responses. Three variables were isolated for analysis:
I: Whether the respondent intended (J = 1) or did not intend (1 = 0) to use the system during the coming four-week period
A: Whether the respondent could (A = 1) or could not (A = 0) recall the transit advertising
U: Whether the respondent had (L7 = 1) or had not (U = 0) used the system during the four-week period preceding the test advertising
The objective was to determine if the advertising was successful in generating intentions to use the system. Thus, the analysis starts by looking at the relationship between advertising and intentions. The first cross-tabulation in Table 15-1 explores this relationship. Notice that the percentage is based on intention, which is assumed to be the variable to be influenced, the dependent variable. The advertising recall variable (A) is thus the causal or independent variable. Table 15-1 shows that the association between advertising (A) and intentions {I) is high. The association can be determined by comparing the percentage who intend to use the system. It is confirmed by the size and statistical significance of the chi-square statistic, here a measure of association. Thus an implication could be that advertising influences intentions via:
A CAUSAL RELATIONSHIP A-»I
The next step in the analysis was to introduce the usage variable. Of the 360 respondents, 160 were transit users and 200 were transit non-users. Of the 160 transit users recall of the advertising did not affect intentions. The percentage who intended to use was about the same for those who saw the advertising (85 percent) and those that did not (80 percent). Similarly, of the 200 transit nonusers intentions were virtually identical (20 percent vs. 18 percent) for the "recall advertising" group and the "not recall advertising" group. The reason that the association between advertising and intentions was found for all respondents is that transit users tended to recall the advertising much more than transit nonusers.
A SPURIOUS RELATIONSHIP
BETWEEN A AND I
A <- U-> I
Figure 15-1 shows the association between advertising recall and intentions in graphic form. It shows that the 63 percent of those with positive intentions who recalled the advertising represented an average of 50 non-users and 100 users. Further, the 36 percent of those not recalling the advertising but who intended to use represented an average of 150 non-users and 60 users. The reason that the association between advertising and intentions was explained by the usage variable was because of the difference between the composition of the "recalled-advertising" group and the "didn't-recall-advertising" group. The users tended to notice and recall the advertising whereas the nonusers were not as interested in the transit system or its advertising. If the composition had been identical, the association between advertising and intentions would not have been affected by the introduction of the usage variable.
When the association between two variables is explained by the identification of a third variable influencing both, the association is termed spurious. The extent to which an association provides evidence of a causal relationship is largely based on the confidence that the association is not spurious. Thus, an effort to discover measured or unmeasured extraneous variables is an important task of analysis.
Experimental Control
To understand this logic completely, it is necessary to recall the concept of experimental control introduced in Chapter 10. Suppose an experiment was designed to test the effect of advertising on intentions by selecting a group of 100 users and another group of 100 nonusers. Half of each of these groups would be selected randomly and exposed to the advertisements. Thus, a user subgroup of size 50 would be exposed to the advertisement and the remaining 50 users would not be exposed. Intentions would then be measured for all 200 subjects. The result would be a randomized block design:
Here, the usage variable is controlled by being introduced as a block. Using the notation of Chapter 10, the symbolic representation is:
[R X 01 n = 50
Nonusers \
^R Oa n = 50
(R X 03 n = 50
Users i
R 04 n = 50
An experiment also can control for variables without introducing them as a block in the design. Assume that the experiment is conducted on 200 randomly selected people, including users and nonusers. The advertisement-exposure treatment is given to half, who are selected randomly. The result would be a simple randomized design:
R 0 n = 100
R X 0 n = 100
Any resulting association between advertising and intentions then could be assumed to be devoid of any spuriousness. It might be that, by accident, the exposed group had an abnormally high representation of non-users, or users. Although that possibility might have to be considered (and could be considered), the fact that the exposed group was selected randomly acts to minimize that possibility, as well as the possibility that other variables besides usage were generating (spuriously) the association between advertising and intentions. This randomization is a way of controlling extraneous variables, such as usage and others that may be unknown.
In the "test" represented in Table 15-1, the exposed and unexposed groups were self-selected. People were assigned randomly to the two groups, and therein lies the possibility that other factors caused the self-selection to occur. One of these factors, as has been illustrated, was usage. There may be others.
Statistical Control
A second method to control for variables that are a source of spurious association is termed statistical control. An example is the process ::' introducing usage in Table 15-1 as a control variable. The association between A and I for the nonuser group then could be interpreted as association with usage held constant—all are nonusers. The association for tie user group has the same interpretation.
Another illustration comes from a chain of retail book stores. Suppose store sales were thought to be influenced by store advertising. Suppose however, that the association between advertising and sales was causer! spuriously by store size. The large stores tended to have higher sales and :: advertise more. A solution might be to control statistically for store size r determining the association between store sales and advertising for large stores and the association for small stores. In Chapter 20 the use of regression analysis to accomplish statistical control will be presented.
Statistical control provides a way to proceed when a true experiment is not conducted, and it often will control adequately for the identified variable. The problem is that there might be other variables that are unmeasured and even unknown. These will remain uncontrolled and will be a source of misinterpretation. The analyst should draw on existing theory and common sense to attempt to identify possible sources of spurious association and to determine the extent to which they may be affecting the analysis. The advantage of randomization in experimental control is that ■ will control for both known and unknown variables.
Combination of Causal Paths
There are times when the association between two variables is completely explained by the introduction of a third variable; however, the more common situation is where the association is reduced but not eliminated by the introduction of the control variable. In Table 15-1 the association was stil positive, although small, after the introduction of the control variable. If 1 were a bit larger, it might have been reasonable to conclude that the original association represented partly a causal link and partly a spurious association:
We now turn to the other three ways in which a third variable can be introduced into the analysis: as an intervening variable, as an additive cause, and as an interactive cause.
Comments
Post a Comment