Hypothesis Testing - Appendix Measures of Association for Nominal Variables

Hypothesis Testing - Appendix Measures of Association for Nominal Variables

We saw earlier in this chapter that the chi-square statistic is seriously flawed as a measure of the association of two variables. The nub of the problem is that the computed value of chi-square can tell us whether there is an association or a relationship but gives us only a weak indication of the strength of the association. The principal purpose of this appendix is to describe a measure, Goodman and Kruskal's Tau, which overcomes many of the problems of chi-square. First we look at some efforts to correct the problems of the chi-square measure. To illustrate these measures we return to Table 14-1 which is reproduced here. According to the chi-square test (x2 = 20), the relationship between location and attendance in Table 14-1 is highly significant. That is, there is a probability of less than .001 that the observed relationship could have happened by chance. Now we wish to know whether there is a sufficiently strong relationship to justify taking action.
Location
Convenient     Not Convenient Row Toted
Often 22 18 40
Attendance at     Occasionally 48 52 100
Symphony (Y)      Never 10 50 60
Column total 80 120 200


Measures Based on Chi-Square
The most obvious flaw of chi-square is that the value is directly proportional to the sample size. If the sample were 2000, rather than 200 in the previous table, and if the distribution of responses were the same (i. e., all cells were 10 times as large) the chi-square would be 200 rather than 20. Two measures have been proposed to overcome this problem:
20
(1) Phi-squared: cj>2 = x2/n =        = 0.10
(2) Contingency coefficient: d> = J—
n + x2
20      = 0.31
200 + 20

Both measures are easy to calculate but unfortunately are hard to interpret. On the one hand, when there is no association they are both zero. But when there is an association between the two variables, there is no upper limit against which to compare the calculated values. There is a special case, when the cross-tabulation has the same number of rows r and columns c, that an upper limit of the contingency coefficient can be computed for two perfectly correlated variables as V(r - l)/r.


Goodman and Kruskal's Tau

The Goodman and Kruskal's Tau is one of a class of measures that permit a "proportional reduction in error" interpretation. That is, a value of Tau between zero and one has a meaning in terms of the contribution of an independent variable—such as location—to explaining variation in a dependent variable such as attendance at the symphony.
The starting point for the calculation is the distribution of the dependent variable. Suppose we were given a sample of 200 people and, without knowing anything more about them, were given the task of assigning them to one of the three categories, as follows:

Number in Category

Attendance at Symphony (Y)
1. Often
2. Occasionally
3. Never

40 (20%
100 (50% I
60 (30%J
200


The only way to complete the task is to randomly draw 40 from the 200 for the first category, and then 100 for the second category, and so forth. The question is, how many of the 40 people who actually belong in the first category would wind up in that category if this procedure were followed? The answer is 40 times (40 -r 200), or 8 people. This is because a random draw of any size from the 200 will on average contain 20 percent who are attending the symphony regularly.
For the purpose of computing Tau, our interest is in the number of errors we would make by randomly assigning the known distribution of responses. For the first category this is 40 - 8 = 32 errors. Similarly, for the second ("occasionally") category we would make 100 (100 + 200) = 50 errors, and for the third category there would be 60 (140 -s- 200) = 42 errors. Therefore, we would expect to make 32 + 50 + 42 = 124 errors in placing the 200 individuals. Of course, we do not expect to make exactly 124 errors, but this would be the best estimate if the process were repeated a number of times.
The next question is whether knowledge of the independent variable will significantly reduce the number of errors. If the two variables are independent we would expect no reduction in the number of errors. To find out we simply repeat the same process we went through for the total sample within each category of the independent variable (that is, within each column). Let us start with the 80 people who said the location of the symphony was convenient. Again we would expect an average, when randomly assigning 22 people from the "convenient location" group to the "often attend" category, to have 22 (58 + 80) = 16 errors. In total, for the "convenient location" group we would expect 16 + 48 (32 80) + 10 (70 -r 80) = 44 errors. We next take the 120 in the "not convenient" group and randomly assign 18 to the first category, 52 to the second category, and 50 to the third category of the dependent variable. The number of errors that would result would be:

18 (102 -5- 120) + 52 (68     120) + 50 (70 -f- 120) = 74 errors
With knowledge of the category of the independent variable X to which each person belonged, the number of errors is 44 + 74 = 118. This is six errors fewer than when we didn't have that knowledge. The Tau measure will confirm that we haven't improved our situation materially by adding an independent variable, X:


We now have a measure that is theoretically more meaningful, since a value of 0 means no reduction in error and a value of 1.0 indicates prediction with no error. However, a value of 1.0 can be achieved only when all cells but one in a column are empty, and such a condition is often impossible given the marginal distributions. Thus, most tables have a ceiling on Tau which is less than 1.0. For example, for the tables on symphony attendance the best relationship we would expect to get—given the marginal distributions of the table—is as follows:

Location (X)

Convenient Not Convenient Total
Often 40 0 40
Occasionally 40 60 100
Never 0 60 60
80 120 200

The Tau for this table is 0.19. This value helps put our calculated value in perspective. The ceiling value of 0.19 for Tau, in a table that satisfies the known marginals, indicates that there is little basis for expecting a strong relationship in this table.

Comments

Popular posts from this blog

Catalog shows

Packing list

Factor Analysis - Factor Rotation