Fundamentals of Data Analysis - Sample Correlation
Fundamentals of Data Analysis - Sample Correlation
The correlation measures the degree to which there is an aggoci^inn_be-tween two intervally scaled variables. A positive correlation will reflect a tendency for a high value of one variable to be associated with a high value in the second. A negative correlation reflects an association between a high value on one variable and a low value on the second variable. Of course, one or both of the intervally scaled variables could be used to define categories such as age and income in Figure 13-3, but that would sacrifice information. If the data base included an entire population, such as all adults in California, the measure would be termed the population correlation. If. however, it is based on a sample, it is termed a sample correlation.
If two variables are plotted on a two-dimensional graph, termed a scatter diagram, the sample correlation reflects the tendency for the points to cluster systematically about a straight line rising or falling from left to right. The sample correlation is termed "r" and always is between 1 and +1. An r of +1 indicates a perfect positive association between the two variables, whereas if r is -1 there is perfect negative association. An r of 0 reflects the absence of any linear association.
Figure 13-4 illustrates five scatter diagrams. In Figure 13-4(a), there is a rather strong tendency for a small Y to be associated with a large X. The sample correlation is -.80. In Figure 13-4(b), the pattern slopes from the lower left to the upper right, and thus the sample correlation would be + .80. Figure 13-4(c) shows an example of a sample correlation of +1. It is a straight line running from the lower left to the upper right. Figure 13-4(d) shows an example in which there is no relationship between X and Y. Figure 13-4(e) shows a plot in which there is a clear relationship between the two variables, but it is not a linear or straightline relationship. Thus, the sample correlation is .00.
A more detailed conceptual explanation of the sample correlation and its calculation is presented in the appendix to this chapter. Further, when regression analysis is discussed in Chapter 19, a rather useful interpretation of the square of the sample correlation (r2) will be presented that will provide additional insights into its interpretation.
The correlation provides a measure of the relationship between two questions or variables. The underlying assumption is that the variables are intervally scaled (recall the discussion in Chapter 9), such as age or income. At issue is to what extent a variable must satisfy that criterion. Does a seven-point agree—disagree scale qualify? The answer depends in part upon the researcher's judgment about the scale. Is the difference between -2 and -1 the same as the difference between +2 and +3? If so, it qualifies. If not, a correlation analysis may still be useful but the results should be tempered with the knowledge that one or both of the scales may not be intervally scaled. In most cases insights gained and judgments made as a result of
correlation analysis will not be affected by departures from intervally scaled data. Of course, if one or both of the variables are 0—1 variables, then the more appropriate approach is the difference between means or cross-tabulation (0—1 variables only take on the values 0 or 1, such as user = 1 and nonuser = 0).
An example of the use of sample correlations comes from a study of donations toward a charitable health organization conducted by Stephen Miller. The organization studied contributions solicited by mail. A sample of 97 zip codes was selected for study. For each, the percentage of mail requests that resulted in donations of $5 or more was determined. The interest is in what characteristics of the zip code area would be correlated with donation behavior. Table 13-3 shows the paired correlations between the percent of families who donated and five other variables. The zip code areas that have a high percentage of families who receive interest income (from savings accounts, for example) have the highest correlation with mail solicitation response.
Hypothesis Testing
The appropriate hypothesis test in this situation is whether the sample correlation (the correlation based upon the sample) would be as large as it is if the population correlation (the correlation using data from the total population) were zero. The hypothesis tested is that the population correlation is zero. If the sample correlation is significant at the 0.10 level, then the probability of getting a sample correlation that large is below 0.10. In Table 13-3, however, another test is mentioned. The footnote in that table indicates that the cited correlations are significantly different, that is, the differences are unlikely to be due to a sampling accident. The hypothesis tested in Table 13-3 is that the two population correlations are the same.
The correlation measures the degree to which there is an aggoci^inn_be-tween two intervally scaled variables. A positive correlation will reflect a tendency for a high value of one variable to be associated with a high value in the second. A negative correlation reflects an association between a high value on one variable and a low value on the second variable. Of course, one or both of the intervally scaled variables could be used to define categories such as age and income in Figure 13-3, but that would sacrifice information. If the data base included an entire population, such as all adults in California, the measure would be termed the population correlation. If. however, it is based on a sample, it is termed a sample correlation.
If two variables are plotted on a two-dimensional graph, termed a scatter diagram, the sample correlation reflects the tendency for the points to cluster systematically about a straight line rising or falling from left to right. The sample correlation is termed "r" and always is between 1 and +1. An r of +1 indicates a perfect positive association between the two variables, whereas if r is -1 there is perfect negative association. An r of 0 reflects the absence of any linear association.
Figure 13-4 illustrates five scatter diagrams. In Figure 13-4(a), there is a rather strong tendency for a small Y to be associated with a large X. The sample correlation is -.80. In Figure 13-4(b), the pattern slopes from the lower left to the upper right, and thus the sample correlation would be + .80. Figure 13-4(c) shows an example of a sample correlation of +1. It is a straight line running from the lower left to the upper right. Figure 13-4(d) shows an example in which there is no relationship between X and Y. Figure 13-4(e) shows a plot in which there is a clear relationship between the two variables, but it is not a linear or straightline relationship. Thus, the sample correlation is .00.
A more detailed conceptual explanation of the sample correlation and its calculation is presented in the appendix to this chapter. Further, when regression analysis is discussed in Chapter 19, a rather useful interpretation of the square of the sample correlation (r2) will be presented that will provide additional insights into its interpretation.
The correlation provides a measure of the relationship between two questions or variables. The underlying assumption is that the variables are intervally scaled (recall the discussion in Chapter 9), such as age or income. At issue is to what extent a variable must satisfy that criterion. Does a seven-point agree—disagree scale qualify? The answer depends in part upon the researcher's judgment about the scale. Is the difference between -2 and -1 the same as the difference between +2 and +3? If so, it qualifies. If not, a correlation analysis may still be useful but the results should be tempered with the knowledge that one or both of the scales may not be intervally scaled. In most cases insights gained and judgments made as a result of
correlation analysis will not be affected by departures from intervally scaled data. Of course, if one or both of the variables are 0—1 variables, then the more appropriate approach is the difference between means or cross-tabulation (0—1 variables only take on the values 0 or 1, such as user = 1 and nonuser = 0).
An example of the use of sample correlations comes from a study of donations toward a charitable health organization conducted by Stephen Miller. The organization studied contributions solicited by mail. A sample of 97 zip codes was selected for study. For each, the percentage of mail requests that resulted in donations of $5 or more was determined. The interest is in what characteristics of the zip code area would be correlated with donation behavior. Table 13-3 shows the paired correlations between the percent of families who donated and five other variables. The zip code areas that have a high percentage of families who receive interest income (from savings accounts, for example) have the highest correlation with mail solicitation response.
Hypothesis Testing
The appropriate hypothesis test in this situation is whether the sample correlation (the correlation based upon the sample) would be as large as it is if the population correlation (the correlation using data from the total population) were zero. The hypothesis tested is that the population correlation is zero. If the sample correlation is significant at the 0.10 level, then the probability of getting a sample correlation that large is below 0.10. In Table 13-3, however, another test is mentioned. The footnote in that table indicates that the cited correlations are significantly different, that is, the differences are unlikely to be due to a sampling accident. The hypothesis tested in Table 13-3 is that the two population correlations are the same.
Comments
Post a Comment