Chapter 1: Introduction to Statistics Section 1-1: Statistical and Critical Thinking ............................................................................1 Section 1-2: Types of Data ........................................................................................................2 Section 1-3: Collecting Sample Data .........................................................................................3 Quick Quiz .................................................................................................................................5 Review Exercises .......................................................................................................................5 Cumulative Review Exercises ...................................................................................................6
Chapter 2: Exploring Data with Tables and Graphs Section 2-1: Frequency Distributions for Organizing and Summarizing Data ..........................7 Section 2-2: Histograms ...........................................................................................................11 Section 2-3: Graphs That Enlighten and Graphs That Deceive ...............................................14 Section 2-4: Scatterplots, Correlation, and Regression ...........................................................17 Quick Quiz ...............................................................................................................................19 Review Exercises .....................................................................................................................19 Cumulative Review Exercises .................................................................................................21
Chapter 3: Describing, Exploring, and Comparing Data Section 3-1: Measures of Center ..............................................................................................23 Section 3-2: Measures of Variation .........................................................................................27 Section 3-3: Measures of Relative Standing and Boxplots......................................................31 Quick Quiz ...............................................................................................................................35 Review Exercises .....................................................................................................................35 Cumulative Review Exercises .................................................................................................36
Chapter 4: Probability Section 4-1: Basic Concepts of Probability .............................................................................39 Section 4-2: Addition Rule and Multiplication Rule ...............................................................40 Section 4-3: Complements, Conditional Probability, and Bayes’ Theorem ............................42 Section 4-4: Risks and Odds ....................................................................................................43 Section 4-5: Rates of Mortality, Fertility, and Morbidity ........................................................45 Section 4-6: Counting ..............................................................................................................47 Quick Quiz ...............................................................................................................................49 Review Exercises .....................................................................................................................50 Cumulative Review Exercises .................................................................................................51
Chapter 5: Discrete Probability Distributions Section 5-1: Probability Distributions .....................................................................................53 Section 5-2: Binomial Probability Distributions......................................................................54 Section 5-3: Poisson Probability Distributions ........................................................................57 Quick Quiz ...............................................................................................................................59 Review Exercises .....................................................................................................................59 Cumulative Review Exercises .................................................................................................60
Copyright © 2024 Pearson Education, Inc.
Chapter 6: Normal Probability Distributions Section 6-1: The Standard Normal Distribution ......................................................................63 Section 6-2: Real Applications of Normal Distributions .........................................................64 Section 6-3: Sampling Distributions and Estimators ...............................................................66 Section 6-4: The Central Limit Theorem .................................................................................70 Section 6-5: Assessing Normality............................................................................................74 Section 6-6: Normal as Approximation to Binomial ...............................................................77 Quick Quiz ...............................................................................................................................79 Review Exercises .....................................................................................................................80 Cumulative Review Exercises .................................................................................................81
Chapter 7: Estimating Parameters and Determining Sample Sizes Section 7-1: Estimating a Population Proportion.....................................................................83 Section 7-2: Estimating a Population Mean.............................................................................88 Section 7-3: Estimating a Population Standard Deviation or Variance ...................................91 Section 7-4: Bootstrapping: Using Technology for Estimates ................................................95 Quick Quiz ...............................................................................................................................96 Review Exercises .....................................................................................................................97 Cumulative Review Exercises .................................................................................................98
Chapter 8: Hypothesis Testing Section 8-1: Basics of Hypothesis Testing ............................................................................101 Section 8-2: Testing a Claim About a Proportion ..................................................................102 Section 8-3: Testing a Claim About a Mean ..........................................................................108 Section 8-4: Testing a Claim About a Standard Deviation or Variance ................................112 Section 8-5: Resampling: Using Technology for Hypothesis Testing...................................115 Quick Quiz .............................................................................................................................117 Review Exercises ...................................................................................................................117 Cumulative Review Exercises ...............................................................................................119
Chapter 9: Chapter 9: Inferences from Two Samples Section 9-1: Two Proportions ................................................................................................121 Section 9-2: Two Means: Independent Samples ....................................................................129 Section 9-3: Matched Pairs ....................................................................................................136 Section 9-4: Two Variances or Standard Deviations .............................................................141 Section 9-5: Resampling: Using Technology for Inferences .................................................143 Quick Quiz .............................................................................................................................145 Review Exercises ...................................................................................................................145 Cumulative Review Exercises ...............................................................................................147
Copyright © 2024 Pearson Education, Inc.
Chapter 10: Correlation and Regression Section 10-1: Correlation .......................................................................................................151 Section 10-2: Regression .......................................................................................................157 Section 10-3: Prediction Intervals and Variation ...................................................................165 Section 10-4: Multiple Regression.........................................................................................167 Section 10-5: Dummy Variables and Logistic Regression ....................................................170 Quick Quiz .............................................................................................................................171 Review Exercises ...................................................................................................................172 Cumulative Review Exercises ...............................................................................................173
Chapter 11: Goodness-of-Fit and Contingency Tables Section 11-1: Goodness-of-Fit ...............................................................................................177 Section 11-2: Contingency Tables .........................................................................................182 Quick Quiz .............................................................................................................................186 Review Exercises ...................................................................................................................186 Cumulative Review Exercises ...............................................................................................187
Chapter 12: Analysis of Variance Section 12-1: One-Way ANOVA ..........................................................................................191 Section 12-2: Two-Way ANOVA .........................................................................................193 Quick Quiz .............................................................................................................................194 Review Exercises ...................................................................................................................195 Cumulative Review Exercises ...............................................................................................196
Chapter 13: Nonparametric Tests Section 13-2: Sign Test ..........................................................................................................199 Section 13-3: Wilcoxon Signed-Ranks Test for Matched Pairs ............................................200 Section 13-4: Wilcoxon Rank-Sum Test for Two Independent Samples ..............................201 Section 13-5: Kruskal-Wallis Test for Three or More Samples ............................................205 Section 13-6: Rank Correlation .............................................................................................207 Quick Quiz .............................................................................................................................208 Review Exercises ...................................................................................................................209 Cumulative Review Exercises ...............................................................................................210
Chapter 14: Survival Analysis Section 14-1: Life Tables .......................................................................................................213 Section 14-2: Kaplan-Meier Survival Analysis .....................................................................215 Quick Quiz .............................................................................................................................217 Review Exercises ...................................................................................................................217 Cumulative Review Exercises ...............................................................................................218
Copyright © 2024 Pearson Education, Inc.
Chapter 1: Introduction to Statistics Section 1-1: Statistical and Critical Thinking 1. The respondents are a voluntary response sample or a self-selected sample. Because those with strong interests in the topic are more likely to respond, it is very possible that their responses do not reflect the opinions or behavior of the general population. 2. a. The sample consists of the 1046 adults who were surveyed. The population consists of all adults. b. When asked, respondents might be inclined to avoid the shame of the unhealthy habit of not washing their hands, so the reported rate of 70% might well be much higher than it is in reality. It is generally better to observe or measure human behavior than to ask subjects about it. 3. Statistical significance is indicated when methods of statistics are used to reach a conclusion that a treatment is effective, but common sense might suggest that the treatment does not make enough of a difference to justify its use or to be practical. It is possible for a study to have statistical significance, but not practical significance. 4. No. Correlation does not imply causation. The example illustrates a correlation that is clearly not the result of any interaction or cause effect relationship between per capita consumption of margarine and the divorce rate in Maine. 5. Yes, there does appear to be a potential to create a bias. 6. No, there does not appear to be a potential to create a bias. 7. No, there does not appear to be a potential to create a bias. 8. Yes, there does appear to be a potential to create a bias. 9. The sample is a voluntary response sample and has strong potential to be flawed. 10. The samples are voluntary response samples and have potential for being flawed, but this approach might be necessary due to ethical considerations involved in randomly selecting subjects and somehow imposing treatments on them. 11. The sampling method appears to be sound. 12. The sampling method appears to be sound. 13. The Ornish weight loss program has statistical significance, because the results are so unlikely (3 chances in 1000) to occur by chance. It does not have practical significance because the amount of lost weight (3.3 lb) is so small. 14. Because there is only one chance in a thousand of getting such success rates by chance, the difference does appear to have statistical significance. The 92% success rate for surgery appears to be substantially better than the 72% success rate for splints, so the difference does appear to have practical significance. 15. The difference between Mendel’s 25% rate and the result of 26% is not statistically significant. According to Mendel’s theory, 145 of the 580 peas would have yellow pods, but the results consisted of 152 peas with yellow pods. The difference of 7 peas with yellow pods among the 580 offspring does not appear to be statistically significant. The difference does not appear to have practical significance. 16. Because there is a 25% chance of getting such results with a program that has no effect, the program does not appear to have statistical significance. Because the average increase is only 3 IQ points, the program does not appear to have practical significance. 17. The sample percentage of males is 49.5%, and it appears to be very close to the percentage expected under normal circumstances. It does not appear to have statistical significance, nor does it appear to have practical significance. 18. Because there is a 15% chance of getting such results with a medication that has no effect, the medication does not appear to have statistical significance. Because the average decrease is only 2 mmHg, the medication does not appear to have practical significance. 19. Because there is an 8% chance of getting such nausea rates by chance, the results do not appear to have statistical significance. Also, they do not appear to have practical significance. 20. Because there is less than a 1% chance of getting the results obtained in this study, the results have statistical significance. They also appear to have practical significance. 21. Yes. Each column of 8 AM and 12 AM temperatures is recorded from the same subject, so each pair is matched. 22. No. The source is from university researchers who do not appear to gain from distorting the data.
Copyright © 2024 Pearson Education, Inc.
23. The data can be used to address the issue of whether there is a correlation between body temperatures at 8 AM and at 12 AM. Also, the data can be used to determine whether there are differences between body temperatures at 8 AM and at 12 AM. 24. Because the differences could easily occur by chance (with a 64% chance), the differences do not appear to have statistical significance. 25. No. The white blood cell counts measure a different quantity than the red blood cell counts, so their differences are meaningless. 26. The issue that can be addressed is whether there is a correlation, or association, between white blood cell counts and red blood cell counts. 27. No. The National Center for Health Statistics has no reason to collect or present the data in a way that is biased. 28. No. Correlation does not imply causation, so a statistical correlation between white blood cell counts and red blood cell counts should not be used to conclude that higher white blood cell counts are the cause of higher red blood cell counts. 29. It is questionable that the sponsor is the Idaho Potato Commission and the favorite vegetable is potatoes. 30. The sample is a voluntary response sample, so there is a good chance that the results do not reflect the larger population of people who have a water preference. 31. The correlation, or association, between two variables does not mean that one of the variables is the cause of the other. Correlation does not imply causation. Clearly, sour cream consumption is not directly related in any way to motorcycle fatalities. 32. The sponsor of the poll is an electronic cigarette maker, so the sponsor does have an interest in the poll results. The source is questionable. 33. The correlation, or association, between two variables does not mean that one of the variables is the cause of the other. Correlation does not imply causation. Common sense suggests that cheese consumption is not directly related in any way to fatalities from bedsheet entanglements. 34. The correlation, or association, between two variables does not mean that one of the variables is the cause of the other. Correlation does not imply causation. 35. The survey results are from subjects who chose to respond, so the results constitute a voluntary response sample. Consequently, the results are questionable. 36. Because the nutritionists are paid such large amounts of money, they might be more inclined to find favorable results. It is very possible that the results represent desired outcomes instead of actual outcomes. 37. a. 0.45(3014) = 1356.3 b. No. The actual count of survey subjects who have at least one chronic condition must be a whole number. c. 1356 adults (Any value from 1342 to 1371 would yield a percentage that rounds to 45%.) d. 1206 / (1206 + 1808) = 0.40, or 40% 38. a. 0.828(198) = 163.944 patients b. No. Because the result is a count of patients, the result must be a whole number. c. 164 patients d. 198 / (199 + 198) = 0.499, or 49.9%, or 50% rounded 39. The wording of the question is biased and tends to encourage negative responses. The sample size of 20 is too small. Survey respondents are self-selected instead of being randomly selected by the newspaper. If 20 readers respond, the percentages should be multiples of 5, so 87% and 13% are not possible results. 40. All percentages of success should be multiples of 5. The given percentages cannot be correct. Section 1-2: Types of Data 1. The population consists of all drug tests of adults in the United States, and the sample is the 10 million drug tests that were analyzed. The value of 4.2% is a statistic because it is obtained from the sample. 2. a. quantitative d. quantitative b. categorical e. quantitative c. categorical
Copyright © 2024 Pearson Education, Inc.
3. 4.
Only part (a) describes discrete data. a. The sample is the 36,000 adults who were surveyed. The population is all adults in the United States. b. statistic c. ratio d. discrete 5. statistic 17. discrete 6. statistic 18. continuous 7. parameter 19. discrete 8. parameter 20. continuous 9. statistic 21. nominal 10. parameter 22. ordinal 11. parameter 23. ordinal 12. parameter 24. ratio 13. continuous 25. interval 14. discrete 26. nominal 15. continuous 27. ratio 16. discrete 28. interval 29. The numbers are not counts or measures of anything. They are at the nominal level of measurement, and it makes no sense to compute the average (mean) of them. 30. The ranks are at the ordinal level of measurement. Differences between medical schools cannot be interpreted, so there is no way to know whether the difference between Harvard and New York University is the same as the difference between New York University and Duke. 31. The numbers on the pain scale are at the ordinal level of measurement. Differences cannot be determined, so there is no way to know whether the first patient has twice as much pain as the second patient. Ratios such as “twice” make no sense with ordinal data. 32. The temperatures are at the interval level of measurement. Because there is no natural starting point with 0ºF representing no heat, it is wrong to state that the second patient is “5% cooler than the first patient.” 33. a. Continuous, because the number of possible values is infinite and not countable. b. Discrete, because the number of possible values is finite. c. Discrete, because the number of possible values is finite. d. Discrete, because the number of possible values is infinite and countable. 34. Interval level of measurement. The direction of north represented by 0° is arbitrary, and 0° does not represent “no direction.” Differences between degrees are meaningful; the difference between 30° and 60° is the same as the difference between 150° and 180°. But ratios are not meaningful; the ratio of 60° to 30° does not result in twice some direction. (These degree measurements are directions, not amounts of rotation.) Section 1-3: Collecting Sample Data 1. The study is an experiment because subjects were given treatments. 2. The subjects in the study did not know whether they were given the magnet treatment or the sham treatment, and those who administered the treatments also did not know. 3. The group sample sizes are large enough so that the researchers could see the effects of the two treatments, but it would have been better to have larger samples. 4. The sample appears to be a convenience sample. Given that the subjects were all patients at a Veterans Affairs hospital, it is not likely that the sample is representative of the population, so it is questionable whether the results can be generalized for the population of subjects with chronic low back pain. 5. The sample appears to be a convenience sample. By e-mailing the survey to a readily available group of Internet users, it was easy to obtain results. Although there is a real potential for getting a sample group that is not representative of the population, indications of which ear is used for cell phone calls and which hand is dominant do not appear to be factors that would be distorted much by a sample bias.
Copyright © 2024 Pearson Education, Inc.
6. 7.
The study is an observational study because the subjects were not given any treatment. With 717 responses, the response rate is 14%, which does appear to be quite low. In general, a very low response rate creates a serious potential for getting a biased sample that consists of those with a special interest in the topic. 8. Answers vary, but the following are good possibilities. a. Obtain a printed copy of the class roster, assign consecutive numbers (integers), then use a computer to randomly generate six of those numbers. b. Select every third student leaving class until six students are chosen. c. Randomly select three males and three females. d. Randomly select a row, and then select the students in that row. (Use only the first six to meet the requirement of a sample of size six.) e. Select the first six students who enter the class. 9. systematic 15. stratified 10. convenience 16. systematic 11. random 17. random 12. stratified 18. cluster 13. cluster 19. convenience 14. random 20. systematic 21. Observational study. The sample is a convenience sample consisting of subjects who decided to respond. Such voluntary response samples have a high chance of not being representative of the larger population, so the sample may well be biased, as it was in this case. 22. Experiment. The sample subjects consist of male physicians only. It would have been better to include females. Also, it would be better to include male and females who are not physicians. 23. Experiment. This experiment would create an extremely dangerous and illegal situation that has a real potential to result in injury or death. It’s difficult enough to drive in New York City while being completely sober. 24. Observational study. The sample of eight respondents is too small. 25. Experiment. The biased sample created by using a small sample of college students cannot be fixed by using a larger sample. The larger sample will still be a biased sample that is not representative of the population of all adults. 26. Experiment. Calling the subjects and asking them to report their weights has a high risk of getting results that do not reflect the actual weights. It would have been much better to somehow measure the weights instead of asking the subjects to report them. 27. Observational study. Respondents are not likely to respond honestly because there is a “social desirability bias” causing respondents to reply in ways that will be viewed favorably. 28. Observational study. The number of responses is very small, and the response rate of only 1.52% is far too small. With such a low response rate, there is a real possibility that the sample of respondents is biased and consists only of those with special interests in the survey topic. 29. prospective study 33. matched pairs design 30. retrospective study 34. randomized block design 31. cross-sectional study 35. completely randomized design 32. prospective study 36. matched pairs design 37. Prospective: The experiment was begun and results were followed forward in time. Randomized: Subjects were assigned to the different groups through a process of random selection, whereby they had the same chance of belonging to each group. Double-blind: The subjects did not know which of the three groups they were in, and the people who evaluated results did not know either. Placebo-controlled: There was a group of subjects who were given a placebo; by comparing the placebo group to the two treatment groups, the effects of the treatments might be better understood.
Copyright © 2024 Pearson Education, Inc.
38. a. Not a simple random sample, but it is a random sample. b. Simple random sample and also a random sample. c. Not a simple random sample and not a random sample. Quick Quiz 1. No. The numbers do not measure or count anything. 2. nominal 6. statistic 3. continuous 7. no 4. quantitative data 8. observational study 5. ratio 9. The subjects did not know whether they were getting aspirin or the placebo. 10. simple random sample Review Exercises 1. The respondents are a voluntary response sample or a self-selected sample. Because those with strong interests in the topic are more likely to respond, it is very possible that their responses do not reflect the opinions or behavior of the general population. 2. a. The survey uses a voluntary response sample, and those with special interests are more likely to respond, so it is very possible that the sample is not representative of the population. b. Because the statement refers to 72% of all Americans, it is a parameter (but it is probably based on a 72% rate from the sample, and the sample percentage is a statistic). c. observational study 3. Randomized: Subjects were assigned to the different groups through a process of random selection, whereby they had the same chance of belonging to each group. Double-blind: The subjects did not know which of the three groups they were in, and the people who evaluated results did not know either. 4. No. Correlation does not imply causality. 5. a. systematic d. convenience b. stratified e. cluster c. simple random sample 6. The survey was sponsored by the American Laser Centers, and 24% said that the favorite body part is the face, which happens to be a body part often chosen for some type of laser treatment. The source is therefore questionable. 7. a. discrete b. ratio c. The mailed responses would be a voluntary response sample, so those with strong opinions or greater interest in the topics are more likely to respond. It is very possible that the results do not reflect the true opinions of the population of all state residents. d. stratified e. cluster 8.
a. 0.628(199) = 125 patients b. 164 /198 = 82.8%
9.
a. interval data; systematic sample b. nominal data; stratified sample c. ordinal data; convenience sample 10. Because there is less than a 1% chance of getting the results by chance, the method does appear to have statistical significance. The result of 239 males in 291 births is a rate of 82% so it is above the 50% rate expected by chance, and it does appear to be high enough to have practical significance. The procedure appears to have both statistical significance and practical significance.
Copyright © 2024 Pearson Education, Inc.
Cumulative Review Exercises
135 + 149 + 145 + 129 + 118 + 119 + 115 + 133 + 107 + 188 + 127 + 131 = 133.0. The IQ score of 188 12 appears to be substantially higher than the other IQ scores.
1.
The mean is
2.
0.513 = 0.000122
3.
203 - 176 = 4.50, which is an unusually high value. 6
4.
5.
98.2 - 98.6 = - 6.64 0.62 106
1.959962 ×0.25 0.032
6.
7.
= 1068
188 - 107 = 20.25 4
(135 - 133.0)2 11
Copyright © 2024 Pearson Education, Inc.
= 0.364
8. 9.
(98.4 - 98.6)2 + (98.6 - 98.6)2 + (98.8 - 98.6)2 3- 1
= 0.04 = 0.20
0.36 = 0.000729
10. 812 = 68,719,476,736 (or about 68,719,477,000) 11. 856 = 377,149,515,625 (or about 377,149,520,000) 12. 0.212 = 0.000000004096
Chapter 2: Exploring Data with Tables and Graphs Section 2-1: Frequency Distributions for Organizing and Summarizing Data 1. No. Instead of frequencies that start low, reach a maximum, and then decrease, the given distribution starts with a maximum frequency of 14 and the frequencies then decrease. 2. Nicotine (mg) 1.0–1.1 1.2–1.3 1.4–1.5 1.6–1.7 1.8–1.9 3.
Relative Frequency 56% 16% 12% 12% 4%
There is overlap in the class limits of 1.0–1.2, 1.2–1.4, and so on. For an amount such as 1.2 mg, it would not be clear which class contains the value. The classes should not overlap. 4. The sum of the relative frequencies is 125%, but it should be 100%, with a small round off error. All of the relative frequencies appear to be roughly the same, but if they are from a normal distribution, they should start low, reach a maximum, and then decrease. 5. Class width: 2.0 Class midpoints: 2.95, 4.95, 6.95, 8.95, 10.95, 12.95, 14.95 Class boundaries: 1.95, 3.95, 5.95, 7.95, 9.95, 11.95, 13.95, 15.95 Number: 147 6. Class width: 2.0 Class midpoints: 2.95, 4.95, 6.95, 8.95, 10.95, 12.95 Class boundaries: 1.95, 3.95, 5.95, 7.95, 9.95, 11.95, 13.95 Number: 153 7. Class width: 100 Class midpoints: 49.5, 149.5, 249.5, 349.5, 449.5, 549.5, 649.5 Class boundaries: –0.5, 99.5, 199.5, 299.5, 399.5, 499.5, 599.5, 699.5 Number: 153 8. Class width: 100 Class midpoints: 149.5, 249.5, 349.5, 449.5, 549.5 Class boundaries: 99.5, 199.5, 299.5, 399.5, 499.5, 599.5 Number: 147 9. Yes. The frequencies start low, reach a maximum, and then decrease. 10. Not normal. The frequencies are not approximately symmetric. The maximum frequency is in the second class. 11. Yes. Except for the single value that lies between 600 and 699, the frequencies start low, reach a maximum of 90, and then decrease. The values below the maximum are very roughly a mirror image of those above it. (That single value between 600 and 699 is an outlier that makes the determination of a normal distribution somewhat questionable, but using a loose interpretation of the criteria for normality, it is reasonable to conclude that the distribution is normal.) 12. Yes. Except for two values that lie between 500 and 599, there is a low frequency of 25, then a maximum frequency of 92, and then a low frequency of 28. The values below and above the maximum are roughly a mirror image. (Those two values between 500 and 599 are outliers that make the determination of a normal distribution somewhat questionable, but using a loose interpretation of the criteria for normality, it is reasonable to conclude that the distribution is normal.)
13. The pulse rates do appear to be from a normal distribution. Pulse Rate 40–49 50–59 60–69 70–79 80–89 90–99
Frequency 1 10 13 9 5 2
14. The pulse rates do appear to be from a normal distribution. Pulse Rate 30–39 40–49 50–59 60–69 70–79 80–89 90–99 100–109
Frequency 1 0 4 7 10 11 6 1
15. The verbal IQ scores do appear to be from a population having a normal distribution. Verbal IQ 50–59 60–69 70–79 80–89 90–99 100–109 110–119 120–129
Frequency 1 2 6 9 4 4 2 2
16. The verbal IQ scores do appear to be from a population having a normal distribution.
Verbal IQ
Frequency
60–69
1
70–79
7
80–89
8
90–99
4
100–109
1
17. No, the pulse rates are not dramatically different. The pulse rates appear to be from a normal distribution. Pulse Rates 40–49 50–59 60–69 70–79 80–89 90–99
Frequency 2 23 53 43 25 5
Copyright © 2024 Pearson Education, Inc.
Section 2-1: Frequency Distributions for Organizing and Summarizing Data 100–109
2
Copyright © 2024 Pearson Education, Inc.
7
18. No, the pulse rates are not dramatically different. The pulse rates appear to be from a normal distribution. Pulse Rates 30–39 40–49 50–59 60–69 70–79 80–89 90–99 100–109
Frequency 1 1 17 33 41 37 13 4
19. The distribution does appear to be a normal distribution. Weight (kg) in September 50–59 60–69 70–79 80–89 90–99
Frequency 2 12 11 3 4
20. No. There does not appear to be such a dramatic weight gain. Weight (kg) in April 50–59 60–69 70–79 80–89 90–99 100–109
Frequency 3 12 8 7 1 1
21. Because there are disproportionately more 0s and 5s, it appears that the heights were reported instead of measured. Consequently, it is likely that the results are not very accurate. Last Digit 0 1 2 3 4 5 6 7 8 9
Frequency 9 2 1 3 1 15 2 0 3 1
Copyright © 2024 Pearson Education, Inc.
Section 2-1: Frequency Distributions for Organizing and Summarizing Data 22. Because there are disproportionately more 0s and 5s, it appears that the heights were reported instead of measured. There does appear to be a gap due to the tendency of respondents to round their heights to values ending in 0 or 5. Because the results appear to be reported instead of measured, it is likely that the results are not very accurate. Last Digit 0 1 2 3 4 5 6 7 8 9
Frequency 26 1 1 2 2 12 1 0 4 1
23. The two distributions appear to be very similar. White Blood Cell Count 2.0–3.9 4.0–5.9 6.0–7.9 8.0–9.9 10.0–11.9 12.0–13.9 14.0–15.9
Females 4.8% 38.1% 31.3% 19.7% 5.4% 0.0% 0.7%
Males 5.9% 39.2% 32.7% 19.0% 2.0% 1.3% 0.0%
24. There do appear to be differences, but overall, they are not very substantial differences. Blood Platelet Count 0–99 100–199 200–299 300–399 400–499 500–599 600–699
Males 0.7% 33.3% 58.8% 6.5% 0.0% 0.0% 0.7%
Females 0.0% 17.0% 62.6% 19.0% 0.0% 1.4% 0.0%
25. White Blood Cell Count of Females Less than 4.0 Less than 6.0 Less than 8.0 Less than 10.0 Less than 12.0 Less than 14.0 Less than 16.0
Cumulative Frequency 7 63 109 138 146 146 147
Copyright © 2024 Pearson Education, Inc.
9
26. White Blood Cell Count of Males Less than 4.0 Less than 6.0 Less than 8.0 Less than 10.0 Less than 12.0 Less than 14.0
Cumulative Frequency 9 69 119 148 151 153
27. Because only the five leading causes of death are listed, we know only that any other cause of death must have fewer than 150,005 deaths. Cause of Death Heart Disease Cancer Accidents Chronic Lower Respiratory Disease Stroke
Relative Frequency 37.9% 34.5% 10.0% 9.0% 8.6%
28. Yes, it appears that births occur on the days of the week with frequencies that are about the same. Day Monday Tuesday Wednesday Thursday Friday Saturday Sunday
Relative Frequency 13.0% 16.5% 18.0% 14.3% 14.3% 10.8% 13.3%
29. The frequencies aren’t very symmetric, but using a very loose interpretation, the measurements do appear to be from a normal distribution. Systolic Blood Pressure (mmHg) of Females 80–99 100–119 120–139 140–159 160–179 180–199
Frequency 8 62 53 21 2 1
Section 2-2: Histograms 1. The histogram should be bell-shaped. 2. Not necessarily. Because the sample subjects themselves chose to be included, the voluntary response sample might not be representative of the population. 3. With a data set that is so small, the true nature of the distribution cannot be seen with a histogram. 4. The outlier will result in a single bar that is far away from all of the other bars in the histogram, and the height of that bar will correspond to a frequency of 1. 5. Approximately 50 6. Approximate values: Class width: 0.5 mm, lower limit of first class: 2.0 mm, upper limit of first class: 2.5 mm 7. The largest possible value is approximately 4.5 mm, which is not an outlier.
Copyright © 2024 Pearson Education, Inc.
Section 2-2: Histograms 8. 9.
11
The histogram very roughly approximates a bell shape, so it appears that the sample is from a population having a normal distribution. The distribution appears to be normal, and there are no outliers.
10. The distribution appears to be normal, and there are no outliers.
11. The distribution appears to be normal, and there are no outliers.
12. The distribution appears to be approximately normal (or perhaps skewed to the right), and there are no outliers.
Copyright © 2024 Pearson Education, Inc.
13. The distribution does not appear to be normal. It appears to be skewed to the right.
14. The distribution does not appear to be normal. It appears to be skewed to the right.
15. The distribution does not appear to be normal. It appears to be skewed to the right.
16. The distribution is dramatically far from normal. It is skewed to the right.
Copyright © 2024 Pearson Education, Inc.
Section 2-2: Histograms
13
17. The digits 0 and 5 appear to occur more often than the other digits, so it appears that the heights were reported and not actually measured. This suggests that the data might not be very useful.
18. The digits 0 and 5 appear to occur more often than the other digits, so it appears that the weights were reported and not actually measured. This suggests that the data might not be very useful.
19. Only part (c) appears to represent data from a normal distribution. Part (a) has a systematic pattern that is not that of a straight line, part (b) has points that are not close to a straight-line pattern, and part (d) is really bad because it shows a systematic pattern and points that are not close to a straight-line pattern. 20. The histogram for the movie lengths suggests that the data have a distribution that is approximately normal with an outlier of 120 minutes. The histogram for the time of tobacco use appears to be similar to the histogram for the times of alcohol use, but both distributions appear to be skewed to the right, with each histogram having an outlier. Section 2-3: Graphs That Enlighten and Graphs That Deceive 1. The data set is too small for a dotplot to reveal important characteristics of the data. Because the data are listed in order for each of the last several years, a time-series graph would be most effective for these data. 2. Yes, the original data values can be found from the stemplot.
5 6 7 8 9 3. 4. 5.
4 8 0369 123 8
No. Graphs should be constructed in a way that is fair and objective. The readers should be allowed to make their own judgments, instead of being manipulated by misleading graphs. No. If the sample is a bad sample, such as one obtained from voluntary responses, there are no graphs or statistical methods that can be used to salvage the data. The pulse rate of 36 beats per minute appears to be an outlier.
Copyright © 2024 Pearson Education, Inc.
6.
There do not appear to be any outliers.
7.
The data are arranged in order from lowest to highest, as 36, 56, 56, and so on.
3 4 5 6 7 8 9 8.
668 044666 6888 02468 4
The two values closest to the middle are 72 mmHg and 74 mmHg.
6 7 8 9 9.
6
0022468 0000246688 22468 00
There was a steep jump in the first four years, but the numbers of triplets have shown a downward trend in the past several years.
10. The first five years don’t show much change, but the trend starts to climb dramatically in 1999.
Copyright © 2024 Pearson Education, Inc.
11. Misconduct includes fraud, duplication, and plagiarism, and it does appear to be a major factor.
12.
13.
14.
15. The distribution appears to be roughly bell-shaped, so the distribution is approximately normal.
Copyright © 2024 Pearson Education, Inc.
16. The distribution appears to be roughly bell-shaped, so the distribution is approximately normal.
17. The two costs are one-dimensional in nature, but the baby bottles are three-dimensional objects. The $4500 cost isn’t even twice the $2600 cost, but the baby bottles make it appear that the larger cost is about five times the smaller cost. 18. The graph is misleading because it depicts one-dimensional data with three-dimensional boxes. See the first and last boxes in the graph. Workers with advanced degrees have annual incomes that are roughly 3 times the incomes of those with no high school diplomas, but the graph exaggerates this difference by making it appear that workers with advanced degrees have incomes that are roughly 27 times the amounts for workers with no high school diploma. 19. 96
96 97 97 98 98 99 99
59 0001112333444 55666666788888999 00000000000002222233444444444444 5555666666666666666777777888888899 001244 56
20. Because the data are listed in order according to day, a time-series graph would be most revealing about the nature of the data. The time-series graph shows two distinct major peaks: in around the beginning of 2021 and in September 2021.
Section 2-4: Scatterplots, Correlation, and Regression 1. The term linear refers to a straight line, and r measures how well a scatterplot of the sample paired data fits a straight-line pattern. 2. No. Finding the presence of a statistical correlation between two variables does not justify any conclusion that one of the variables is a cause of the other. 3.
A scatterplot is a graph of paired ( x, y ) quantitative data. It helps us by providing a visual image of the data plotted as points, and such an image is helpful in enabling us to see patterns in the data and to recognize that there may be a correlation between the two variables. Copyright © 2024 Pearson Education, Inc.
4.
a. 1 b. 0
5.
c. 0 d. - 1 There does not appear to be a linear correlation between brain volume and IQ score.
6.
There does appear to be a linear correlation between the chest sizes and weights of bears.
7.
There does not appear to be a linear correlation between the heights of fathers and the heights of their first sons.
8.
There does not appear to be a correlation between pulse rates of females and males. The major flaw with this exercise is that the data are not paired as required. The results are therefore meaningless.
Copyright © 2024 Pearson Education, Inc.
With n = 5 pairs of data, the critical values are ±0.878. Because r = 0.127 is between –0.878 and 0.878, there is not sufficient evidence to conclude that there is a linear correlation. 10. With n = 7 pairs of data, the critical values are ±0.754. Because r = 0.980 is in the right tail region beyond 0.754, there is sufficient evidence to conclude that there is a linear correlation. 11. With n = 10 pairs of data, the critical values are ±0.632. Because r = - 0.017 is between –0.632 and 0.632, there is not sufficient evidence to conclude that there is a linear correlation. 12. With n = 10 pairs of data, the critical values are ±0.632. Because r = - 0.076 is between –0.632 and 0.632, there is not sufficient evidence to conclude that there is a linear correlation. The data are not paired, so the results are meaningless. 13. Because the P-value of 0.839 is not small (such as 0.05 or less), there is a high chance of getting the sample results when there is no correlation. There is not sufficient evidence to conclude that there is a linear correlation. 14. Because the P-value of 0.0001 is small (such as 0.05 or less), there is a small chance of getting the sample results when there is no correlation. There is sufficient evidence to conclude that there is a linear correlation. 15. Because the P-value of 0.963 is not small (such as 0.05 or less), there is a high chance of getting the sample results when there is no correlation. There is not sufficient evidence to conclude that there is a linear correlation. 16. Because the P-value of 0.835 is not small (such as 0.05 or less), there is a high chance of getting the sample results when there is no correlation. There is not sufficient evidence to conclude that there is a linear correlation. 17. The scatterplot shows a pattern suggesting, not too surprisingly, that as the numbers of new cases of COVID-19 increase, the numbers of deaths also increase. There appears to be a correlation. Among the 800 points, there appears to be a somewhat strong correlation except for the 45 points farthest to the right. A flaw in the paired data is that vaccinations began roughly a year into the pandemic, so the paired data of new cases and deaths consists of a population that is changing instead of being constant. 9.
Quick Quiz 1. The class width is 20. 2. The class boundaries are 19.5 and 39.5. 3. No, it is impossible to determine the original values. 4. 153, 154, 158 7. time-series graph 5. The histogram will be bell-shaped. 8. scatterplot 6. variation 9. Pareto chart 10. A frequency distribution is in the format of a table; a histogram is a graph. Review Exercises 1. Reported Weight (lb) 135–159 160–184 185–209 210–234 235–259
Frequency 2 7 5 5 1
Copyright © 2024 Pearson Education, Inc.
2.
The data appear to be from a population with a distribution that is approximately normal. The bars start low, reach a maximum, and then decrease, and the left half of the histogram is approximately a mirror image of the right half. The graph is approximately bell-shaped.
3.
By using fewer classes, the histogram does a better job of illustrating the distribution.
4.
There are no outliers.
13 14 15 16 17 18 19 20 21 22 23 24
8 0 0079 00 2 0038 5 26 355 7
5.
Yes. There is a pattern suggesting that there is a relationship.
6.
a. time-series graph c. Pareto chart
b. scatterplot
Copyright © 2024 Pearson Education, Inc.
7.
By using a vertical scale that starts at 49.0% instead of 0%, the difference is greatly exaggerated. The graph creates the false impression that female enrollees outnumber male enrollees by a ratio of roughly 2.5 to 1, but the actual percentages of 50.5% and 49.5% are very much closer than that. If the graph had been created with a vertical axis starting at 0%, the difference between the two bars would be almost imperceptible. Cumulative Review Exercises 1. Grooming Time (min) 0–9 10–19 20–29 30–39 40–49
Frequency 2 3 9 4 2
2.
The histogram is approximately bell-shaped. The frequencies increase to a maximum and then decrease, and the left half of the histogram is roughly a mirror image of the right half. The data do appear to be from a population with a normal distribution.
3.
0 1 2 3 4
4.
There are disproportionately more last digits of 0 and 5. Fourteen of the 20 times have last digits of 0 or 5. It appears that the subjects reported their own results and they tended to round the results. The data do not appear to be very accurate.
05 255 024555778 0055 05
Last Digit 0 1 2 3 4 5 6 7 8 9
Frequency 5 0 2 0 1 9 0 2 1 0
5.
6.
a. ratio b. continuous c. No. The grooming times are quantitative data. d. statistic The scatterplot helps address the issue of whether there is a correlation between the heights of mothers and the heights of their first daughters. The scatterplot does not reveal a clear pattern suggesting that there is a correlation.
Chapter 3: Describing, Exploring, and Comparing Data Section 3-1: Measures of Center 1. The term average is not used in statistics. The term mean should be used for the result obtained by adding all of the sample values and dividing the total by the number of sample values. 2. No. The 50 amounts are all weighed equally in the calculation that yields the mean of 17.5%, but states have different populations, so the mean should be calculated using a weighted mean that takes into account the populations in the different states. (The CDC reported a value of 17.1% for all of the states.) 3.
4.
80 + 94 + 58 + 66 + 56 = 70.8 bpm and the median is 66.0 bpm. Using the 5 six values that include the outlier, x = 182.3 bpm and the median is 73.0 bpm. The outlier caused the mean to change by a substantial amount, but the median did not change by very much. The median is resistant to the effect of the outlier; the mean is not resistant. They all use different approaches for providing a value (or values) of the center or middle of the sorted list of data. Using the five values, the mean is x =
1.8 + 1.7 + 1.7 + 1.6 + 1.5 + 1.4 + 1.4 + 1.3 + 1.3 + 1.2 + 1.2 + 1.1 = 1.43 mg. 12 1.4 + 1.4 = 1.40 mg. The median is 2 The modes are 1.7 mg, 1.4 mg, 1.3 mg, and 1.2 mg. 1.1 + 1.8 = 1.45 mg. The midrange is 2 Apart from the fact that the other nicotine amounts are lower than those given, nothing meaningful can be known about the sample of all such amounts.
5.
The mean is x =
6.
The mean is x =
1 + 14 + 4 + 16 + 2 + 15 + 3 + 15 + 19 + 5 + 11 + 13 + 14 + 9 = 10.1 g. 14 11 + 13 = 12.0 g. The median is 2 The modes are 14 g and 15 g. Copyright © 2024 Pearson Education, Inc.
1 + 19 = 10.0 g. 2 Americans consume some brands of cereal much more often than others, but the 14 brands are all weighted equally in the calculations, so the statistics are not necessarily representative of the population of all servings of cereal consumed by Americans. The midrange is
7.
The mean is x =
88.3 + 86.5 + 71.3 + 81.6 + 75.6 +
+ 83.1 + 90.4 + 78.6 + 75.1 + 69.2 15
= 79.50 kg.
The median is 82.40 kg. There is no mode. 59.9 + 90.4 = 75.15 kg. The midrange is 2 Because the measurements were made in 1988, they are not necessarily representative of the current population of all males in the Army.
Copyright © 2024 Pearson Education, Inc.
8.
The mean is x =
1560 + 1665 + 1711 + 1660 + 1572 +
+ 1668 + 1654 + 1666 + 1599 + 1673 16
= 1641.8 mm.
1660 + 1665 = 1662.5 mm. 2 The mode is 1707 mm. 1521 + 1711 = 1616.0 mm. The midrange is 2 The Army has height restrictions, so females in the Army are not necessarily representative of the population of all adult females. The median is
9.
The mean is x =
524 + 607 + 266 + 485 + 405 + 723 +
+ 540 + 249 + 692 + 675 + 545 + 373 18
= 491.6.
524 + 540 = 532.0. 2 There is no mode. 249 + 723 = 486.0. The midrange is 2 The identification numbers are nominal data that are just replacements for names. They do not measure or count anything, so the resulting statistics are meaningless. The median is
10. The mean is x =
69 + 70 + 67 + 68 + 70 + 73 +
+ 68 + 69 + 70 + 73 + 72 + 72 20
= 70.5 in.
70 + 70 = 70.0 in. 2 The mode is 72 in. 66 + 81 = 73.5 in. The midrange is 2 Because the heights were reported instead of being measured, it is very possible that the results are not likely to represent the population. The median is
11. The mean is x =
41,945 + 42,196 + 43, 005 +
+ 37, 473 + 36,560 + 36,120 20
= 38, 000.0 deaths
37, 423 + 37, 473 = 37, 448 deaths. 2 There is no mode. 32, 479 + 43,510 = 37,994.5 deaths. The midrange is 2 The data are time-series data, but the measures of center do not reveal anything about a trend consisting of a pattern of change over time. The median is
2 +1 +1 +1 +1 +1 +1 + 4 +1 + 2 + 2 +1 + 2 + 3 + 3 + 2 + 3 + 1+ 3 + 1+ 3 + 1+ 3 + 2 + 2 = 1.9. 25 The median is 2.0. The mode is 1. 1+ 4 = 2.5. The midrange is 2 The mode of 1 correctly indicates that the smooth-yellow peas occur more than any other phenotype, but the other measures of center do not make sense with these data at the nominal level of measurement.
12. The mean is x =
Copyright © 2024 Pearson Education, Inc.
13. The mean is x =
0+0+0+
+ 34 + 36 + 38 + 41 + 41 + 41 + 20
+ 53 + 54 + 55
= 32.6 mg.
38 + 41 = 39.5 mg. 2 The mode is 0 mg. 0 + 55 = 27.5 mg. The midrange is 2 Americans consume some brands much more often than others, but the 20 brands are all weighted equally in the calculations, so the statistics are not necessarily representative of the population of all cans of the same 20 brands consumed by Americans. The median is
7.9 + 7.7 + 8 + 7.9 + 7.9 + 7.8 + 7.7 + 7.8 + 7.4 + 6.8 + 7.2 + 7 + 6.8 = 7.53%. 13 The median is 7.70%. The mode is 7.9%. 6.8 + 8 = 7.40%. The midrange is 2 The data are time-series data, but the measures of center do not reveal anything about a trend consisting of a pattern of change over time.
14. The mean is x =
9 + 10 + 10 + 20 + 40 + 50 + + 0 + 0 + 0 + 0 + 0 + 0 = 2.8 cigarettes. 50 The median is 0.0 cigarettes. The mode is 0 cigarettes. 0 + 50 = 25.0 cigarettes. The midrange is 2 Because the selected subjects report the number of cigarettes smoked, it is very possible that the data are not at all accurate. And what about that person who smokes 50 cigarettes (or 2.5 packs) a day? What are they thinking?
15. The mean is x =
1.18 + 1.41 + 1.49 + 1.04 + 1.45 + 0.74 + 0.89 + 1.42 + 1.45 + 0.51 + 1.38 = 1.178 W/kg. 11 The median is 1.380 W/kg. The mode is 1.45 W/kg. 0.51 + 1.49 = 1.000 W/kg. The midrange is 2 If concerned about radiation absorption, you might purchase the cell phone with the lowest radiation level. All of the cell phones in the sample have radiation levels below the FCC maximum of 1.6 W/kg. 17. Systolic: 96 + 116 + 118 + 120 + 122 + 126 + 128 + 136 + 156 + 158 = 127.6 mmHg, The mean is x = 10 122 + 126 = 124.0 mmHg. The median is 2 Diastolic: 52 + 58 + 64 + 72 + 74 + 76 + 80 + 82 + 88 + 90 = 73.6 mmHg. The mean is x = 10 74 + 76 = 75.0 mmHg. The median is 2 Given that systolic and diastolic blood pressures measure different characteristics, a comparison of the measures of center doesn’t make sense. Because the data are matched, it would make more sense to investigate whether there is an association or correlation between systolic blood pressure measurements and diastolic blood pressure measurements. (FYI: A helpful measure is the mean arterial pressure, which is (1 / 3)(systolic) + (2 / 3)diastolic.) 16. The mean is x =
Copyright © 2024 Pearson Education, Inc.
18. White blood cells:
8.7 + 4.9 + 6.9 + 7.5 + 6.1 + 5.7 + 4.1 + 8.1 + 8 + 5.6 + 8.3 + 6.9 = 6.73 (1000 cells mL). 12 6.9 + 6.9 = 6.90 (1000 cells mL). The median is 2 Red blood cells: 4.8 + 4.7 + 4.5 + 4.3 + 5 + 4 + 4.7 + 4.6 + 4.1 + 5.5 + 4.4 + 4.2 = 4.57 ( million cells mL). The mean is x = 12 4.5 + 4.6 = 4.55 ( million cells mL). The median is 2 Given that the white and red blood cell counts measure different characteristics and use different units of measurement, a comparison of the measures of center doesn’t make much sense. Because the data are matched, it would make more sense to investigate whether there is an association or correlation between white blood cell counts and red blood cell counts. 19. Males: 4.9 + 7.5 + 6.1 + 5.7 + 4.1 + 5.6 + 8.3 + 5.1 + 9.5 + 6.1 + 5.7 + 5.4 = 6.17 (1000 cells mL). The mean is x = 12 5.7 + 5.7 = 5.70 (1000 cells mL). The median is 2 Females: 8.7 + 6.9 + 8.1 + 8 + 6.9 + 8.1 + 6.4 + 6.3 + 10.9 + 4.8 + 5.9 + 7.2 = 7.35 (1000 cells mL). The mean is x = 12 6.9 + 7.2 = 7.05 (1000 cells mL). The median is 2 Females appear to have higher white blood cell counts than males. 20. Single Line: 390 + 396 + 402 + 408 + 426 + 438 + 444 + 462 + 462 + 462 = 429.0 seconds. The mean is x = 10 426 + 438 = 432.0 seconds. The median is 2 Individual Lines: 252 + 324 + 348 + 372 + 402 + 462 + 462 + 510 + 558 + 600 = 429.0 seconds. The mean is x = 10 402+462 = 432.0 seconds. The median is 2 Although the measures of center are the same, the times with individual lines are much more varied than those with a single line. The mean is x =
21. Using all values, the mean is x = 53.7 mg/dL and the median is 52.0 mg/dL. The highest value of 138 mg/dL appears to be an outlier. Excluding 138 mg/dL, the mean is x = 53.4 mg/dL and the median is 52.0 mg/dL. Excluding the outlier does not cause much of a change in the mean, and the median remains the same. 22. The mean is x = 99,342.2 new cases and the median is 56,515.5 new cases. Because of vaccinations, the numbers of new cases were changing over time, so the population was not constant. 23. The mean is x = 98.20°F and the median is 98.40°F. These results suggest that the mean is less than 98.6°F. 24. The mean is x = 3152.0 g and the median is 3300.0 g. All of the weights end in 00, so they are all rounded to the nearest 100 grams. This suggests that the results should be rounded as follows: x = 3150.0 g and the median is 3300 g.
Copyright © 2024 Pearson Education, Inc.
1(49.5) + 51(149.5) + 90 (249.5) + 10 (349.5) + 0 (449.5) + 0 (549.5) +1(649.5) 1 + 51 + 90 + 10 + 0 + 0 + 1 = 224.0 (1000 cells / mL). The mean from the frequency distribution is quite close to the mean of
25. The mean is x =
224.3 (1000 cells / mL) obtained by using the original list of values. 25 (149.5) + 92 (249.5) + 28 (349.5) + 0 (449.5) + 2 (549.5) = 255.6 (1000 cells / mL). The mean 25 + 92 + 28 + 0 + 2 from the frequency distribution is quite close to the mean of 225.1 (1000 cells / mL) obtained by using the original list of values.
26. The mean is x =
æ63 + 91 + 88 + 84 + 79 ö + 0.10 (86) + 0.15 (90) + 0.15 (70) = 81.2, 27. The mean is x = 0.60 ç ÷ è ø 5 so the student earned a B.
59, 470(64,966) + 89,310 (7,924) + 77, 000 (55,176) + 60, 780 (35, 430) + 106,950 (337, 738) = $93,957 64,966 + 7,924 + 55,176 + 35, 430 + 337, 738 The weighted mean salary of the 501,234 nurses is $93,957. 59, 470 + 89,310 + 77, 000 + 60, 780 + 106,950 = $78, 702. b. The mean of the five listed salaries is 5 c. The results differ by a considerable amount. The mean salary of $93,957 is better because it is a weighted mean that takes into account the different numbers of nurses in the five different states.
28. a.
29. a. The missing value is 5 (78.0) - 82 - 78 - 56 - 84 = 90 beats per minute. b. n - 1 30. If we include the censored values, the mean is 15.9 years or greater. The results of 15.5 years and 15.9 years do not differ by a considerable amount. 31. Mean: 113.7 mg/dL; 10% trimmed mean: 112.6 mg/dL; 20% trimmed mean: 112.2 mg/dL. The 10% trimmed mean and 20% trimmed mean are fairly close, but the untrimmed mean of 113.7 mg/dL differs from them because it is more strongly affected by the outliers. Section 3-2: Measures of Variation 1.
193.9 - 155.0 = 9.58 cm, which is in the general ballpark of the standard deviation of 7.10 cm calculated 4 using the 153 heights. The range rule of thumb does not necessarily give an estimate of s that is very accurate s»
2.
Significantly low values are less than or equal to 174.12 - 2(7.10) = 159.92 cm and significantly high values are greater than or equal to 174.12 + 2(7.10) = 188.32 cm. A height of 190 cm is significantly high.
3.
(7.10 cm)2 = 50.41 cm2
4.
(a) s, (b) s , (c) s 2, (d) s 2 ; If sample data consist of weights measured in grams, s and s would also be in
( )
units of grams (g), but s 2 and s 2 would be in units of grams squared g 2 . 5.
The range is 1.8 - 1.1 = 0.70 mg. The variance is s 2 =
(1.8 - 1.43) 2 + (1.7 - 1.43) 2 +
+ (1.2 - 1.43) 2 + (1.1 - 1.43) 2 12
= 0.05 mg.
The standard deviation is 0.05 = 0.23 mg. Because the top 12 sample amounts are used, the variation is not at all typical for the entire sample.
Copyright © 2024 Pearson Education, Inc.
6.
The range is 19 - 1 = 18.0 g.
(1 - 10.1) 2 + (14 - 10.1) 2 + + (14 - 10.1) 2 + (9 - 10.1) 2 = 35.8 g 2 . 14 - 1 The standard deviation is s = 35.8 = 6.0 g. Because the top 12 sample amounts are used, the variation is not at all typical for the entire sample. Americans consume some brands of cereal much more often than others, but the 14 brands are all weighted equally in the calculations, so the statistics are not necessarily representative of the population of all servings of cereal consumed by Americans. The range is 90.4 - 59.9 = 30.50 kg. The variance is s 2 =
7.
(88.3 - 79.5) 2 + (86.5 - 79.5) 2 + + (75.1 - 79.5) 2 + (69.2 - 79.5) 2 = 65.32 kg 2 . 15 - 1 The standard deviation is s = 65.32 = 8.08 kg. Because the measurements were made in 1988, they are not necessarily representative of the current population of all males in the Army. The range is 1711 - 1521 = 190.0 mm. The variance is s 2 =
8.
The variance is s 2 =
9.
16(43,173,572) - (26, 268)2 = 3205.5 mm2 . 16(16 - 1)
The standard deviation is s = 3205.5 = 56.6 mm. The Army has height restrictions, so females in the Army are not necessarily representative of the population of all adult females. The range is 723 - 249 = 474.0. The variance is s 2 =
18(4,784,320) - (8848)2 = 25,590.4. 18(18 - 1)
The standard deviation is s = 25,590.4 = 160.6. The identification numbers are nominal data that are just replacements for names. They do not measure or count anything, so the resulting statistics are meaningless. 10. The range is 81 - 66 = 15.0 in. The variance is s 2 =
20(99,602) - (1410)2 = 10.4 in.2. 20(20 - 1)
The standard deviation is s = 10.4 = 3.2 in. Because the heights were reported instead of being measured, it is very possible that the results are not likely to represent the population. 11. The range is 43,510 - 32, 479 = 11,031.0 deaths. The variance is s 2 =
20(29, 204,367, 422) - (759,990) 2 = 17,111,969.3 deaths 2 . 20(20 - 1)
The standard deviation is s = 17,111,969.3 = 4136.7 deaths. The data are time-series data, but the measures of variation do not reveal anything about a trend consisting of a pattern of change over time. 12. The range is 4 - 1 = 3.0. 2
The variance is s 2 =
25 (109) - (47 ) = 0.9. 25 (25 - 1)
The standard deviation is s = 0.9 = 0.9. Because the data are at the nominal level of measurement, the measures of variation are meaningless.
Copyright © 2024 Pearson Education, Inc.
13. The range is 55 - 0 = 55.0 mg. 2
20 (29,045) - (651) = 413.4 mg 2 . The variance is s = 20 (20 - 1) 2
The standard deviation is s 2 = 413.4 = 20.3 mg. Americans consume some brands much more often than others, but the 20 brands are all weighted equally in the calculations, so the statistics are not necessarily representative of the population of all cans of the same 20 brands consumed by Americans. 14. The range is 8 – 6.8 = 1.20%. The variance is s 2 =
13(739.57) - (97.9)2 = 0.19%2. 13 (13 - 1)
The standard deviation is s = 0.19%2 = 0.44%. The data are time-series data, but the measures of variation do not reveal anything about a trend consisting of a pattern of change over time. 15. The range is 50 - 0 = 50.0 cigarettes. The variance is s 2 =
50(4781) - (139)2 = 89.7 (cigarettes)2 50 (50 - 1)
The standard deviation is s = 89.7 = 9.5 cigarettes. Because the selected subjects report the number of cigarettes they smoke, it is very possible that the data are not at all accurate, so the results might not reflect the actual smoking behavior of California adults. 16. The range is 1.49 - 0.51 = 0.980 W/kg. The variance is s 2 =
11(16.4078) - (12.96)2 = 0.114 (W/kg)2 11(11 - 1)
The standard deviation is s = 0.114 = 0.337 W/kg. If concerned about radiation absorption, you might purchase the cell phone with the lowest absorption rate. 17. Systolic: x = 127.6 mmHg, s = 18.6 mmHg; The coefficient of variation is
18.6 mmHg ×100% = 14.6%. 127.6 mmHg
Diastolic: x = 73.6 mmHg, s = 12.5 mmHg; The coefficient of variation is
12.5 mmHg ×100% = 16.9%. 73.6 mmHg
The variation is roughly about the same. 18. All units are (1000 cells mL). White Blood Cell Counts: x = 6.73, s = 1.46; The coefficient of variation is
1.46 ×100% = 21.7%. 6.73
All units are (million cells mL). Red Blood Cell Counts: x = 4.57, s = 0.42; The coefficient of variation is White blood cell counts appear to vary more than red blood cell counts.
Copyright © 2024 Pearson Education, Inc.
0.42 ×100% = 9.2%. 4.57
19. All units are (1000 cells mL). Male: x = 6.17, s = 1.53; The coefficient of variation is
1.53 ×100% = 24.8%. 6.17
1.57 ×100% = 21.3%. 7.35 The variation is roughly about the same for females and males. Female: x = 7.35, s = 1.57; The coefficient of variation is
20. Single Line: x = 429.0 sec, s = 28.6 sec; The coefficient of variation is
28.6 sec ×100% = 6.7%. 429.0 sec 109.3 sec ×100% = 25.5%. 429.0 sec
Individual Lines: x = 429.0 sec, s = 109.3 sec; The coefficient of variation is The single line has much less variation than with individual lines. 21. Using all data: 2
Range = 112.0 mg/dL, s 2 = 238.3 (mg/dL ) , s = 15.4 mg/dL 2
Excluding possible outlier of 138: Range = 87.0 mg/dL, s 2 = 215.2 (mg/dL) , s = 14.7 mg/dL The measures of variation change, but not by substantial amounts. 2
22. Using all data: Range = 212.0 mg/dL, s 2 = 1238.3 (mg/dL ) , s = 35.2 mg/dL 2
Excluding possible outlier of 251: Range = 184.0 mg/dL, s 2 = 1179.0 (mg/dL) , s = 34.3 mg/dL Only the range changes by a fairly large amount. 23. Range = 3.10 F, s 2 = 0.39 ( F)2 , s = 0.62 F 24. Range = 4600.0 g, s 2 = 480,848.1 g 2 , s = 693.4 g; All of the weights end in 00, so they are all rounded to the nearest 100 g. This suggests that the results should be rounded as follows: Range = 4600.0 g, s 2 = 480,850 g 2 , s = 690 g. 25. The rule of thumb standard deviation is s »
138 - 26 = 28.0 mg/dL, which is far from s = 15.4 mg/dL found 4
by using all of the data. 26. The rule of thumb standard deviation is s »
251 - 39 = 53.0 mg/dL, which is far from s = 35.2 mg/dL found by 4
using all of the data. 27. The rule of thumb standard deviation is s »
99.6 - 96.5 = 0.78 F, which is not substantially different from 4
s = 0.62 F found by using all of the data. 4900 - 300 = 1150.0 g, which differs from s = 693.4 g found by 4 using all of the data by a considerable amount. Several of the lowest weights correspond to premature births, and they cause the range to be larger, with the resulting estimate being larger.
28. The rule of thumb standard deviation is s »
29. Significantly low values are less than or equal to 74.0 - 2 (12.5) = 49.0 beats per minute, and significantly high values are greater than or equal to 74.0 + 2 (12.5) = 99.0 beats per minute. A pulse rate of 44 beats per minute is significantly low.
Copyright © 2024 Pearson Education, Inc.
30. Significantly low values are less than or equal to 69.6 - 2 (11.3) = 47.0 beats per minute, and significantly high values are greater than or equal to 69.6 + 2 (11.3) = 92.2 beats per minute. A pulse rate of 50 beats per minute is neither significantly low or high. 31. Significantly low values are less than or equal to 77.32 - 2 (1.29) = 24.74 cm, and significantly high values are greater than or equal to 77.32 + 2 (1.29) = 29.90 cm. A foot length of 30 cm is significantly high. 32. Significantly low values are less than or equal to 98.20 - 2 (0.62) = 96.96 F, and significantly high values are greater than or equal to 98.20 + 2 (0.62) = 99.44 F. A body temperature of 100 F is significantly high.
(
103 1×49.52 + 33. s =
)
+ 1×649.52 - (1×49.5 +
2
+ 1×649.5)
103 (103 - 1)
= 68.4, which is somewhat far from the exact
value of 59.5.
(
147 25 ×149.52 + 34. s =
)
+ 2 ×549.52 - (25 ×149.5 +
2
+ 2 ×549.5)
147 (147 - 1)
= 69.5, which is not very far from the exact
value of 65.4. 2
2
2
(9 - 13.0) + (10 - 13.0) + (20 - 13.0) 9 + 10 + 20 = 13 cigarettes and s 2 = 35. a. m = = 24.7 cigarettes2 3 3 b. The nine possible samples of two values are the following: {(9, 9), (9, 10), (9, 20), (10, 9), (10, 10), (10, 20), (20, 9), (20, 10), (20,20)} which have the following corresponding sample variances:{0, 0.5, 60.5, 0.5, 0, 50, 60.5, 50, 0}, that have a mean of s 2 = 2.47 cigarettes 2 . c. The population variances of the nine samples above are {0, 0.25, 30.25, 0.25, 0, 25, 30.25, 25, 0} that have a mean of s 2 = 12.3 cigarettes2 . d. Part (b), because repeated samples result in variances that target the same value (24.7 cigarettes 2 ) as the population variance. Use division by n - 1. e. No. The mean of the sample variances (24.7 cigarettes 2 ) equals the population variance (24.7 cigarettes 2 ), but the mean of the sample standard deviations (3.5 cigarettes) does not equal the population standard deviation (5.0 cigarettes). 36. The mean absolute deviation of the population is 4.7 cigarettes. With repeated samplings of size 2, the nine different possible samples have mean absolute deviations of 0, 0, 0, 0.5, 0.5, 5, 5, 5.5, 5.5. With many such samples, the mean of those nine results is 2.4 cigarettes, showing that the sample mean absolute deviations tend to center about the value of 2.4 cigarettes instead of the mean absolute deviation of the population, which is 4.7 cigarettes. The sample mean deviations do not target the mean deviation of the population. This is not good. This indicates that a sample mean absolute deviation is not a good estimator of the mean absolute deviation of a population. Section 3-3: Measures of Relative Standing and Boxplots 1. Brady’s height is 2.66 standard deviations above the mean. 2. The waiting line represented by the bottom boxplot is better because the times have much less variation, so all customers have wait times that are closer together. 3. It appears that weights of U.S. Army males increased from 1988 to 2012. All values of the 5-number survey increased, except for the minimum. 4. 2.00 should be preferred, because it is 2.00 standard deviations above the mean and would correspond to the highest of the five different possible scores.
Copyright © 2024 Pearson Education, Inc.
5.
6.
7.
8.
9.
a. The difference is 98 - 70.2 = 27.8 mmHg. 27.8 = 2.48 standard deviations b. 11.2 c. z = 2.48 d. The diastolic blood pressure of 98 mmHg is significantly high. a. The difference is 71.3 - 40 = 31.3 mmHg. 31.3 = 2.61 standard deviations b. 12.0 c. z = - 2.61 d. The diastolic blood pressure of 40 mmHg is significantly low. a. The difference is 17.2 - 5 = 12.2. 12.2 = 0.91 standard deviation b. 13.4 c. z = - 0.91 d. The measurement of 5 is neither significantly low nor significantly high. a. The difference is 75 - 14.4 = 60.6. 60.6 = 5.05 standard deviations b. 12 c. z = 5.05 d. The measurement of 75 is significantly high. Significantly low scores are less than or equal to 21.1 - 2 (5.2) = 10.7, and significantly high scores are greater than or equal to 21.1 + 2 (5.2) = 31.5. Scores that are not significant are between 10.7 and 31.5.
10. Significantly low scores are less than or equal to 100 - 2 (15) = 70, and significantly high scores are greater than or equal to 100 + 2 (15) = 130. Scores that are not significant are between 70 and 130. 11. Significantly low knee heights are less than or equal to 21.4 - 2 (1.2) = 19.0 in., and significantly high knee heights are greater than or equal to 21.4 + 2 (1.2) = 23.8 in. Values that are not significant are between 19.0 in. and 23.8 in. 12. Significantly low hip breadths are less than or equal to 36.6 - 2 (2.5) = 31.6 cm, and significantly high hip breadths are greater than or equal to 36.6 + 2 (2.5) = 41.6 cm. Hip breadths that are not significant are between 31.6 cm and 41.6 cm.
272 - 174.12 54.6 - 174.12 = 13.79 and the shortest man’s z score is z = 7.10 7.10 = - 16.83. Chandra Bahadur Dangi has the more extreme height because his z score of –16.83 is farther from the mean than the z score of 13.79 for Robert Wadlow.
13. The tallest man’s z score is z =
14. The female has a higher pulse rate because her z score is z = the z score of z =
99 - 74.0 = 2.00, which is a higher number than 12.5
50 - 69.6 = - 1.73 for the male. 11.3
15. The male has a more extreme birth weight because his z score is z = number than the z score of z =
1500 - 3272.8 = - 2.69, which is a lower 660.2
1500 - 3037.1 = - 2.18 for the female. 706.3
Copyright © 2024 Pearson Education, Inc.
16. The egg put in the wren’s nest is more extreme since its z score is z = than the z score of z =
24.080 - 22.575 = 2.20 for the robin’s nest. 0.685
17. For 0.48 W/kg,
3 ×100 = 6, so it is the 6th percentile. 50
18. For 1.47 W/kg,
46 ×100 = 92, so it is the 92nd percentile. 50
19. For 1.10 W/kg,
20 ×100 = 40, so it is the 40th percentile. 50
20. For 98 W/kg,
22.916 - 21.130 = 2.40, which is larger 0.744
16 ×100 = 32, so it is the 32nd percentile. 50
21. L =
30 ×50 0.93 + 0.97 = 15, so P30 = = 0.95 W/kg (Tech: Minitab: 0.942 W/kg; Excel: 0.958 W/kg) 100 2
22. L =
25 ×50 = 12.5, so Q1 = P25 = 0.91 W/kg 100
23. L =
75 ×50 = 37.5, so Q3 = P75 = 1.28 W/kg (Tech: Minitab: 1.285 W/kg) 100
24. L =
40 ×50 1.09 + 1.10 = 20, so P40 = = 1.095 W/kg (Tech: Minitab: 1.094 W/kg; Excel: 1.096 W/kg) 100 2
25. L =
50 ×50 1.15 + 1.16 = 25, so P50 = = 1.155 W/kg 100 2
26. L =
75 ×50 = 37.5, so Q3 = P75 = 1.28 W/kg (Tech: Minitab: 1.285 W/kg) 100
27. L =
25 ×50 = 12.5, so Q1 = P25 = 0.91 W/kg 100
85 ×50 = 42.5, so P85 = 1.38 W/kg (Tech: Minitab: 1.387 W/kg; Excel: 1.3765 W/kg) 100 29. The 5-number summary is 1.5 mg, , 5.35 mg, 6.60 mg, 7.25 mg, 11 mg. 28. L =
30. The 5-number summary is 4.8 mg, 5.60 mg, 6.30 mg, 7.70 mg, 8.8 mg.
31. The 5-number summary is 128 mBq, 140.0 mBq, 150.0 mBq, 158.5 mBq, 172 mBq. (Tech: Minitab yields Q1 = 139.0 mBq and Q3 = 159.75 mBq. Excel yields Q1 = 141.0 mBq and Q3 = 157.25 mBq.)
32. The 5-number summary is 120 mmHg, 130.0 mmHg, 132.5 mmHg, 140.0 mmHg, 150 mmHg. (Tech: Minitab yields Q1 = 128.75 mmHg and Q3 = 140.75 mmHg.)
Copyright © 2024 Pearson Education, Inc.
33. The top boxplot represents males. Males appear to have slightly lower pulse rates than females. (Tech: For males, Minitab yields Q3 = 77.) Male Pulse
Female Pulse 34. The top boxplot represents the tar in king-size cigarettes, and the bottom boxplot represents the tar in menthol cigarettes. From the boxplots, it appears that there is substantially more tar in king-size cigarettes than in menthol cigarettes. (Tech: For the menthol cigarettes, Excel yields Q1 = 14.5 and Q3 = 15.5.) King-size
Menthol 35. The top boxplot represents the weights from ANSUR I 1988 and the bottom boxplot represents weights from ANSUR II 2012. It appears that the weights of male Army personnel increased somewhat from 1988 to 2012. (Tech: For ANSUR II 2012, Minitab yields Q1 = 75.575 and Q3 = 94.425. TI data: The values of the fivenumber summary for ANSUR I are 49.8, 70.7, 77.7, 86.0, 116.9, and for ANSUR II those values are 47.8, 75.2, 84.0, 93.2, 137.1.) ANSUR I ANSUR II 36. The low lead level group represented in the top boxplot has much more variation and the IQ scores tend to be higher than the IQ scores from the high lead level group. (Tech: For the low lead level group, Minitab yields Q1 = 84.75 and Q3 = 101.25. For the high lead level group, Minitab yields Q3 = 93.50. The results suggest that greater exposure to lead corresponds to lower full IQ scores (although a direct cause/effect link is not established). Low Lead
High Lead
Copyright © 2024 Pearson Education, Inc.
37. The top boxplot represents males. Males appear to have slightly lower pulse rates than females. The outliers for males are 40 beats per minute, 102 beats per minute, and 104 beats per minute. The outlier for females is 36 beats per minute. Males
Females
Quick Quiz
5.
8+ 7 +5+ 7 + 4 + 7 + 6 + 7 +8 +8+8+ 6 = 6.8 hr. 12 3. The modes are 7 hr and 8 hr. 7+7 = 7.0 hr. The median is 2 4. The variance is (1.3 hr)2 =1.7 hr 2 . The sleep time of 0 hr appears to be an outlier because it is substantially less than all the other sleep times.
6.
z=
7.
About 75% or 0.75 (80) = 60 sleep times are less than Q3.
8.
minimum, first quartile Q1, second quartile Q2 (or median), third quartile Q3, maximum
1. 2.
The sample mean is x =
5 - 6.3 = - 0.93; No, the sleep time of 5 hr is not significantly low, so it is not an outlier. 1.4
10. x , m, s, s , s 2, s 2
10 - 4 = 1.5 hr 4 Review Exercises 1. The differences are given in the table below. 9.
s»
Reported Measured
68 67.9
71 69.9
63 64.9
70 68.3
71 70.3
60 60.6
65 64.5
64 67
54 55.6
63 74.2
66 65
72 70.8
Difference
0.1
1.1
–1.9
1.7
0.7
–0.6
0.5
–3
–1.6 –11.2
1
1.2
0.1 + 1.1 + (- 1.9) + 1.7 + 0.7 + (- 0.6) + 0.5 + (–3) + (- 1.6) + (- 11.2) + 1 + 1.2 = - 1.00 in. 12 0.1 + 0.5 = 0.30 in. b. The median is 2 c. There is no mode. - 11.2 + 1.7 = - 4.75 in. d. The midrange is 2 e. The range is 1.7 - (- 11.2) = 12.90 in. a. The mean is x =
f.
s=
(0.1 - (- 1.0))2 + (1.1 - (- 1.0))2 +
2
2
+ (1 - (- 1.0) ) + (1.2 - ( - 1.0) )
12 - 1
= 3.52 in.
g. s 2 = 3.522 = 12.39 in.2
(- 1.9) + (- 1.6) 25 ×12 = 3, so Q1 = = - 1.75 in. (Tech: Minitab: Q1 = - 1.83 in. Excel: Q1 = - 1.675 in.) 100 2 75 ×12 1 + 1.1 L= = 9, so Q3 = = 1.05 in. (Tech: Minitab: Q3 = 1.07 in. Excel: Q3 = 1.025 in.) 100 2
h. L = i.
Copyright © 2024 Pearson Education, Inc.
2.
3.
The difference of –11.2 in. appears to be an outlier. If that outlier is excluded, the mean changes from –1.00 in. to –0.07 in., the median changes from 0.30 in. to 0.50 in., and the standard deviation changes from 3.52 in. to 1.51 in. The outlier has a strong effect on the mean and standard deviation, but very little effect on the median.
- 11.2 - (- 1.00) = - 2.90; The difference of –11.2 in. is significantly low (because its z score is less than or 3.52 equal to –2). z=
4.
The 5-number summary is –11.2 in., –1.75 in., 0.30 in., 1.05 in., 1.70 in. (Tech: Minitab yields Q1 = - 1.83 in. and Q3 = 1.07 in. Excel yields Q1 = - 1.675 in. and Q3 = 1.025 in.)
5.
The mean is x =
6.
Significantly low scores are less than or equal to 504.7 - 2 (9.4) = 485.9, and significantly high scores are
12 + 14 + 22 + 27 + 40 = 23.0. The numbers don’t measure or count anything. They are used as 5 replacements for the names of the categories, so the numbers are at the nominal level of measurement. In this case the mean is a meaningless statistic. greater than or equal to 504.7 + 2 (9.4) = 523.5.
7. 8.
The minimum value is 119 mm, the first quartile is 128 mm, the second quartile (or median) is 131 mm, the third quartile is 135 mm, and the maximum value is 141 mm. The outlier is 646. The mean and standard deviation with the outlier included are x = 267.8 and s = 131.6. Those statistics with the outlier excluded are x = 230.0 and s = 42.0. Both statistics changed by a substantial amount, so here the outlier has a very strong effect on the mean and standard deviation.
9.
Significantly low heights are 97.5 - 2 (6.9) = 83.7 cm or less; significantly high heights are
97.5 + 2 (6.9) = 111.3 cm or greater. The height of 87.8 cm is not significant, so the physician should not be concerned. 3400 - 3272.8 3200 - 3037.1 = 0.19 . The female z score is z = = 0.23 . The female has 660.2 706.3 the larger relative birth weight because the female has the larger z score. Cumulative Review Exercises 1. a. quantitative b. ratio level of measurement c. continuous d. sample e. statistic 10. The male z score is z =
2.
a. The mean is x = b. The median is
0.72 + 0.9 + 0.84 + 0.68 + 0.84 +
+ 0.95 + 0.86 + 0.88 + 0.85 + 0.87 = 0.830 mm. 20
0.84 + 0.84 = 0.840 mm. 2 2
c. The standard deviation is s =
20 (13.8901) - (16.59) = 0.082 mm. 20 (20 - 1)
2
d. The variance is s 2 = (0.082) = 0.007 mm 2 . e. The range is 0.95 - 0.64 = 0.310 mm. 3.
For 0.76 mm,
4 ×100 = 20, so it is the 20th percentile. 20
Copyright © 2024 Pearson Education, Inc.
4. Thorax Length (mm) 0.60–0.65 0.66–0.71 0.72–0.77 0.78–0.83 0.84–0.89 0.90–0.95
Frequency 1 1 3 1 9 5
5.
6. 7. 8.
The vertical scale does not begin at 0, so the differences among the outcomes are exaggerated. Because the distribution is roughly bell-shaped, it does appear that the sample data are from a population with a normal distribution. Based on the scatterplot, there does appear to be a correlation between heights of fathers and heights of their first sons. Because the points are not very close to a straight-line pattern, the correlation does not appear to be very strong.
Copyright © 2024 Pearson Education, Inc.
Chapter 4: Probability Section 4-1: Basic Concepts of Probability 1. The probability of randomly selecting someone with blue eyes is 0.35. 2. The probability of a baby being male is 1/2 or 0.5. 3.
P( A) = 1 - P( A) = 1 - 0.512 = 0.488
4.
The answers vary, but a high answer in the neighborhood of 0.999 is reasonable.
5.
0, 3 5, 1, 0.135
11. 1 4 , or 0.25
6.
1 5 , or 0.2
12. 0.292
7.
1 10 , or 0.1
13. 1 10 , or 0.1
8.
{bb, bg, gb, gg}
14. 1 2 , or 0.5
9.
1 2 , or 0.5
15. 0 16. 1
10. 1 5 , or 0.2 17.
239 239 = , or 0.821; Yes, the technique appears to be effective. 239 + 52 291
18.
117 117 = , or 0.159; Headaches do appear to be a substantial adverse reaction. 117 + 617 734
19.
428 , or 0.738; Yes, it is reasonable. 580 c. He already knew. d. 0
20. a. 1 365
b. yes 3732 311 = , or 0.730; Yes, it is likely for someone to use a social networking site. 21. 1380 + 3732 426 22.
83, 600 418 = , or 0.0160; Yes, in a passenger car crash, a rollover is unlikely. 83, 600 + 5,127, 400 26, 055
23. a. brown/brown, brown/blue, blue/brown, blue/blue b. 1 4 c. 3 4 24. a. brown/blue, brown/blue, blue/blue, blue/blue b. 1 2 c. 1 2 25. 3 8, or 0.375
26. 3 8, or 0.375
27. {mmmm, fmmm, mfmm, mmfm, mmmf, ffmm, fmfm, fmmf, mffm, mfmf, mmff, mfff, fmff, ffmf, fffm, ffff}; 4 16 = 1 4, or 0.25 28. 2 16 = 1 8, or 0.125 29. 53 females is neither significantly low nor significantly high. 30. 35 females is significantly low. 31. 75 females is significantly high. 32. 48 females is neither significantly low nor significantly high. 33. Because the probability of 0.364 is not small (less than or equal to 0.05), it appears that getting a result of 91 or higher can easily occur. The result of 91 or higher is not significantly high (nor is it significantly low). 34. Because the probability of 0.000 is small, and because 856 is higher than expected, it appears that the result of 856 is significantly high.
Copyright © 2024 Pearson Education, Inc.
35. The probability of 0.873 is not small, suggesting that 6062 deaths is not significantly low (nor is it significantly high). (Because there were more deaths in the week before Thanksgiving than in the week after, the claim of postponing death does not appear to be valid.) 36. The probability of 0.265 is not small (less than or equal to 0.05), so the result of 152 yellow peas could easily occur by chance. The result of 152 yellow peas is not significantly high (nor is it significantly low). 37. If pregnant women have no ability to predict the sex of their babies, then among 104 predictions, we expect about half of them (or 52) to be correct. The 57 correct predictions is greater than 52. The high probability of 0.189 is greater than 0.05, so 57 correct predictions is not significantly high (and it is not significantly low). It does not appear that pregnant women can correctly predict the sex of their babies. 38. Because the probability of 0.00000000978 is so low (less than or equal to 0.05), the result of 148 is significantly low. This suggests that the claim that the rate is less than 50% appears to be supported. 39. Because the probability of getting 604 or more respondents who have made new friends online is 0.00000306 (less than or equal to 0.05), it appears that 604 is significantly high. This suggests that the true rate is not 50%. 40. Because 0.000235 is so low (less than or equal to 0.05), and because 37 deaths is greater than half of the 49 deaths, it appears that 37 male deaths is significantly high. It appears that male selfie deaths are more likely than female selfie deaths. 41. In the following, the first letter represents the chromosome contributed by the father and the second letter represents the chromosome contributed by the mother. Let X1 and X2 represent the possible X chromosomes contributed by the mother. a. 0; The possible outcomes are {xX1, xX2, YX1, YX2}, neither son will have the disease. b. 0; The possible outcomes are {xX1, xX2, YX1, YX2}, neither daughter will have the disease. c. 1 2, or 0.5; The possible outcomes are {Xx1, XX2, Yx1, YX2}, one of the two sons will have the disease. d. 0; The possible outcomes are {Xx1, XX2, Yx1, YX2}, neither daughter will have the disease. Section 4-2: Addition Rule and Multiplication Rule 1.
P( B) represents the probability that when an adult is randomly selected, the person selected has blue eyes.
P( B ) represents the probability that when an adult is randomly selected, the person selected does not have blue eyes. 2.
P ( M | B) represents the probability of getting a male, given that someone with blue eyes has been selected.
Because P ( B | M ) is the probability of selecting someone with blue eyes given that the selected person is a male, P ( M | B) is not the same as P( B | M ). 3.
4.
Because the selections are made without replacement, the events are dependent. Because the sample size of 1068 is less than 5% of the population size of 30,488,983, the selections can be treated as being independent (based on the 5% guideline for cumbersome calculations). Because the probability of getting a result of 47 or lower is only 0.000000987, it appears that 47 is significantly low. It does appear the claim is supported by the data.
6.
580 - 428 152 = , or 0.262 580 580 1 - 0.0025 = 0.9975, or 99.75%
8.
P( I ) denotes the probability of screening a driver and finding that he or she is not intoxicated, and
5.
7.
617 617 = , or 0.841 617 + 117 734
P( I ) = 0.99112, or 0.991 when rounded. Use the following table for Exercises 9–20
Texted While Driving No Texting While Driving Total
Drove When Drinking Alcohol? Yes No 731 3054 156 4564 887 7618
Copyright © 2024 Pearson Education, Inc.
Total 3785 4720 8505