CONTENTS
Part One: Gathering and Exploring Data Chapter 1: Statistics: The Art and Science of Learning from Data Section 1.1: Using Data to Answer Statistical Questions ................................................1 Section 1.2: Sample Versus Population ...........................................................................1 Section 1.3: Using Calculators and Computers ...............................................................3 Chapter Problems: Practicing the Basics .........................................................................4 Chapter Problems: Concepts and Investigations..............................................................5 Chapter Problems: Student Activities..............................................................................5
Chapter 2: Exploring Data with Graphs and Numerical Summaries Section 2.1: Different Types of Data ...............................................................................7 Section 2.2: Graphical Summaries of Data......................................................................8 Section 2.3: Measuring the Center of Quantitative Data ...............................................14 Section 2.4: Measuring the Variability of Quantitative Data ........................................16 Section 2.5: Using Measures of Position to Describe Variability .................................20 Section 2.6: Recognizing and Avoiding Misuses of Graphical Summaries ..................25 Chapter Problems: Practicing the Basics .......................................................................26 Chapter Problems: Concepts and Investigations............................................................35 Chapter Problems: Student Activities............................................................................39
Chapter 3: Association: Contingency, Correlation, and Regression Section 3.1: The Association Between Two Categorical Variables ..............................41 Section 3.2: The Association Between Two Quantitative Variables.............................45 Section 3.3: Predicting the Outcome of a Variable........................................................49 Section 3.4: Cautions in Analyzing Associations ..........................................................55 Chapter Problems: Practicing the Basics .......................................................................62 Chapter Problems: Concepts and Investigations............................................................72 Chapter Problems: Student Activities............................................................................75
Chapter 4: Gathering Data Section 4.1: Experimental and Observational Studies...................................................77 Section 4.2: Good and Poor Ways to Sample ................................................................79 Section 4.3: Good and Poor Ways to Experiment .........................................................81 Section 4.4: Other Ways to Conduct Experimental and Nonexperimental Studies.......82 Chapter Problems: Practicing the Basics .......................................................................84 Chapter Problems: Concepts and Investigations............................................................89 Chapter Problems: Student Activities............................................................................91
Part Two: Probability, Probability Distributions, and Sampling Distributions Chapter 5: Probability in Our Daily Lives Section 5.1: How Probability Quantifies Randomness..................................................93 Section 5.2: Finding Probabilities..................................................................................94 Section 5.3: Conditional Probability..............................................................................98 Section 5.4: Applying the Probability Rules................................................................101 Chapter Problems: Practicing the Basics .....................................................................105 Chapter Problems: Concepts and Investigations..........................................................112 Chapter Problems: Student Activities..........................................................................115
Chapter 6: Probability Distributions Section 6.1: Summarizing Possible Outcomes and Their Probabilities.......................117 Section 6.2: Probabilities for Bell-Shaped Distributions.............................................120 Section 6.3: Probabilities When Each Observation Has Two Possible Outcomes ......124 Chapter Problems: Practicing the Basics .....................................................................128 Chapter Problems: Concepts and Investigations..........................................................135 Chapter Problems: Student Activities..........................................................................138
Chapter 7: Sampling Distributions Section 7.1: How Sample Proportions Vary Around the Population Proportion ........139 Section 7.2: How Sample Means Vary Around the Population Mean ........................142 Chapter Problems: Practicing the Basics .....................................................................145 Chapter Problems: Concepts and Investigations..........................................................148 Chapter Problems: Student Activities..........................................................................151
Part Three: Inferential Statistics Chapter 8: Statistical Inference: Confidence Intervals Section 8.1: Point and Interval Estimates of Population Parameters...........................153 Section 8.2: Constructing a Confidence Interval to Estimate a Population Proportion ...............................................................................................................154 Section 8.3: Constructing a Confidence Interval to Estimate a Population Mean.......157 Section 8.4: Choosing the Sample Size for a Study.....................................................160 Section 8.5: Using Computers to Make New Estimation Methods Possible...............162 Chapter Problems: Practicing the Basics .....................................................................163 Chapter Problems: Concepts and Investigations..........................................................170 Chapter Problems: Student Activities..........................................................................174
Chapter 9: Statistical Inference: Significance Tests About Hypotheses Section 9.1: Steps for Performing a Significance Test ................................................175 Section 9.2: Significance Tests About Proportions .....................................................175 Section 9.3: Significance Tests About Means .............................................................179 Section 9.4: Decisions and Types of Errors in Significance Tests ..............................182 Section 9.5: Limitations of Significance Tests ............................................................184 Section 9.6: The Likelihood of a Type II Error and the Power of a Test ....................185 Chapter Problems: Practicing the Basics .....................................................................187 Chapter Problems: Concepts and Investigations..........................................................193 Chapter Problems: Student Activities..........................................................................196
Chapter 10: Comparing Two Groups Section 10.1: Categorical Response: Comparing Two Proportions............................197 Section 10.2: Quantitative Response: Comparing Two Means ..................................200 Section 10.3: Other Ways of Comparing Means, Including a Permutation Test.........206 Section 10.4: Analyzing Dependent Samples..............................................................209 Section 10.5: Adjusting for the Effects of Other Variables .........................................213 Chapter Problems: Practicing the Basics .....................................................................214 Chapter Problems: Concepts and Investigations..........................................................223 Chapter Problems: Student Activities..........................................................................226
Part Four: Analyzing Association and Extended Statistical Methods Chapter 11: Analyzing the Association Between Categorical Variables Section 11.1: Independence and Dependence (Association) .......................................227 Section 11.2: Testing Categorical Variables for Independence...................................229 Section 11.3: Determining the Strength of the Association.........................................232 Section 11.4: Using Residuals to Reveal the Pattern of Association...........................234 Section 11.5: Fisher’s Exact and Permutation Tests....................................................235 Chapter Problems: Practicing the Basics .....................................................................236 Chapter Problems: Concepts and Investigations..........................................................241 Chapter Problems: Student Activities..........................................................................244
Chapter 12: Analyzing the Association Between Quantitative Variables: Regression Analysis Section 12.1: Modeling How Two Variables Are Related ..........................................245 Section 12.2: Inference About Model Parameters and the Association.......................247 Section 12.3: Describing the Strength of Association .................................................251 Section 12.4: How the Data Vary Around the Regression Line ..................................254 Section 12.5: Exponential Regression: A Model for Nonlinearity .............................256 Chapter Problems: Practicing the Basics .....................................................................258 Chapter Problems: Concepts and Investigations..........................................................261 Chapter Problems: Student Activities..........................................................................267
Chapter 13: Multiple Regression Section 13.1: Using Several Variables to Predict a Response .....................................269 Section 13.2: Extending the Correlation and R2 for Multiple Regression ...................273 Section 13.3: Inferences Using Multiple Regression...................................................274 Section 13.4: Checking a Regression Model Using Residual Plots.............................276 Section 13.5: Regression and Categorical Predictors ..................................................280 Section 13.6: Modeling a Categorical Response .........................................................282 Chapter Problems: Practicing the Basics .....................................................................285 Chapter Problems: Concepts and Investigations..........................................................289 Chapter Problems: Student Activities..........................................................................292
Chapter 14: Comparing Groups: Analysis of Variance Methods Section 14.1: One-Way ANOVA: Comparing Several Means...................................293 Section 14.2: Estimating Differences in Groups for a Single Factor...........................296 Section 14.3: Two-Way ANOVA................................................................................298 Chapter Problems: Practicing the Basics .....................................................................301 Chapter Problems: Concepts and Investigations..........................................................306 Chapter Problems: Student Activities..........................................................................309
Chapter 15: Nonparametric Statistics Section 15.1: Compare Two Groups by Ranking ........................................................311 Section 15.2: Nonparametric Methods for Several Groups and for Matched Pairs.....313 Chapter Problems: Practicing the Basics .....................................................................315 Chapter Problems: Concepts and Investigations..........................................................318
Chapter 1: The Art and Science of Learning from Data 1
Section 1.1 Using Data to Answer Statistical Questions 1.1 Aspirin and heart attacks a) Aspects of the study that have to do with design include the sample of 22,000 physicians, the randomization of the halves of the sample to the two groups (aspirin and placebo), and the plan to obtain percentages of each group that have heart attacks. b) Aspects having to do with description include the actual percentages of the people in the sample who have heart attacks (i.e., 0.9% for those taking aspirin and 1.7% for those taking placebo). c) Aspects that have to do with inference include the use of statistical methods to conclude that taking aspirin reduces the risk of having a heart attack. 1.2 Poverty and race a) The aspects referring to description are the percentages of the 68,000 households (18.0% of whites, 37.5% of blacks, and 13.4% of Asians) who had incomes below the poverty level. b) The statistical method that predicted that the percentage of all black households in the United States that had income below the poverty level was between 35.6% and 39.4% is an example of inference. 1.3 GSS and heaven Yes, definitely: 64.6%; Yes, probably: 20.8%; No, probably not: 8.7%; No, definitely not: 5.9% 1.4 GSS and heaven and hell a) Yes, definitely: 64.3%; Yes, probably: 20.8%; No, probably not: 8.8%; No, definitely not: 6.0% b) Yes, definitely: 52.6%; Yes, probably: 20.3%; No, probably not: 14.8%; No, definitely not: 12.3%; The percentage of “yes, definitely” responses was higher for belief in heaven in 2008. 1.5 GSS for subject you pick The results for this item will be different depending on the topic that you chose.
Section 1.2 Sample Versus Population 1.6 Description and inference a) With description, we are summarizing a group of numbers. We can use description with either samples or populations. With inferences, we use data from samples to make conclusions or predictions about populations. For example, if we ask a sample of adults how many pets they own, and take the mean number of pets, that number is a description. If we use that number to predict the mean number of pets owned by the whole population, the predicted mean (or the predicted range for the mean) would be an inference. b) Descriptive statistics would be useful to summarize data from a population. With a census, it would be unwieldy to examine everyone’s ages, for example, but it would be useful to know a mean age. Inferential statistics are not needed, however, because we already have information about the population; we don’t need to predict it. 1.7 Censorship a) The sample is the 3077 people who responded. b) The population is all adults in the United States. c) The statistic is 23% of respondents said antireligious books should be removed. 1.8 Concerned about global warming? a) The sample is the set of polled Floridians. The population is the set of all adult Florida residents. b) The percentages quoted are statistics since they are summaries of the sample. 1.9 Graduate school information a) Each student in the program is a subject. b) The sample is the students identified for an interview from the given program. c) The population is all students in the program. 1.10 Is globalization good? a) The samples are those people selected from each country to participate in the survey. The populations are all adults in Africa and all adults in North America. b) These are statistics because they represent a summary of the sample data. Copyright © 2017 Pearson Education, Inc.
2 Statistics: The Art and Science of Learning from Data, 4th edition 1.11 Graduating seniors’ salaries a) These are descriptive statistics. They are summarizing data from a population – all graduating seniors at a given school. b) These analyses summarize data on a population – all graduating seniors at a given school; thus, the numerical summaries are best characterized as parameters. 1.12 At what age did women marry? a) The mean age of 24.1 years for this sample is descriptive. b) The historian estimates the age for the whole population of brides in early 19th century New England, estimating the average age to fall between 23.5 and 24.7. This is inferential. c) The inference refers to the population of all New England brides between the years of 1800 and 1820. d) The average of 24.1 years is based on a sample and is therefore a statistic. 1.13 Age pyramids as descriptive statistics a) The bar graph for 1750 shows shorter and shorter bars as age increases indicating that there were few Swedish people who were old in 1750. b) For every age range, the bars are much longer for both men and women in 2010 than in 1750. c) The bars for women in their 70’s and 80’s in 2010 are longer than those for men of the same age in the same year. d) The first manned space flight took place in 1961 so that people born during this era would fall in the 45–49 year old category. This is the largest five-year group for both men and women. 1.14 Gallup polls Responses to this exercise will differ depending on the studies that students choose. a) The descriptive statistic will be a summary of data, without any prediction or population estimate. It might be a mean rating for a given attitude, for example. b) The inferential statistical analysis will have some kind of prediction or estimation; for example, the inferential statistic might include the margin of error for a mean, indicating that the population mean likely falls somewhere in a given range. 1.15 National service a) Yes, the populations are the same in the two studies. For both, it’s all students at your school. b) It is very unlikely that you will choose the same 20 students. c) Although it is most likely that the sample proportions will not be the same, they should be close to each other. 1.16 Samples vary less with more data a) It would be more surprising to flip a coin 500 times and observe all heads. b) As the sample size increases, the amount by which sample proportions tend to vary decreases. The estimates from larger samples, therefore, tend to be more accurate than estimates from smaller samples. When the coin is flipped just 5 times, it’s easy to see that we could get a sample with all heads. However, when the number of flips is increased to 500, it is much more likely that the sample proportion is near the population proportion of 0.5. It would be extremely unlikely to observe very few heads or almost all heads in 500 flips of a fair coin. 1.17 Comparing polls a)
1 n 100% 1 1000 100% 0.0316 100% 3.16% (rounds to 3.2%)
b) The first four polls are all within the margin of error; however, Rand favored Obama slightly, and Fox underestimated Obama’s margin. Generally, the polls are fairly accurate. 1.18 Margin of error and n a)
1 n 100% 1 100 100% 0.1 100% 10%, which suggests that between 50% and 70% of
Americans favored offshore drilling as a means of reducing U.S. dependence on foreign oil. b)
1 n 100% 1 400 100% 0.05 100% 5%, which suggests that between 55% and 65% of
Americans favored offshore drilling as a means of reducing U.S. dependence on foreign oil.
Copyright © 2017 Pearson Education, Inc.
Chapter 1: The Art and Science of Learning from Data 3 1.18 (continued) c)
1 n 100% 1 1600 100% 0.025 100% 2.5%, which suggests that between 57.5% and
62.5% of Americans favored offshore drilling as a means of reducing U.S. dependence on foreign oil. As n increases, the sample becomes a more accurate reflection of the population, and the margin of error decreases. 1.19 Smoking cessation a) iii b) Yes. Because the employees were assigned to treatments randomly, the study provides us with convincing evidence that the difference was due to the effect of the financial incentive.
Section 1.3 Using Calculators and Computers 1.20 Data file for friends The results for this exercise will be different for each person who does it. The data files, however, should all look like this: Friend Characteristic 1 Characteristic 2 1 2 3 4 For each friend, you’ll have a number or label under characteristics 1 and 2. For example, if you asked each friend for gender and hours of exercise per week, the first friend might have m (for male) under Characteristic 1, and 6 (for hours exercised per week) under Characteristic 2. 1.21 Shopping sales data file Customer Clothes Sporting goods Books Music CDs 1 $49 $0 $0 $16 2 $0 $0 $0 $0 3 $0 $0 $0 $0 4 $0 $0 $92 $0 5 $0 $0 $0 $0 1.22 Sample with caution A sample of individuals with children who read the Ann Landers column is not a random sample of individuals with children because every member of the population does not have the same chance of being in the sample. Many individuals with children may not read Ann Landers while others who do read the column may choose not to participate in the survey. The feelings of those who choose to participate usually are not representative of the general population. In general, one should not rely much on the information contained in such samples. 1.23 Create a data file with software Your MINITAB data (from Exercise 1.21) will be in the following format, although it will reside in the cells of the MINITAB worksheet. Customer Clothes Sporting goods Books Music CDs 1 49 0 0 16 2 0 0 0 0 3 0 0 0 0 4 0 0 92 0 5 0 0 0 0 1.24 Use a data file with software See solution for Exercise 1.21 for format of data in MINITAB.
Copyright © 2017 Pearson Education, Inc.
4 Statistics: The Art and Science of Learning from Data, 4th edition 1.25 Simulate with the Sampling Distribution for the Sample Proportion web app a) These will be different each time this exercise is completed. b) Regardless of the specific graphs constructed in (a), you will see that the amounts by which sample percentages tend to vary get smaller as the sample size n gets larger. c) The practical implication of this is that larger sample sizes tend to provide more accurate estimates of the true population percentage value. 1.26 Margin of error a) Answers will vary. b)
1 n 100% 1 1000 100% 0.0316 100% 3.16% (rounds to 3%)
c) Answers will vary. d) Answers will vary. 1.27 Ebola outbreaks The answer to this problem is based on a random process. This leads to potentially different answers each time it is performed. The binomial distribution (see Section 6.3) says that 14 or fewer people who died should occur in only about 1 in 100 simulations, so most students will likely not see any of these situations.
Chapter Problems: Practicing the Basics 1.28 UW Student survey a) The population is the entire UW student body of 40,858. The sample is the 100 students who were asked to complete the questionnaire. b) This value would not necessarily equal the value for the entire population of UW students. It is quite possible that the sample of 100 is not exactly representative of the whole student body. This percentage is only an estimate of the percentage of all students who would respond this way. It is unlikely that any single sample of 100 would have a percentage that was exactly the percentage of the entire population. c) The numerical summary is a sample statistic because it only summarizes for a sample, not for a population. 1.29 Euthanasia a) The population is all American adults. b) The sample data are summarized by a proportion, 0.598. c) The population proportion who would commit suicide. 1.30 Sleep disorders among college students It is very likely that between 25% and 29% of students are at risk for at least one sleep disorder. 1.31 Breaking down Brown versus Whitman a) The results summarize sample data because not every voter in the 2010 California gubernatorial election was polled. b) The percentages reported here are descriptive in that they describe the exact percentages of the sample polled who were Democrat and voted for Brown, who were Republican and voted for Brown and who were Independent and voted for Brown. c) The inferential aspect of this analysis is that the exit poll results were used to predict what percentage of each of the three parties (Democrat, Republican and Independent) voted for Brown in the 2010 California gubernatorial election. The margins of error give a likely range for the population percentages for each of the three parties. 1.32 Online learning a) The sample is the 100 students surveyed. The population is all students in this school. b) (i) Descriptive statistics would give us information about the preferences of the 100 students in the sample. (ii) Inferential statistics allow us to draw a conclusion about the preferences of the student body.
Copyright © 2017 Pearson Education, Inc.
Chapter 1: The Art and Science of Learning from Data 5 1.33 Marketing study For the study on the marketing of digital media, the population is all Facebook users, and the sample is the 1000 Facebook users to whom the ad was displayed. Example 5 suggests that we might determine that the average sales per person equaled $0.90. This would be a descriptive statistic in that it describes the average sales per person in the sample of 1000 potential customers. If one were to use this information to make a prediction about the population, this would be an inferential statistic. 1.34 Support of labor unions a)
1 n 100% 1 1540 100% 0.025 100% 2.5% .
b) Between 50.5% and 55.5% c) ii. Inferential statistics 1.35 Multiple choice: Use of inferential statistics? The best answer is (c). 1.36 True or false? False: We often want to describe the sample AND make inferences about the population.
Chapter Problems: Concepts and Investigations 1.37 Statistics in the news If your article has numbers that summarize for a given group (sample or population), it’s using descriptive statistics. If it uses numbers from a sample to predict something about a population, it’s using inferential statistics. 1.38 What is statistics? Answers will vary. 1.39 Surprising suicide data? The likelihood of getting this result is extremely small. 1.40 Create a data file See solution for Exercise 1.23 for format of data in MINITAB.
Chapter Problems: Student Activities 1.41 Getting to know the class Answers will vary.
Copyright © 2017 Pearson Education, Inc.
Chapter 2: Exploring Data with Graphs and Numerical Summaries
7
Section 2.1 Different Types of Data 2.1 Categorical/quantitative difference a) Categorical variables are those in which observations belong to one of a set of categories, whereas quantitative variables are those on which observations are numerical. b) An example of a categorical variable is religion. An example of a quantitative variable is temperature. 2.2 U.S. married-couple households The variable summarized is categorical. The variable is type of U.S. married-couple households, and there are four types: traditional, dual-income with children, dual-income with no children, and other. These types are the categories. 2.3 Identify the variable type a) quantitative c) categorical b) categorical d) quantitative 2.4 Categorical or quantitative? a) categorical c) categorical b) quantitative d) quantitative 2.5 Discrete/continuous a) A discrete variable is a quantitative variable for which the possible values are separate values such as 0, 1, 2, …. A continuous variable is a quantitative variable for which the possible values form an interval. b) Example of a discrete variable: the number of children in a family (a given family can’t have 2.43 children). Example of a continuous variable: temperature (we can have a temperature of 48.659). 2.6 Discrete or continuous? a) continuous c) continuous b) discrete d) discrete 2.7 Discrete or continuous 2 a) continuous c) discrete b) discrete d) continuous 2.8 Number of children a) The variable, number of children, is quantitative. b) The variable, number of children, is discrete. c) No. children 0 1 2 3 4 5 6 7 8+ Count 521 323 524 344 160 77 30 19 22 Proportion 0.258 0.160 0.259 0.170 0.079 0.038 0.015 0.009 0.011 Percentage 25.8 16.0 25.9 17.0 7.9 3.8 1.5 0.9 1.1 2.9 Fatal Shark Attacks a) Florida Location Count 2 Proportion 0.032 Percentage 3.2
Hawaii 2 0.032 3.2
California 4 0.063 6.3
Australia 15 0.238 23.8
Reunion Island Brazil Bahamas Other Location Count 6 4 6 11 Proportion 0.095 0.063 0.095 0.175 Percentage 9.5 6.3 9.5 17.5 b) Australia is the modal category. c) The regions with most frequent fatal shark attacks are Australia and South Africa.
Copyright © 2017 Pearson Education, Inc.
South Africa 13 0.206 20.6
8 Statistics: The Art and Science of Learning from Data, 4th edition
Section 2.2 Graphical Summaries of Data 2.10 Generating Electricity a) Electricity Generation 40
Percent
30 20 10 0 oa C
l
a ur at N
as lG
uc N
ar
er ow
le le ab ab w w e e n en re yd H rR on e N th er O th O Source
le
p ro
b) Sketching a bar chart would be easier. Sketching the precise areas corresponding to the percentages is more challenging in a pie chart. c) It is straightforward to judge the relative sizes when comparing the bars corresponding to the percentages. d) Coal is the modal category. 2.11 What do alligators eat? a) Primary food choice is categorical. b) The modal category is “fish.” c) Approximately 43% of alligators ate fish as their primary food choice. d) This is an example of a Pareto chart, a chart that is organized from most to least frequent choice. 2.12 Weather stations a) The slices of the pie portray categories of a variable (i.e., regions). b) The first number is the frequency, the number of weather stations in a given region. The second number is the percentage of all weather stations that are in this region. c) It is easier to identify the modal category using a bar graph than using a pie chart because we can more easily compare the heights of bars than the slices of a piece of pie. For example, in this case, the slices for Midwest and West look very similar in size, but it would be clear from a bar graph that West was taller in height than Midwest. 2.13 France is most popular holiday spot a) Country visited is categorical. b) A Pareto chart would make more sense because it allows the viewer to easily locate the categories with the highest and lowest frequencies. c) A dot plot or stem-and-leaf plot do not make sense because the data are categorical; these two types of plots are used with quantitative data (and also with data that have relatively few observations).
Copyright © 2017 Pearson Education, Inc.
Chapter 2: Exploring Data with Graphs and Numerical Summaries 2.14 Pareto chart for fatal shark attacks (i) Alphabetically
(ii) Pareto chart Shark Fatalities 25
20
20 Percent
Percent
Shark Fatalities 25
15 10
15 10
5
5
0
0
l a i h er an d r ica zi nia r ida lia as l f tr a am Bra if or lo Haw Ot Is s A F l h u n th A io Ca Ba u n u So Re
i l ai z i nia id a lia ica er a s nd tr a A f r Oth am Isla Br a if or lor aw s F l h H u th n A Ca Ba nio u u So Re
Location
Location
With a Pareto chart, it is straightforward to identify the few regions with the largest number of fatal shark attacks. 2.15 Sugar dot plot a) The minimum sugar value is zero grams, and the maximum is 18 grams. b) The sugar outcome that occurs most frequently is called the mode. For this data set there are five modes: three, four, eleven, twelve and fourteen grams. 2.16 Spring break hotel prices a) 1 | 24677999 2 | 133445 3 | 1338 b) 1 | 24 1 | 677999 2 | 13344 2|5 3 | 133 3|8 The plot with split stems gives a clearer picture of the shape of the distribution. c) Most hotels charge between $150 and $250 per night, with a few charging more. The distribution of prices is right-skewed. Histogram of Hotel Price 6
Frequency
5 4 3 2 1 0
100
150
200 250 300 Hotel Price
350
Copyright © 2017 Pearson Education, Inc.
400
9
10 Statistics: The Art and Science of Learning from Data, 4th edition 2.17 Graphing exam scores a) There are 33 students in the class; the minimum score is 65 and the maximum is 98. b) Dotplot of Exam Scores
65
70
75
80 85 Exam Scores
90
95
c) Histogram of Exam Scores 14 12
Frequency
10 8 6 4 2 0
60
70
80 90 Exam Scores
100
2.18 Fertility rates a) 1 | 3333445677778899 2 | 04 A disadvantage of this plot is that it is too compact making it difficult to visualize where the data fall. b) 1 | 333344 1 | 5677778899 2 | 04
Copyright © 2017 Pearson Education, Inc.
Chapter 2: Exploring Data with Graphs and Numerical Summaries
11
2.18 (continued) c) Histogram of Fertility 9 8
Frequency
7 6 5 4 3 2 1 0
1.1
1.4
1.7 2.0 Fertility
2.3
2.6
2.19 Split Stems a) smallest 0 g, largest 18 g b) 10 g, 11 g, 11 g c) Six cereals have less than 5 grams of sugar with 0 g, 1 g, 3 g, 3 g, 4 g, and 4 g. 2.20 Histogram for sugar a) –1 to 1,1 to 3, 3 to 5, 5 to 7, 7 to 9, 9 to 11, 11 to 13, 13 to 15, 15 to 17, and 17 to 19 b) The distribution is bimodal; child cereals, on average, have more sugar than adult cereals have. c) The dot and stem-and-leaf plots allow us to see all the individual data points. d) The relative differences among bars would remain the same. 2.21 Shape of the histogram a) Assessed value of houses in a large city – skewed to the right (a long right tail) because of some very expensive homes. b) Number of times checking account overdrawn in the past year for the faculty at the local university – skewed to the right because of the few faculty who overdraw frequently. c) IQ for the general population – symmetric because most would be in the middle, with some higher and some lower; there is no reason to expect more to be higher or lower (particularly because IQ is constructed as a comparison to the general population’s “norms”). d) The height of female college students – symmetric because most would fall in the middle, going down to a few short students and up to a few tall students. 2.22 More shapes of histograms a) The scores of students (out of 100 points) on a very easy exam in which most score perfectly or nearly so, but a few score very poorly – skewed to the left because of the few who score poorly. b) The weekly church contribution for all members of a congregation, in which the three wealthiest members contribute generously each week – skewed to the right because of the few wealthy members’ contributions. c) Time needed to complete a difficult exam (maximum time is 1 hour) – skewed to the left because most take almost or all of the whole time, whereas a few finish very quickly. d) Number of music CDs (compact discs) owned, for each student in your school – skewed to the right because of a few students’ huge CD collections.
Copyright © 2017 Pearson Education, Inc.
12 Statistics: The Art and Science of Learning from Data, 4th edition 2.23 Gestational Period a) 9 8
Frequency
7 6 5 4 3 2 1 0
0
100 200 300 400 500 600 700 gestational period
b) The elephant, with a gestational period of 624 days, is unusual. c) The distribution is right-skewed. d) Neither of the two histograms accurately summarizes the distribution. The one with 4 intervals is too coarse, the one with 30 intervals too fine. 14
4
12 3 Frequency
Frequency
10 8 6 4
2
1
2 0
0
160 320 480 gestational period
640
0
0
100 200 300 400 500 600 700 gestational period
2.24 How often do students read the newspaper? a) This is a discrete variable because the value for each person would be a whole number. One could not read a newspaper 5.76 times per week, for example. b) (i) The minimum response is zero. (ii) The maximum response is nine. (iii) Two students did not read the newspaper at all. (iv) The mode is three. c) This distribution is unimodal and somewhat skewed to the right. 2.25 Blossom widths a) The distribution is slightly right-skewed (or roughly symmetric). Most blossoms have a width between 3.2 and 3.6 in. There is one blossom with an unusual small width for that species of less than 2.4 in. b) The distribution is left-skewed. Most blossoms have a width between 2.8 and 3.2 in. 6 15 24 c) 0.90 or 90% 50 d) No. We don’t know how many blossoms in the interval from 2.8 to 3.2 in. are actually wider than 3 in.
Copyright © 2017 Pearson Education, Inc.
Chapter 2: Exploring Data with Graphs and Numerical Summaries
13
2.26 Central Park temperatures a) The distribution is somewhat skewed to the left. b) A time plot connects the data points over time to show time trends. c) A histogram shows the number of observations at each level more easily than does the time plot. We also can see the shape of the distribution from the histogram but not from the time plot. 2.27 Is whooping cough close to being eradicated? a) One can see in the time plot below that after an initial slight increase, there was a sharp and steady decrease in incidence of whooping cough starting around 1940. The decrease leveled off starting around 1960. These data suggest that the whooping cough vaccination was proving effective in reducing the incidence of whooping cough. Scatterplot of Rate per 100,000 vs Year 160
Rate per 100,000
140 120 100 80 60 40 20 0 1920
1930
1940 1950 Year
1960
1970
b) The incidence rate stayed low until about 2000, after which a sharp increase can be observed. No, the United States is not close to eradicating whooping cough. Potential reasons for this include fewer people deciding to get vaccinated and less efficient vaccinations. c) A histogram would not address this question because it does not show the rates for each year; we would not be able to see changes over time. 2.28 Warming in Newnan, GA? Overall, the time plot (below) does seem to show a decrease in temperature over time. Scatterplot of Temperature vs Year 66 65
Temperature
64 63 62 61 60 59 58 1900
1920
1940
1960 Year
Copyright © 2017 Pearson Education, Inc.
1980
2000
14 Statistics: The Art and Science of Learning from Data, 4th edition
Section 2.3 Measuring the Center of Quantitative Data 2.29 Median versus mean a) Median (The distribution would be right-skewed.) b) Median (The distribution would be left-skewed.) c) Mean (The distribution would be symmetric.) 2.30 More median versus mean a) Median (The distribution would be right-skewed.) b) Mean (The distribution would be symmetric.) c) Median (The distribution would be left-skewed.) 2.31 More on CO2 emissions a)
Mean: x
x 8.0 5.3 1.8 1.7 1.2 0.8 0.6 0.5 0.4 0.4 20.7 2.07 n
10 10 n 1 10 1 Median: Find the middle value: 5 12 th position 2 2 0.4, 0.4, 0.5, 0.6, 0.8, 1.2, 1.7, 1.8, 5.3, 8.0 0.8 1.2 1. The median is 2 b) Comparing absolute emission values for nations with different population sizes might be misleading because nations with larger populations tend to have larger total emissions. When viewed per capita, a different picture might emerge. 2.32 Resistance to an outlier a) The median for all three data sets is ten. The values for all three sets of observations are already arranged in numerical order, and the middle number for each is 10. x 8 9 10 11 12 50 10 b) Set 1: x n 5 5 x 8+9+10+11+100 138 27.6 Set 2: x n 5 5 x 8+9+10+11+1000 1038 207.6 Set 3: x n 5 5 c) As the highest value becomes more and more of an extreme outlier, the median is unaffected, whereas the mean increases as the outlier becomes more extreme. 2.33 Income and health insurance The distributions for both will be skewed to the right because the mean is much larger than median. 2.34 Labor dispute Management would want to use the mean because it would be skewed right by the outliers – the few members of management who make a whole lot of money. The mean income would be higher because of the outliers. The workers would prefer the median because it is not affected by the large outliers. It is a more accurate measure of the actual typical income. 2.35 Cereal sodium The moderate skewness to the left causes the mean to be lower than the median. 2.36 Center of plots a) The mean and median would be the same for the dot plots to the middle and to the right because the distributions are symmetric. b) The distribution to the left is skewed to the right, and the mean would be higher than the median would. The mean would be pulled toward the higher, atypical values.
Copyright © 2017 Pearson Education, Inc.
Chapter 2: Exploring Data with Graphs and Numerical Summaries
15
2.37 Public transportation – center a) The mean is 2, the median is 0, and the mode is 0. Thus, the average score is 2, the middle score is 0 (indicating that the mean is skewed by outliers), and the most common score also is 0. x 0 0 4 0 0 0 10 0 6 0 20 2 Mean: x n 10 10 Median: middle score of 0, 0, 0, 0, 0, 0, 0, 4, 6, 10 Mode: the most common score is zero. b) Now the mean is 10, but the median is still 0. x 0 0 4 0 0 0 10 0 6 90 110 10 Mean: x n 10 10 Median: middle score of 0, 0, 0, 0, 0, 0, 0, 4, 6, 10, 90 The median is not affected by the magnitude of the highest score, the outlier. Because there are so many zeros, even though we’ve added one score, the median remains zero. The mean, however, is affected by the magnitude of this new score, an extreme outlier. 2.38 Public transportation – outlier a) The mean versus median applet confirms that the median is not affected by the magnitude of the highest score. Because there are so many zeros, even though we’ve added one score, the median remains zero. The mean, however, is affected by the magnitude of this new score, an extreme outlier. b) The applet demonstrates that the outlier has a weaker effect when there are more scores near the original mean. 2.39 Baseball salaries There are a few valuable players who receive exorbitant salaries, whereas the typical player is paid much less (although still a lot by most people’s standards!). The very high salaries of the few affect the mean, but not the median. 2.40 More baseball salaries Answers will vary. 2.41 European fertility a) The median fertility rate is 1.7. Thus, about half of the countries listed have mean fertility rates at or below 1.7 with the remaining countries having fertility rates above 1.7. b) The mean of the fertility rates is 1.65. c) Since the population of adult women can vary greatly among the countries, it is necessary to calculate an overall fertility rate for the country in order to make comparisons. This rate is found by calculating the mean number of children per adult woman. The mean for a variable need not be one of the possible values for the variable. Although the number of children born to each adult woman is a whole number, the mean number of children born per adult woman need not be a whole number. For example, the mean number of children per adult woman is considerably higher in Mexico than in Canada. 2.42 Sex partners a) Number of partners Number of respondents 0 1 2 3 4 5
102 233 18 9 2 1
Total
365
If the data is sorted from smallest to largest, the median is the number in the (365 + 1)/2 = 183rd position. Since 102 respondents answered 0 and 233 answered 1, the median is 1. Copyright © 2017 Pearson Education, Inc.
16 Statistics: The Art and Science of Learning from Data, 4th edition 2.42 (continued) b) Mean: x
c)
xi 102 0 233 1 18 2 9 3 2 4 15 309 0.85
n 365 365 Since the total number of respondents is 365, the median is still the value in the 183rd place when the data are sorted from smallest to largest. Since 233 respondents gave an answer of 1, the median is still 1. However, the value of the mean changes: Mean: x
xi 0 0 233 1 18 2 9 3 2 4 103 5 819 2.2
n 365 365 2.43 Marriage statistics for 20–24-year-olds a) Women: The mean is 0.274, the median is 0. xi 7350 0 2587 1 80 2 2747 0.274 x n 10, 017 10,017 The median is the middle score. With 10,017 scores, the median is the score in the 5009th position. Thus, the median is 0. Men: The mean is 0.161, the median is 0. xi 8418 0 1594 1 10 2 1614 0.161 x n 10,022 10,022 The median is the middle score. With 10,022 scores, the median is the score between the 5011th and 5012th positions. Thus, the median is 0. b) Using the medians, it seems that there is no difference. Using the mean, in this age group, women have, on average, been married more often. 2.44 Knowing homicide victims a) The mean is 0.16. xi 3944 0 279 1 97 2 40 3 23 4.5 696.5 0.16 x n 4383 4383 b) The median is the middle score. With 4383 scores, the median is the score in the 2192nd position. Thus, the median is 0. c) The median would still be 0, because there are still 2200 people who gave 0 as a response. The mean would now be 1.95. xi 2200 0 279 1 97 2 40 3 1767 4.5 8544.5 1.95 x n 4383 4383 d) The median is the same for both because the median ignores much of the data. The data are discrete; hence, a high proportion of the data falls at only one or two values. The mean is better in this case because it uses the numerical values of all of the observations, not just the ordering. 2.45 Accidents In this case, the mean is likely to be more useful because it uses the numerical values of all of the observations, not just the ordering. Because so many people would report 0 motor accidents, the median is not very useful. It ignores too much of the data.
Section 2.4 Measuring the Variability of Quantitative Data 2.46 Sick leave a) The range is 6; this is the distance from the smallest to the largest observation. In this case, there are six days separating the fewest and most sick days taken (6 – 0 = 6).
Copyright © 2017 Pearson Education, Inc.
Chapter 2: Exploring Data with Graphs and Numerical Summaries
17
2.46 (continued) b) The standard deviation is the typical distance of an observation from the mean (which is 1.25). s2
( x x )2 (0 1.25)2 (0 1.25)2 (4 1.25)2 (6 1.25)2
n 1 39.5 5.643 7
c)
7
s s 2 5.643 2.38 The standard deviation of 2.38 indicates a typical number of sick days taken is 2.38 days from the mean of 1.25. Redo (a) and (b). a) The range is 60; this is the distance from the smallest to the largest observation. In this case, there are sixty days separating the fewest and most sick days taken (60 – 0 = 60). b) The standard deviation is the typical distance of an observation from the mean (which is 8). s2
( x x )2 0 8 0 8 4 8 60 8 2
2
n 1 3104 443.43 7
2
2
7
s s 2 443.43 21.06 The standard deviation of 21.06 indicates a typical number of sick days taken is 21.06 days from the mean of 8. The range and mean both increase when an outlier is added. 2.47 Life expectancy a) Upon examination of the data, the countries in Africa will have a larger standard deviation since the spread of the data is greater for this group than for the countries in Western Europe. b) Western Europe: x 81 80 80 81 80 82 82 83 1220 81.3333 x n 15 15
s Africa:
2 2 2 x x 81 81.3333 83 81.3333 1.05
n 1
15 1
x
x 47 50 51 57 64 63 62 61 914 57.125
s
2 x x 47 57.125 61 57.125 5.18
n
16
2
16
2
n 1 16 1 Note that the standard deviation for the Western Europe group, 1.0 (rounded), is much smaller than for the Africa group, 5.2. 2.48 Life expectancy including Russia We would expect the standard deviation to be larger since the value for Russia is significantly smaller than the rest of the group adding additional spread to the data. The standard deviation including Russia is, in fact, 3.01. 2.49 Shape of home prices? The most plausible value is $60,000. –$15,000 is not possible because a standard deviation cannot be negative. $1,000 and $1,000,000 are unlikely because they are too small or too big, respectively, for a typical deviation. One would not expect the typical deviation to be that far from the median for home prices.
Copyright © 2017 Pearson Education, Inc.
18 Statistics: The Art and Science of Learning from Data, 4th edition 2.50 Exam standard deviation The most realistic value is 12. There are problems with all the others. –10: We can’t have a negative standard deviation. 0: We know there is spread because the scores ranged from 35 to 98, so the standard deviation is not 0. 3: This standard deviation seems very small for this range. 63: This standard deviation is too large for a typical deviation. In fact, no score differed from the mean by this much. 2.51 Heights a) According to the Empirical Rule, 68% of men would be within one standard deviation of the mean between 71 – 1(3) = 68 and 71 + 1(3) = 74 inches. 95% of men would be within two standard deviations of the mean, between 71 – 2(3) = 65 and 71 + 2(3) = 77 inches. All or nearly all men would be within three standard deviations of the mean, between 71 – 3(3) = 62 and 71 + 3(3) = 80 inches. b) The mean for women is lower than the mean for men. Because each gender’s heights would tend to be closer to that gender’s mean than to the overall mean, the standard deviation would be smaller when we compared them with the appropriate gender group than when we compared them to the overall group. Would not expect unimodal but more bimodal. 2.52 Histograms and standard deviation a) (i) The sample on the right has the largest standard deviation since is the most spread out. (ii) The sample in the middle has the smallest standard deviation since it has no spread. b) The Empirical Rule is relevant only for the distribution on the left because the distribution is bellshaped. 2.53 Female strength According to the Empirical Rule, 68% of women would be able to lift within one standard deviation from the mean, between 79.9 – 1(13.3) = 66.6 and 79.9 + 1(13.3) = 93.2 pounds. 95% of women would be able to lift within two standard deviations from the mean, between 79.9 – 2(13.3) = 53.3 and 79.9 + 2(13.3) = 106.5 pounds. All or nearly all women would be able to lift within three standard deviations from the mean, between 79.9 – 3(13.3) = 40.0 and 79.9 + 3(13.3) = 119.8 pounds. 2.54 Female body weight a) 95% of weights would fall within two standard deviations from the mean, between 133 – 2(17) = 99 and 133 – 2(17) = 167. b) An athlete who is three standard deviations above the mean would weight 133 + 3(17) = 184 pounds. This would be an unusual observation because typically all or nearly all observations fall within three standard deviations from the mean. In a bell-shaped distribution, this would likely be about the highest score one would obtain. 2.55 Shape of cigarettes taxes With a bell-shaped distribution, we expect scores to extend about three standard deviations from the mean 0 73 1.52, or 1.52 standard in either direction. The lowest possible value of 0, however, is only 48 deviations below the mean, and so the distribution likely is skewed to the right. 2.56 Empirical rule and skewed, highly discrete distribution a)
x
x 8418 0 1594 1 10 2 1614 0.16
s
x x
10022
n
2
n 1
10022
8418 0 0.16 1594 1 0.16 10 2 0.16 0.37 10022 1 2
2
Copyright © 2017 Pearson Education, Inc.
2
Chapter 2: Exploring Data with Graphs and Numerical Summaries
19
2.56 (continued) b) Observations
Predicted by Empirical Rule
One standard deviation from the mean is between 0.16 84.0% 68% – 1(0.37) = –0.21 and 0.16 + 1(0.37) = 0.53 Two standard deviations from the mean is between 84.0% 95% 0.16 – 2(0.37) = –0.58 and 0.16 + 2(0.37) = 0.90 Three standard deviations from the mean is between 99.9% About 100% 0.16 – 3(0.37) = –0.95 and 0.16 + 3(0.37) = 1.27 There are more observations within one standard deviation of the mean and fewer within two standard deviations than would be predicted by the Empirical Rule. c) The Empirical Rule is only valid when used with data from a bell-shaped distribution. This is not a bell-shaped distribution; rather, it is highly skewed to the right. Most observations have a value of 0, and hardly any have the highest value of 2. 2.57 How much TV? These statistics suggest that this distribution is highly skewed toward the right for two main reasons. The mean is larger than the median, and the standard deviation is almost as large as the mean. In fact, the 0 3.09 1.08, or 1.08 standard deviations below the mean. lowest possible value of 0 is only 2.87 2.58 How many friends? a) The standard deviation is larger than the mean; in addition, the mean is higher than the median. In 0 7.4 0.67, or 0.67 standard deviations below the fact, the lowest possible value of 0 is only 11.0 mean. These situations occur when the mean and standard deviation are affected by an outlier or outliers. It appears that this distribution is skewed to the right. b) The Empirical Rule does not apply to these data because they do not appear to be bell-shaped. 2.59 Judging skew using x and s The largest observation is 86 is less than one standard deviation above the mean of 70.4. Specifically, 86 is 86 70.4 0.93 standard deviations above the mean. The smallest observation is 35, which is only 16.7 35 70.4 2.12 , or 2.12 standard deviations below the mean. This distribution, therefore, is likely 16.7 skewed to the left. Dotplot of Exam Scores
40
48
56 64 Exam Scores
72
Copyright © 2017 Pearson Education, Inc.
80
88
20 Statistics: The Art and Science of Learning from Data, 4th edition 2.60 Youth unemployment in the EU a) 30
unemployment
25 20 15 10 5 0
l ria n y rg lta d s l ic r k ia m e n nd m i a ia ry c e nd i a ia l y ri a nd i a u s a ti a in ce s t a ou Ma la n u b m a a n gd o e d n la l gi u to n v en ga ra n o la ua n atv I ta lg a e la v a k pr rtug ro a S pa ree r e p n o m in w F i e s l o un F P th L r lo y o C u I A u erm mb e G S B B E S H h R De R K S C P Li G xe et d N ec h Lu te z ni C U
country
b) Mean = 11.1%, median = 10.1%, s = 5.6% c) The distribution of the youth unemployment rate in the EU is skewed to the right, with two countries (Greece and Spain) showing an unemployment rate of more than 20%. The mean unemployment rate is 11.1% (median 10.1%). The variability in the unemployment rate is relatively large, with a standard deviation of 5.6%, but this may be inflated due to the two outliers and right-skewness of the distribution. 2.61 Create data with a given standard deviation a) One possible answer: 30, 50, 80 b) One possible answer: 10, 50, 90. c) The largest standard deviations results from two 0s and two 100s, with s = 57.74.
Section 2.5 Using Measures of Position to Describe Variability 2.62 Vacation days a) Median: Find the middle value of 13, 25, 26, 28, 34, 35, 37, 42. The median is 31, the average of the two middle values, 28 and 34. b) The first quartile is the median of 13, 25, 26, 28. The first quartile is 25.5, the average of the two middle values, 25 and 26. c) The third quartile is the median of 34, 35, 37, 42. The third quartile is 36, the average of the two middle values. d) 25% of countries have residents who take fewer than 25.5 vacation days, half of countries have residents who take fewer than 31 vacation days, and 75% of countries have residents who take fewer than 36 vacation days per year. The middle 50% of countries have residents who take an average of between 25.5 and 36 vacation days annually. 2.63 Youth unemployment a) The median is 10.15, the average of the 14th and 15th values, 10.0 and 10.3. In 2013, half of the European Union nations had an unemployment rate less than 10.15%. b) The first quartile is 7.15, the average of the 7th and 8th values, 7.0 and 7.3. In 2013, 75% of the European Union nations had an unemployment rate larger than 7.15% (or 25% had rate less than 7.15%). c) The third quartile is 13.05, the average of the 21st and 22nd values, 13.0 and 13.1. In 2013, the unemployment rate was larger than 13.05% for 25% of the European Union nations. d) The 10th percentile will be around 6% because Q1 = 7.15%.
Copyright © 2017 Pearson Education, Inc.
Chapter 2: Exploring Data with Graphs and Numerical Summaries
21
2.64 Female strength a) One fourth of the females had a maximum bench press less than 70 pounds, and one fourth had a maximum bench press greater than 90 pounds. b) The mean and median are the about the same, and the first and third quartiles are equidistant from the median. These are both indicators of a roughly symmetric distribution. 2.65 Female body weight a) One quarter of the females had weight below 119 and one quarter had weight above 144. b) The mean and median are the about the same, and the first and third quartiles are approximately equidistant from the median. These are both indicators of a roughly symmetric distribution. 2.66 Ways to measure variability a) The range is even more affected by an outlier than is the standard deviation. The standard deviation takes into account the values of all observations and not just the most extreme two. b) With a very extreme outlier, the standard deviation will be affected both because the mean will be affected and because the deviation of the outlier (and its square) will be very large. The IQR would not be affected by such an outlier. c) The standard deviation takes into account the values of all observations and not just the two marking 25% and 75% of observations. 2.67 Variability of cigarette taxes a) (i) Q1, marking the lowest 25% of states, has a value of 36. Thus, 75% of states have cigarette taxes greater then 36 cents. (ii) Q3, marking the highest 25% of states, has a value of 100. Thus, 25% of states have cigarette taxes greater than $1.00. b) The two values that demarcate the middle 50% are Q1 = 36 cents and Q3 = 100 cents (one dollar). c) The interquartile range (IQR) is the difference between Q1 and Q3. IQR = Q3 – Q1 = 100 – 36 = 64 cents. For the middle 50% of state cigarette taxes, $0.64 is the distance between the largest and smallest cigarette tax amount. d) With a bell-shaped distribution, we expect Q1 and Q3 to be roughly equidistant from the median which is not the case here. The maximum value is also quite far from Q3. Thus, it appears that the distribution is skewed to the right. 2.68 Sick leave a) The range is six; this is the distance from the smallest to the largest observation. In this case, there are six days separating the fewest and most sick days taken (6 – 0 = 6). b) The interquartile range is the difference between Q3 and Q1. IQR = Q3 – Q1 = 2. c) Redo (a) and (b). a) The range is sixty; this is the distance from the smallest to the largest observation. In this case, there are sixty days separating the fewest and most sick days taken (60 – 0 = 60). b) Q1, the median of all scores below the median, is still 0. Q3, the median of all scores above the median, is still 2 (the average of 0 and 4). The interquartile range remains the same: IQR = Q3 – Q1 = 2. The IQR is least affected by the outlier because it doesn’t take the magnitudes of the two extreme scores into account at all, whereas the range and s do. 2.69 Infant mortality Africa a) Q1 is the median of the lower half of the sorted data: 54, 63, 68, 76, 78, 79, 80. It is 76. Q3 is the median of the upper half of the sorted data: 81, 84, 96, 101, 110, 121, 154. It is 101. b) IQR = Q3 – Q1 = 25. For the middle half of the infant mortality rates, the distance between the largest and smallest rates is 25. 2.70 Infant mortality Europe Q1 is the median of the lower half of the sorted data: 3, 3, 3, 4, 4, 4, 4. It is 4. Q2, the median, is the 44 4 . Q3 is the median of the upper half of the sorted data: 4, average of the middle two data values 2 4, 4, 4, 5, 5, 5. It is 4.
Copyright © 2017 Pearson Education, Inc.
22 Statistics: The Art and Science of Learning from Data, 4th edition 2.71 Computer use a) This five-number summary suggests that the distribution is skewed to the right. The distance between the minimum and the median is much smaller than the distance between the median and the maximum. b) In this case, outliers would be those values more than 1273.5 points from the first and third quartiles: IQR = Q3 – Q1 = 1105 – 256 = 849 and 1.5(IQR) = 1.5(849) = 1273.5 The lower boundary: Q1 – 1.5(IQR) = 256 – 1273.5 = –1017.5 The upper boundary: Q3 + 1.5(IQR) = 1105 + 1273.5 = 2378.5 In the current example, the lowest score is 4, so there are no scores below –1017.5. The highest score, on the other hand, is 320,000, much higher than 2378.5. Thus, there are potential outliers according to this criterion. 2.72 Central Park temperature distribution revisited a) We would expect it to be skewed to the left because the maximum is closer to the median than is the minimum. b) Numbers are approximate: Minimum: 49.0, Q1: 52.5, Median: 53. 5, Q3: 55.0, Maximum: 57.0 These approximations support the premise that the distribution is skewed to the left if it is skewed. The median is closer to the maximum and Q3 than it is to the minimum and Q1. 2.73 Box plot for exam The minimum, Q1, median, Q3, and maximum are used in the box plot. Boxplot of Exam Score 100
Score
90
80
70
60
2.74 Public transportation a) Minimum: 0, Q1: 0, Median: 0, Q3: 4, Maximum: 10 Boxplot of Miles per day 10
Miles per day
8
6
4
2
0
Copyright © 2017 Pearson Education, Inc.
Chapter 2: Exploring Data with Graphs and Numerical Summaries
23
2.74 (continued) b) Q1 and the median share the same line in the boxplot because so many employees have a score of zero that the middle score of the whole set of data is zero and the middle score of the lower half of the data also is zero. c) There is no whisker because the minimum score also is zero. This situation resulted because there are so many people with the lowest score. 2.75 Energy statistics a) Numbers are approximate: Minimum: 50, Q1: 130, Median: 160, Q3: 250, Maximum: 650 One country was a potential outlier, the one around 650. b) We can know how far Italy was from the mean in terms of standard deviations by calculating its zscore. It is 0.47 standard deviations below the mean of 195. x x 139 195 z 0.47 120 s c) The U.S. is 1.16 standard deviations above the mean. x x 334 195 z 1.16 120 s 2.76 European Union youth unemployment rates a) In a box plot, Q1 = 7.15 and Q3 = 13.05, would be the outer edges of the box. 1.5(IQR) = 1.5(13.05 – 7.15) = 8.85. The whisker on the left would extend to the minimum 4.9, since it is larger than 7.15 – 8.85 = –1.7. The whisker on the right would extend to the value of 17.2 (Croatia), since it is the largest value below 7.15 + 8.85 = 21.9. b) Greece (27.3) and Spain (26.4) have values larger than 21.9, so would be considered outliers. c) Greece’s score is 2.89 standard deviations above the mean, and thus, is not an outlier according to the three standard deviation criterion. x x 27.3 11.1 z 2.89 s 5.6 d) A z-score of 0 indicates that the country’s unemployment rate is zero standard deviations from the mean; hence, the unemployment rate is equal to the mean. In this case, a country with an unemployment rate of 11.1 would have a z-score of 0. 2.77 Air pollution a) Finland’s pollution is exactly one standard deviation above the mean pollution of all countries in the EU. x x 11.5 7.9 z 1 s 3.6 b) Sweden’s pollution is 0.64 standard deviation below the mean pollution of all countries in the EU. x x 5.6 7.9 z 0.64 s 3.6 c) The United Kingdom’s pollution is exactly equal to the mean pollution of all countries in the EU. x x 7.9 7.9 z 0 s 3.6 2.78 Female heights x x 56 65.3 3.1 s 3.0 b) The negative sign indicates that the height of 56 inches is below the mean. c) Because the height of 56 inches is more than three standard deviations from the mean, it is a potential outlier.
a)
z
Copyright © 2017 Pearson Education, Inc.
24 Statistics: The Art and Science of Learning from Data, 4th edition 2.79 Hamburger sales This z-score indicates that the sales for this day are more than three standard deviations above the mean, and thus, would be a potential outlier – in other words, an unusually good day. x x 2000 1165 3.80 220 s 2.80 Florida students again a) The distribution depicted in the box plot is skewed to the right. Most observations fall between about 0 and 15 but there are a few outliers representing very large values. Minimum: 0, Q1: 3, Median: 6, Q3: 10, Maximum: 37 z
Boxplot of TV 40
TV
30
20
10
0
b) Since IQR = 10 – 3 = 7 and 1.5(IQR) = 1.5(7) = 10.5, the 1.5(IQR) criterion would indicate that all data should fall between about 3 – 10.5 = –7.5 and 10 + 10.5 = 20.5. Because some data points fall beyond this range, it appears that there are potential outliers. 2.81 Females or males watch more TV? Based on the Florida survey data, females tend to watch more TV. The median, Q1 and Q3 are higher for females than for males. Boxplot of TV 40
TV
30
20
10
0 f
m gender
2.82 CO2 comparison a) The two outliers for Central and South America have a value of roughly 12 metric tons. b) The distributions would be skewed to the right. The median sits low in the box (pulled toward the first quartile), and the upper whisker stretches out more from the box to the maximum value than the left whisker stretches to minimum value.
Copyright © 2017 Pearson Education, Inc.
Chapter 2: Exploring Data with Graphs and Numerical Summaries
25
2.82 (continued) c) The median CO2 emission is much larger in Europe than the median for Central and South America. The spread of the middle 50% of the distribution of emissions, as measured by the IQR, seems to be about the same for Europe and Central and South America. 75% of the distribution of emissions for Europe is higher than the lower 75% of the distribution of emissions for Central and South America. Overall, emissions are higher for Europe than for Central and South America.
Section 2.6 Recognizing and Avoiding Misuses of Graphical Summaries 2.83 Great pay (on the average) a) The mean is $43,700 and the median is $9300. b) It is misleading because the mean is so heavily influenced by the outlier (her own salary) that it is not a typical value. The median would be a much more accurate summary of these salaries. 2.84 Market share for food sales a) One problem with this chart is that the percentages do not add up to 100. Second, the Tesco slice seems too large for 27.2%. A third problem is that contiguous colors are very similar. This increases the difficulty in easily reading this chart. b) It would be easier to identify the mode with a bar graph because one would merely have to identify the highest bar. 2.85 Enrollment trends a) This graph shows an overall decrease in enrollment in STEM majors at first, with what appears to be somewhat of a plateau toward the end of the time span. Time Series Plot of STEM Majors 2000
STEM Majors
1950
1900
1850
1800 2004
2005
2006
2007
2008 Year
2009
2010
2011
2012
b) This graph shows a gradual decrease over time in the percentage of students who are enrolled in STEM majors. Time Series Plot of Percent of enrolled students in STEM
Percent of enrolled students
0.0675 0.0650 0.0625 0.0600 0.0575 0.0550 2004
2005
2006
2007
2008 Year
2009
Copyright © 2017 Pearson Education, Inc.
2010
2011
2012
26 Statistics: The Art and Science of Learning from Data, 4th edition 2.85 (continued) c) The graphs in (a) and (b) tell us that although there are some fluctuations in the numbers of students enrolling in STEM majors over the years, there is a steady decrease in the percentage of enrolling students who are enrolled in STEM majors over the years. We cannot learn this from Figures 2.18 and 2.19. 2.86 Terrorism and war in Iraq a) This graph is misleading. Because the vertical axis does not start at 0, it appears that six times as many people are in the “no, not” column than in the “yes, safer” column, when really it’s not even twice as many. b) With a pie chart, the area of each slice represents the percentage who fall in that category. Therefore the relative sizes of the slices will always represent the relative percentages in each category. 2.87 BBC license fee The 2013 projection is shown where the observation would be plotted for the year 2007, not 2013. 2.88 Federal government spending The slices do not seem to have the correct sizes, for instance the slice with 16% seems larger than the slice with 19%. 2.89 Bad graph Answers will vary.
Chapter Problems: Practicing the Basics 2.90 Categorical or quantitative? a) Number of children in family: quantitative b) Amount of time in football game before first points scored: quantitative c) Choice of major (English, history, chemistry, …): categorical d) Preference for type of music (rock, jazz, classical, folk, other): categorical 2.91 Continuous or discrete? a) Age of mother: continuous b) Number of children in a family: discrete c) Cooking time for preparing dinner: continuous d) Latitude and longitude of a city: continuous e) Population size of a city: discrete 2.92 Young non-citizens in the U.S. a) Region of Birth is Categorical. Noncitizens aged 18 to 24 in the United States Region of Birth Number (in Thousands) Percentage
Africa
115
Asia
590
Europe
148
Latin America & Caribbean
1666
Other
49
Total
2568
Copyright © 2017 Pearson Education, Inc.
115 0.045 4.5% 2568 590 0.230 23.0% 2568 148 0.058 5.8% 2568 1116 0.649 64.9% 2568 49 0.019 1.9% 2568
Chapter 2: Exploring Data with Graphs and Numerical Summaries
27
2.92 (continued) b) Mode. Most young noncitizens are from Latin America and the Caribbean. c) Latin America and Caribbean, Asia, Europe, Africa, Other. One immediately sees that most young noncitizens are from two regions, Latin America and Caribbean and from Asia. 2.93 Cool in China a) The variable being measured is the personality trait that defines “cool.” b) This is a categorical variable. c) Because the data are categorical with unordered categories, we could use only the bar chart and the modal category. 2.94 Chad voting problems a) We first locate the dot directly above 11.6% on the horizontal x axis. We then look at the vertical y axis across from this point to determine the label for that dot: Optical scanning with a two-column ballot. This tells us that the over-vote was highest among those using optical scanning with a twocolumn ballot. b) We first locate the dots above the lowest percentages on the x axis. We then determine the labels across from these dots on the y axis to determine the lowest two combinations: optical, one column, and votomatic, one column. Thus, the lowest over-voting occurred when voters had a ballot with only one column that was registered either using optical scanning or votomatic (manual punching of chads). c) We could summarize these data further by using a bar for each combination: optical, one column; optical, two column; votomatic, one column, etc. For each bar, we could then plot the average overvote of all counties in that category. To do this, we would need the exact percentages of each county in each category. 2.95 Number of children a) A histogram would be most appropriate since the upper interval is 8+, which may contain digits above 8. b) The distribution is skewed to right. 600
Frequency
500 400 300 200 100 0
0
1
2 3 4 5 6 7 Number of Children
2.96 Longevity a) 0 | 0 | 57789 1 | 0011112234 1|9 2 | 23 2| 3|0 3|5
Copyright © 2017 Pearson Education, Inc.
8
28 Statistics: The Art and Science of Learning from Data, 4th edition 2.96 (continued) b) 9 8
Frequency
7 6 5 4 3 2 1 0
4
8
12
16 20 24 longevity
28
32
36
c) The distribution of longevity is right-skewed. Most animals live to be between 5 and 15 years old. 2.97 Newspaper reading a) Dotplot of Newspaper Reading
0
2
4 6 Times per Week
b)
8
0 | 00 1 | 0000 2 | 0000 3 | 00000000 4 | 0000 5 | 00000 6 | 00 7 | 0000 8 | 00 9|0 The leaf unit is identified above. The stems are the whole numbers, 0 through 9. c) The median is the middle number. There are 36 numbers, so the median is between the 18th and 19th which have the values 3 and 4, respectively. Thus, the median is 3.5. d) The distribution is slightly skewed to the right.
Copyright © 2017 Pearson Education, Inc.
Chapter 2: Exploring Data with Graphs and Numerical Summaries
29
2.98 Match the histogram a) symmetric and bimodal c) skewed to the left b) skewed to the right d) symmetric and unimodal 2.99 Sandwiches and protein a) 0 | 8 1| 1 | 7889 2 | 113 2 | 666 b) A stem-and-leaf plot allows one to see the individual amounts. c) The protein amounts are mostly between 17 and 21 grams with a few sandwiches having a higher protein value of 26 grams. There appears to be one outlier having only 8 grams of protein. 2.100 Sandwiches and cost a) (The data values have been truncated.) 2|4 2 | 999 3 | 1444 3 | 688 b) A stem-and-leaf plot allows one to see the individual prices. c) Most of the sandwiches cost between $2.90 and $3.89. The prices are skewed to the left with one sandwich costing only $2.49. 2.101 What shape do you expect? a) Number of times arrested in past year – skewed to the right because most values are at 0 but there are some large values. b) Time needed to complete difficult exam (maximum time is 1 hour) – skewed to the left because most values are at 1 hour or slightly less, but some could be quite a bit less. c) Assessed value of home – skewed to the right because there are some extremely large values. d) Age at death – skewed to the left because most values are high, but some very young people die. 2.102 Sketch plots NOTE: Plots will vary, but should have the following characteristics. a) It would be skewed to the right, and the mean would be greater than the median because of a few mansions that sell for millions. b) It would be skewed to the right, and the mean would be higher than the median. Most women do not give birth over age 40. Thus, the median would be zero. The mean, however, would be positive, because some women do give birth over the age of 40. c) It would be skewed to the left, and the mean would be lower than the median. The mean would be pulled down by the outlier of 50. The standard deviation is only 10, so there probably aren’t lots of low scores. Moreover, the highest possible score of 100 is only 12/10 = 1.2 standard deviations above the mean. d) It would be skewed to the left, and the mean would likely be lower than the median. Most people with cars drive them every month, but a few drive them less, and some hardly or not at all. These outliers would pull the mean, but not the median, lower. The median and mode probably would be 12. 2.103 Median versus mean sales price of new homes We would expect the mean sales price to have been higher due to the distribution being skewed to the right. A few very expensive homes will greatly affect the mean, but not the median sales price. 2.104 Household net worth a) The distribution of these families’ net worth is likely to be skewed to the right because relatively few families would have very high net worth so that we expect the mean to be greater than the median. b) When assets such as homes and retirement savings decline due to a recession, it is typical for the highest valued assets to be affected the most. Thus, we would expect the mean net worth to drop more than the median net worth. Copyright © 2017 Pearson Education, Inc.
30 Statistics: The Art and Science of Learning from Data, 4th edition 2.105 Golfers’ gains a) The data for the 90 players would be skewed to the right with the majority of the golfers earning between $1 and $3 million and a few earning over $3 million. b) Since the data is skewed to the right, the mean would be the higher value of $2,090,012 and the median the lesser value of $1,646,853. 2.106 Hiking The classification into easy, medium or hard is categorical and the length classification is quantitative. 2.107 Lengths of hikes a) One example is 1, 2, 4, 6, 7. Both the mean and median are 4. b) One example is 2, 2, 3, 5 and 6. 2.108 Central Park monthly temperatures a) Both distributions are fairly symmetric and bell-shaped, with January having greater variability than July. b) The mean temperature for January is around 32º and the mean temperature for July is around 76º. The average monthly temperature in January is approximately 44º less than the average monthly temperature in July. c) The average monthly temperature in January is more variable than in July. The range of average temperatures for January is approximately 22º to 43º and the standard deviation is approximately 5º. The range of average temperatures for July is approximately 71º to 81º and the standard deviation is approximately 2º. It may be a bit surprising to see how much more variable are the average monthly temperatures in January than in July. 2.109 What does s equal? a) Given the mean and range, the most realistic value is 12. –10 is not realistic because standard deviation must be 0 or positive. Given that there is a large range, it is not realistic that there would be almost no spread; hence, the standard deviation of 1 is unrealistic. 60 is unrealistically large; the whole range is hardly any more than 60. b) –20 is impossible because standard deviations must be nonnegative. 2.110 Female heights a) According to the Empirical Rule, 95% of scores in a bell-shaped distribution fall within two standard deviations of the mean. x 2 s 65 2 3.5 58 x 2 s 65 2 3.5 72
Thus, 95% of heights likely fall between 58 and 72 inches. b) The height for a woman who is three standard deviations below the mean is 54.5. x 3s 65 3 3.5 54.5 This is on the cusp of what would be considered an outlier according to the z-score criterion. Scores that are beyond three standard deviations from the mean are considered to be potential outliers. So, yes, this height is bordering on unusual. 2.111 Energy and water consumption a) The distribution is likely skewed to the right because the maximum is much farther from the mean than the minimum is, and also because the standard lowest possible value of 0 is only 780/506 = 1.54 standard deviations below the mean. b) The distribution is likely skewed to the right because the standard deviation is almost as large as the mean, and the smallest possible value is zero, only 1.15 standard deviation below the mean. 2.112 Hurricane damage a) The distribution is skewed to the right. b) The median should be used since the distribution is skewed to the right. c) The values are correct.
Copyright © 2017 Pearson Education, Inc.
Chapter 2: Exploring Data with Graphs and Numerical Summaries
31
2.112 (continued) d) The distribution of hurricane damage is skewed to the right, with the damage for the costliest hurricane Katrina (more than 100 billion) far exceeding all others. The median damage was 7.9 billion. The 25% most costly hurricanes had a cost of over 11.8 billion, whereas the 25% fewest damaging hurricanes cost no more than 5.7 billion. 2.113 More hurricane damage a) The percentage differs from 68% because of the extreme right skew of the distribution. b) The mean and standard deviation would get much smaller due to the removal of the extreme value. The median, IQR, and 10th percentile would not change much. 2.114 Student heights a) For a bell-shaped distribution, such as the heights of all men, the Empirical Rule states that all or nearly all scores will fall within three standard deviations of the mean ( x 3s ). In this case, that means that nearly all scores would fall between 62.2 and 79.6. In this example, almost all men’s scores do fall between these values. b) The center for women is about five inches less than the center for men. The variability, however, is very similar. These distributions are likely very similar in shape; they are just centered around different values. x x 62 70.9 c) The lowest score for men is 62. This would have a z-score of z 3.07 . Thus, it s 2.9 falls 3.07 standard deviations below the mean. 2.115 Cigarette tax a) 12
Frequency
10 8 6 4 2 0
0
40
80
120 Cigarette Tax
160
200
The histogram shows a unimodal distribution that is skewed to the right. If there are any outliers, they would be the most extreme scores, such as the one around 200. b) The mean is 72.85 and the median is 60. The mean is inflated relative to the median as one would expect from the distribution depicted in the histogram that is skewed to the right. The few high scores would pull the mean higher, but not the median. c) The standard deviation is 48.00. This indicates that the typical score falls about 48.0 from the mean. 2.116 Cereal sugar values a) Numbers are approximate: Minimum: 0, Q1: 4, Median: 9.5, Q3: 13.5, Maximum: 18 b) Because the median is closer to Q3 and the maximum than it is to Q1 or the minimum, it appears that this distribution is slightly skewed to the left. c) This sugar value falls 1.64 standard deviations below the mean. x x 0 8.75 z , z 1.64 s 5.32
Copyright © 2017 Pearson Education, Inc.