Skip to main content

SOLUTIONS MANUAL FOR Analyzing Data and Making Decisions Statistics for Business Microsoft Excel 201

Page 1

Solutions Manual

For

Analyzing Data and Making Decisions

Statistics for Business Microsoft Excel 2010 Updated Second Edition

Judith Skuce

Part 1: Page 1-454 Part 2: Page 455-732


Contents Part I: Introduction Chapter 1: Using Data to Make Better Decisions................................................................1

Part II: Descriptive Statistics Chapter 2: Using Graphs and Tables to Describe Data .......................................................9 Chapter 3: Using Numbers to Describe Data.....................................................................51

Part III: Building Blocks for Inferential Statistics Chapter 4: Calculating Probabilities ..................................................................................65 Chapter 5: Probability Distributions ..................................................................................88 Chapter 6: Using Sampling Distributions to Make Decisions .........................................108

Part IV: Making Decisions Chapter 7: Making Decisions with a Single Sample .......................................................129 Chapter 8: Estimating Population Values ........................................................................161 Chapter 9: Making Decisions with Matched Pairs Samples, Quantitative or Ranked Data..........................................................................194 Chapter 10: Making Decisions with Two Independent Samples, Quantitative or Ranked Data..........................................................................234 Chapter 11: Making Decisions with Three or More Samples, Quanitative Data—Analysis of Variance (ANOVA) ....................................275 Chapter 12: Making Decisions with Two or More Samples, Qualitative Data .............................................................................................321

Part V: Analyzing Relationships Chapter 13: Analyzing Linear Relationships, Two Quantitative Variables ....................351 Chapter 14: Analyzing Linear Relationships, Two or More Variables ...........................389


Instructor’s Solutions Manual - Chapter 1

Chapter 1 Solutions Develop Your Skills 1.1 1. You would have to collect these data directly from the students, by asking them. This would be difficult and time-consuming, unless you are attending a very small school. You might be able to get a list of all the students attending the school, but privacy protection laws would make this difficult. No matter how much you tried, you would probably find it impossible to locate and interview every single student (some would be absent because of illness or work commitments or because they do not attend class regularly). Some people may refuse to answer your questions. Some people may lie about their music preferences. It would be difficult to solve some of these problems. You might ask for the school's cooperation in contacting students, but it is unlikely they would comply. You could offer some kind of reward for students who participate, but this could be expensive. You could enter participants' names in a contest, with a music-related reward available. None of these approaches could guarantee that you could collect all the data, or that students would accurately report their preferences. One partial solution would be to collect data from a random sample of students, as you will see in the discussion in Section 1.2 of the text. Without a list of all students, it would be difficult to ensure that you had a truly random sample, but this approach is probably more workable than a census (that is, interviewing every student). 2.

Because you need specific data on quality of bicycle components, you would need to collect primary data. Customer complaints about quality are probably the only source of secondary data that you would have.

3.

Statistics Canada has a CANSIM Table 203-0010, Survey of household spending (SHS), household spending on recreation, by province and territory, annual, which contains information on purchases of bicycles, parts and accessories. There is a U.S. trade publication called "Bicycle Retailer & Industry News", which provides information about the industry. See http://www.bicycleretailer.com/. Access is provided through the Business Source Complete database. Industry Canada provides a STAT-USA report on the bicycle industry in Canada, at http://strategis.ic.gc.ca/epic/internet/inimr-ri.nsf/en/gr105431e.html. Somewhat outdated information is also available at http://www.ic.gc.ca/eic/site/sgas.nsf/eng/sg03430.html. Canadian Business magazine has a number of articles on the bicycle industry. One of the most recent describes the purchase of the Iron Horse Co. of New York by Dorel Industries (a Montreal firm). http://www.canadianbusiness.com/markets/headline_news/article.jsp?content=b1560 9913

Copyright © 2011 Pearson Canada Inc.

1


Instructor’s Solutions Manual - Chapter 1

4.

Although Statistics Canada takes great care in its data collection, errors do still occur, and data revisions are required. An interesting overview of GDP data quality for seven OECD countries is available at http://www.oecd.org/dataoecd/20/26/34350524.pdf You should be able to locate other information about data revisions. See also http://www.statcan.ca/english/about/policy/infousers.htm which describes Statistics Canada’s policy on informing users about data quality.

5.

At least some of the secondary data sources listed in Section 1.1 should help you. If you cannot locate any secondary data, get help from a librarian.

Develop Your Skills 1.2 6. The goal for companies is to create population data, but it is unlikely that every customer is captured in any CRM database. There are many examples of companies using CRM data. A search of the CBCA database on August 7, 2009 produced a list of 102 articles (for 2009) that contained “customer relationship management” as part of their citation and indexing. For example, the publication called "Direct Marketing" regularly writes about database marketing, data mining, and web analytics. See http://www.dmn.ca/index.html. 7.

This is a nonstatistical sample, and could be described as a convenience sample. The restaurant presumably has diners on nights other than Friday, and none of these could be selected for the sample. The owner should not rely on the sample data to describe all of the restaurant's diners, although the sample might be useful to test reaction to a new menu item, for example.

8.

These are sample statistics, as they are based on sample data. It would be impossible to collect data from all postsecondary students.

9.

Follow the instructions for Example 1.2c. The random sample you get will be different, but here is one example of the 10 names selected randomly. AVERY MOORE EMILY MCCONNELL HARRIET COOGAN DYLAN MILES TERRY DUNCAN GEORGE BARTON JAMES BARCLAY AVA WORTH PAIGE EATON JORDAN BOCK

10. First, Calgary Transit will probably find it impossible to establish a frame for its target population, which is people with disabilities who use Calgary Transit. It will also have to carefully define what it means by “people with disabilities”. If this

Copyright © 2011 Pearson Canada Inc.

2


Instructor’s Solutions Manual - Chapter 1

means “people in wheelchairs”, then it will at least be possible to identify such riders when an interviewer visits a bus or a bus stop. However, it will be quite difficult for Calgary Transit to obtain a truly random sample of the opinions of people in wheelchairs who use Calgary Transit’s services. Coverage errors will be practically unavoidable. As well, if interviewers are approaching only those riders in wheelchairs, the survey respondents may be unhappy about being singled out because of their wheelchairs. They may refuse to answer the interviewer’s questions, leading to nonresponse errors. Interviewers will have to be trained carefully to overcome any resulting resistance of survey subjects. Because data will probably be collected on buses or at bus stops, with interviewers recording information while a bus is in motion, or possibly during bad weather at a bus stop, processing errors may occur. Finally, Calgary Transit will have to be sure that suitably qualified people are doing the analysis, to avoid estimation errors. Develop Your Skills 1.3 11. This is impossible. A price cannot decrease by more than 100% (and a 100% decrease would mean the price was 0). It is likely that the company means that the old price is 125% of the current price. So, for example, if the old price was $250, then the new price would be $200. You can see that 250/200 = 1.25 or 125%. 12. The graph with the y-axis that begins at 7,000 is misleading, because it makes the index decline at the end of 2008 look more dramatic than it actually was. While the fall in the stock market index was significant, using a y-axis that begins with zero puts it in better perspective. 13. “Jane Woodsman’s average grade has increased from 13.8% last semester to 16.6% this semester.” The provocative language of the initial statement (“astonishing progress”, “substantial 20%”) is inappropriate. As well, the 20% figure, used as it is here, suggests something different from the facts. Jane’s grades did increase by 20%, but this is only 20% of the original grade of 13.8%, so it is not much of an improvement. 14. Aside from the fact that you should be suspicious of anyone who will not share the actual data with you, the local manager’s assurance that all is well may not be borne out by fact. Notice that the decrease claimed is in maximum wait times, not average wait times. It is entirely possible that average waiting times have increased. You need to see the data! 15. Yes. There is no distortion in how the data are represented, and the graph is clearly labelled and easy to understand. Develop Your Skills 1.4 16. No! With such an observational study, this kind of conclusion about cause and effect is not justified. There could be many factors (other than income) that explain why children in wealthier families are better off. For example, parents in wealthier

Copyright © 2011 Pearson Canada Inc.

3


Instructor’s Solutions Manual - Chapter 1

families may have more confidence, and this may provide a very positive environment in which children flourish. 17. No. Even if the study was randomized (no information is provided), it would not be legitimate to make this conclusion. While taller men are more likely to be married, we cannot conclude that they are more likely to be married because they are taller. There could be many other factors at work. 18. It may be that the diary system contributed to increased sales. Because the data compares the same people before and after use of the diary system, there is some support for this conclusion. However, notice that only poor performers were selected for the trial. These people may have worked harder simply because it was clear their poor performance had been noticed. 19. If you had compared the sales performance of a randomly-selected group of salespeople (not only poor performers), you would be able to come to a stronger conclusion about the diary system’s impact on increased sales. 20. There may be a cause-and-effect relationship here, but any conclusions should be made cautiously. For example, hotter weather or a nearby fair for children could have increased foot traffic (and sales) during the period. Develop Your Skills 1.5 21.a. In this case, the national manager of quality control probably has a good grasp of statistical approaches. While you should still strive for clarity and simplicity, you can include more of the technical work in the body of the report. Printouts of computer-based analysis would be included in the appendix. b. While human resources professionals probably have some understanding of statistical analysis, they are less likely to understand the details. In this case, you should write your report with a minimum of statistical jargon. The body of your report should contain key results, but the details of your analysis should be saved for the appendix. c. In this case, you can assume no statistical expertise in your readers. While you should still report on how your analysis was conducted, and how you arrived at your conclusions, you probably would not send this part of the report to your customers. The report you send your boss should be easily understandable to everyone. The challenge here will be to make your conclusions easy to understand, while not oversimplifying, or suggesting that your results are stronger than they actually are. 22. It is incorrect to suggest that study of a random sample “proves” anything. This statement is much more definitive than can be justified. As well, the study was done on past customers, and may not apply to future customers. Nevertheless, such a study could be persuasive about what segment of the market the company should focus on.

Copyright © 2011 Pearson Canada Inc.

4


Instructor’s Solutions Manual - Chapter 1

23. The average amount of paint in a random sample of 30 cans was 3.012 litres, compared with the target level of 3 litres. This sample mean is within control chart limits, indicating no need to adjust the paint filling line. 24. Analysis of output for a random sample of 50 workers showed an increase in worker output from 52 units per hour, on average, to 56 units per hour after the training. This evidence is sufficient to suggest that the worker output increased after training. This may mean that the training caused the increase in output, but this can be determined only after an examination of the circumstances, to determine if there were other possible causes of the increase in output. 25. In fact, some studies have shown that there is a positive relationship between height and income (that is, taller people tend to have higher incomes). However, all such studies must be observational (there is no ethical way to control height!), and so the cause-and-effect conclusion suggested here is not valid. The statement could be rewritten as follows: “A study has shown a strong positive relationship between height and income.” You might even go on to discourage the unsophisticated reader from jumping to conclusions, as follows: “Of course, this should not be interpreted as meaning that greater height guarantees higher income, or that you cannot earn a high income if you are short.” Chapter Review Exercises 1. Collecting data usually leads to a better understanding of the question and a better decision. 2.

Businesses may not have all of the data available, because it is impossible or too costly to collect. For example, it is unlikely that a business would have detailed information available about every customer. However, reliable decisions can be made on the basis of detailed information about a random sample of customers.

3.

It is generally not valid to draw conclusions about a population on the basis of a convenience sample. Because there is no way to estimate the probability of any particular element of the population being selected for a convenience sample, we cannot control or estimate the probability of error.

4.

Students often use their average grade to summarize their performance in a semester.

5.

Decision-makers may not be statisticians. Statistical analysis is powerful only if it is communicated so that those making the decision can understand the story the data are telling.

6.

One of the difficulties with gathering data through personal interviews is that those being surveyed are sometimes charmed by the interviewer. This sometimes results in interviewees acting to please the interviewer, rather than providing honest or informed responses to questions. Respondents can also be misled by the interviewer, as Rick Mercer so successfully illustrates in the “Talking to Americans” segments.

Copyright © 2011 Pearson Canada Inc.

5


Instructor’s Solutions Manual - Chapter 1

For example, Rick Mercer managed to persuade Americans to congratulate Canadians on legalizing VCRs. 7.

It is possible to make reliable conclusions on the basis of 1,003 responses. The sample size may seem small, when you consider there are millions of adult Canadians who have retirement funds. However, the sample size required depends not on the size of the population, but on its variability. If all Canadians were exactly the same, a sample of one would be sufficient. The more variable Canadians are, the larger the sample size required for estimates of a desired level of accuracy. This is something you will explore more in Chapter 8.

8.

No. The applications are filled out by employers, not employees. Employers have a vested interest in portraying themselves as “top” employers. The sample is not representative. The Top 100 are selected from a self-selected sample.

9.

Done correctly, such a study could identify a positive relationship between height and income. However, it is not correct to suggest the study means that height is the cause of the differences in incomes, although if the study was well-designed, other causes could have been randomized. Still, the leap to evolution as the explanation is not at all justified by the study. While this may be the explanation, it is only one of many possible explanations.

10. You collected your sample results at only one physical location at the school. Since it was near the smoking area, it is likely that your sample contained a disproportionate number of smokers. It is likely that smokers’ opinions about the new smoking policy would differ from nonsmokers’ opinions. Your sample is almost certainly not representative of the entire school community. 11. There are many examples of loyalty programs: Airmiles, President’s Choice Financial rewards, American Express rewards, HBC rewards, PetroCanada’s PetroPoints, Sears Club points, Aeroplan. Enter “loyalty rewards programs” into an Internet search engine, and you will find references to many such programs. 12. This is an observational study. It is not possible to draw a strong conclusion that drinking alcohol causes higher earnings. In fact, the causation may run the other way: higher income may cause more drinking (because those with higher incomes can afford to drink alcohol). The study measured income, not social networks, so the explanation provided is speculative. As well, there are other factors that could explain the differences in incomes, and these were not controlled in the study. 13. The article describes the New Coke story as "the greatest marketing disaster of all time". The research failed to uncover the attachment people felt to the original Coke. A question such as "Would you switch to the New Coke?" might have revealed how loyal customers were to the original Coke.

Copyright © 2011 Pearson Canada Inc.

6


Instructor’s Solutions Manual - Chapter 1

14. The Conference Board included the detail so that anyone reading the study could draw their own conclusions about the reliability of the data, and possible biases in the study results. 15. While the title used in the report is accurate, the percentage decrease is relatively small, and the actual number of drivers has increased. The title could easily be misunderstood. 16. The author has used Excel to generate Lotto 6/49 quick picks, but she didn't win the lottery either! 17. Of course, because your samples are randomly-selected, your samples cannot be predicted. When the author did this exercise, the sample averages were as shown in the table below. The population average is 65.0. The sample averages ranged from 60.7 to 73.7. Some were quite close to the population average, and the largest difference between a sample average and the true value was 8.7.

Average Mark from 10 Randomly-Selected Samples of Size 10 60.7 61.8 65.5 65.6 67.4 69.8 70.3 71.3 73.0 73.7

18. Again, the values you obtain cannot be predicted. These averages should be closer to the true population average, because they are based on more data (15 data points instead of 10). The author's results are shown in the table below. In general, the sample averages are closer to the true population average. The largest difference between a sample average and the true value is -6.4.

Average Mark from 10 Randomly-Selected Samples of Size 15 58.6 61.2 62.6 64.0 66.2 66.9 67.2 67.7 68.4 70.8

19. Based on the author's results for exercises 17 and 18 above (yours will be different): the average of the sample averages, when the sample size is 10, is 67.9. The average of the sample averages, when the sample size is 15, is 65.4. The average of the sample averages is closer to the true population average value when the sample size is larger. You will investigate this more in Chapter 6.

Copyright © 2011 Pearson Canada Inc.

7


Instructor’s Solutions Manual - Chapter 1

20. a. Because students are generally quite mobile, often moving from the place where they attended to school to their place of work after graduating, it may be quite difficult to contact them for the graduation employment and satisfaction measures. Some graduates will almost certainly be missed, and if their opinions differ from those surveyed, the results will not be truly representative. The same kinds of problem arise with surveys of current students. Such students are surveyed in their classes, so those who are absent on the day of the survey are missed. It could be that the opinions of those who are absent are different from the opinions of those who are present in class (for example, it could be the case that students are absent because they do not find their classes relevant to their future careers). Employers may also be missed, so the potential for coverage errors exists throughout. Of course, as with any survey, all the other nonsampling errors are possibilities: nonresponse errors, response errors, processing and estimation errors. This is a large undertaking, and so there are more possibilities for error. b.

Although it could be the case that all colleges improved their services significantly between 1999-2000 and later years, this seems unlikely. There may be another explanation for the shift in the percentage of students satisfied with college services. In fact, a closer inspection reveals that about a half-dozen colleges improved their ratings quite significantly. Without more information, there is no way to know why this happened.

c.

The report uses summary measures (percentages of responses for each category, for each year) and graphs (line graphs and bar graphs) to summarize the data.

d.

In 1999-2000, an additional capstone question was included in the calculation. Since this question was not used in subsequent years, the average student satisfaction rates are not directly comparable.

Copyright © 2011 Pearson Canada Inc.

8


Instructor’s Solutions Manual - Chapter 2

Chapter 2 Solutions Develop Your Skills 2.1 1. The number of dented cans is a count of qualitative data (think of it this way—the original data might be recorded as “yes” or “no” to the question: is the can dented?). These are also time-series data, as they are collected over successive periods of time. The data are discrete. 2.

Stock price data are quantitative data. These are also time-series data, as they were collected over three years. Prices are treated as continuous data.

3.

The employees' final average grades from college are continuous quantitative data. The scores assigned by the supervisors are ranked data. Although the questionnaire will help, such a ranking is somewhat subjective. If different supervisors assign the ranks, they may not be comparable.

4.

The price data are continuous quantitative data. They are also organized according to qualitative data on the size of the coffee. These are cross-sectional data, as they would be collected at around the same period in time.

5.

Postal codes are qualitative data.

Copyright © 2011 Pearson Canada Inc.

9


Instructor’s Solutions Manual - Chapter 2

Develop Your Skills 2.2 6. The three different histograms are shown below.

Survey of Drugstore Customers: Customer Ages 18 16

Number of Customers

14 12 10 8 6 4 2 0

Age of Customers

Survey of Drugstore Customers: Customer Ages Number of Customers

30 25 20 15 10 5 0

Age of Customers

Survey of Drugstore Customers: Customer Ages 35

Number of Customers

30 25 20 15 10 5 0

Age of Customers

Copyright © 2011 Pearson Canada Inc.

10


Instructor’s Solutions Manual - Chapter 2

All three histograms clearly show that the distribution of customer ages is skewed to the right, that is, while most customers are under 40 years old, there are some customers who are much older, in fact as old as 85. [Note that when you are describing the distribution, it is not sufficient to stop at “skewed to the right”—you should explain what this means, in the context of this particular data set.] A class width of 5 is not a good choice for this data set. There are too many classes, many of which have only a very few data points. A class width of 10 or 15 would be a better choice. 7.

A frequency distribution and histogram are shown below.

Survey of Drugstore Customers Customer Income Number of Customers $30,000 to <$35,000 2 $35,000 to <$40,000 9 $40,000 to <$45,000 14 $45,000 to <$50,000 7 $50,000 to <$55,000 5 $55,000 to <$60,000 6 $60,000 to <$65,000 3 $65,000 to <$70,000 4

Survey of Drugstore Customers: Customer Incomes 16 14

Number of Customers

12 10 8 6 4 2 0

Customer Income

Choosing class widths is a bit tricky in this instance. The class width template suggests class widths of 5155, 8905, or 8170. None of these numbers is that comfortable for incomes. Class widths of $5,000 and $10,000 were considered. A class width of $10,000 was discarded because it would have resulted in only four classes (five is a good minimum number of classes).

Copyright © 2011 Pearson Canada Inc.

11


Instructor’s Solutions Manual - Chapter 2

The distribution of customer incomes in the drugstore survey is skewed to the right. Most customer incomes are in the $35,000 to $50,000, but there are a number of customers with higher incomes, the highest being $68,800. 8.

The stem and leaf display is shown below.

1 2 3 4

9 4 1 0

7 6 1 1

4 3 4

8 9 1

2 4 6

7 9

7

6

5

9

4

6

The lowest daily customer count in the random sample from Downtown Automotive is 12, and the largest is 41. Most days the shop deals with 20-some customers. It is unusual for the shop to deal with more than 35 customers. 9.

This histogram totally fails at its job of summarizing the accompanying data set. 1. The graph does not have meaningful titles or labels. It completely fails to communicate what it’s about. 2. There are gaps between the bars, which there should not be. 3. It appears that the creator of this graph used bin numbers correctly, but s/he forgot to round them for presentation. The graph should show lower class limits along the x-axis in the proper location, that is, aligned under the left-hand side of each bar.

Copyright © 2011 Pearson Canada Inc.

12


Instructor’s Solutions Manual - Chapter 2

10.

These graphs have not been properly set up, so Patty probably deserved her low mark in Statistics. The titles are not correct. They are not Ms. Nice’s marks, they are marks from a random sample of students in Ms. Nice’s statistics class, and the title should tell us that. The title for the marks from Mr. Mean’s class should be similarly adjusted. In both graphs, the label on the x-axis should say something like “Final Grade”. The label on the y-axis should say something like “Number of Students”. In the graph of the marks from Ms. Nice’s class, the labels are at the centre of each class, and should be adjusted. Horizontal grid lines would also help the reader. The graphs should be set up properly for comparison, with the same classes and scales on each axis. A quick glance at these graphs might lead you to think that the marks are lower in Mr. Mean’s class, but in fact, the opposite is the case.

Copyright © 2011 Pearson Canada Inc.

13


Instructor’s Solutions Manual - Chapter 2

Develop Your Skills 2.3 11. Either a bar graph or a pie chart would be appropriate.

Number of Customers

Survey of Drugstore Customers: Speed of Service Ratings 20 18 16 14 12 10 8 6 4 2 0 Excellent

Good

Fair

Poor

Rating

Survey of Drugstore Customers: Speed of Service Ratings Poor 18%

Excellent 6%

Good 38%

Fair 38%

The graphs indicate that over a third of customers (38%) rated the speed of service as good, but only 6% rated it as excellent. Over a third of customers (again, 38%) rated the speed of service as only fair, while 18% rated it as poor. These ratings indicate that there may be some room for improvement in speed of service at the drugstore.

Copyright © 2011 Pearson Canada Inc.

14


Instructor’s Solutions Manual - Chapter 2

12. The most effective graph would be a bar chart, showing actual and desired relative frequencies for each colour. First, create a table for the data, and then create the bar chart.

Actual and Observed Colours in Candy Red Green Blue Observed sample values Desired Percentage Sample Percentage

305

265

Yellow

201

96

40% 30% 20% 10% 35.179% 30.565% 23.183% 11.073%

Candy Colours, Random Sample After Reorganization of Production Process 45% 40% 35% 30% 25% 20% 15% 10% 5% 0%

Desired Percentage Sample Percentage

Red

Green

Blue

Yellow

This graph makes it easy to see that the most important differences in the candy colour distribution are in the red candies (fewer than desired) and in the blue candies (more than desired).

Copyright © 2011 Pearson Canada Inc.

15


Instructor’s Solutions Manual - Chapter 2

13. Since we want to compare the number of defects by shift, it is appropriate to compare the categories for the number of defects across the horizontal axis. Since each shift produced a different number of items 1 , it makes sense to use relative frequencies. For each shift, calculate the percentage of items with no defects, one minor defect, more than one minor defect, and then create a bar graph.

Percentage of Total Number of Items Produced

Defects Observed at a Manufacturing Plant, by Shift 100% 8:00 a.m. – 4 p.m.

90% 80% 70%

4:00 p.m. – midnight

60% 50% 40%

midnight – 8:00 a.m.

30% 20% 10% 0% Items With No Apparent Defects

Items With One Minor Defect

Items With More Than One Minor Defect

Across all three shifts, the percentage of items produced with more than one minor defect is small. For all three shifts, by far the greatest percentage of items produced has no apparent defects. The midnight-8:00 a.m. shift has the greatest percentage of defects, and the 4:00 p.m. – midnight shift has the lowest percentage of defects.

1

This too is interesting, and the fact should be included in any accompanying report.

Copyright © 2011 Pearson Canada Inc.

16


Instructor’s Solutions Manual - Chapter 2

14.

Financial Services Company Customer Survey: "The staff at my local branch can provide me with good advice on my financial affairs." Number of Customers

60 50 40 30 20 10 0 Strongly Agree

Agree

Neither Agree nor Disagree

Disagree

Strongly Disagree

The majority of the customers surveyed agree with the statement that staff at the local branch can provide good advice on financial affairs. A significant number of the customers neither agreed nor disagreed with the statement, and it might be worthwhile to investigate why these customers appeared to have no opinion (was it lack of knowledge?). There were customers who disagreed or strongly disagreed. It might be worthwhile to investigate further (why was this the case? Were these customers disappointed in past advice, or do they just have an impression that local staff cannot provide good advice?).

Copyright © 2011 Pearson Canada Inc.

17


Instructor’s Solutions Manual - Chapter 2

15.

Survey of a Random Sample of People Walking Around Kempenfelt Bay 14

Number of People

12 10 8 6 4 2 0 Vanilla

Chocolate Strawberry Maple Chocolate Pralines Walnut Chip

Other

Favourite Flavour of Ice Cream

Vanilla and chocolate were tied as the most frequently-mentioned favourite flavours of ice cream among the people surveyed at Kempenfelt Bay, followed by chocolate chip. The fourth most popular flavour was strawberry. Only a few people cited maple walnut as their favourite flavour, and no one called pralines their favourite. One person had a favourite flavour other than the ones cited specifically in the survey. Of course, there are other options for this graphical display. It might be helpful to arrange the categories from most preferred to least-preferred. As well, a pie graph is an option.

Survey of a Random Sample of People Walking Around Kempenfelt Bay 14

Number of People

12 10 8 6 4 2 0 Vanilla

Chocolate Chocolate Strawberry Maple Chip Walnut

Other

Pralines

Favourite Flavour of Ice Cream

Copyright © 2011 Pearson Canada Inc.

18


Instructor’s Solutions Manual - Chapter 2

Survey of a Random Sample of People Walking Around Kempenfelt Bay Other 2%

Pralines 0%

Chocolate Chip 21%

Vanilla 28%

Maple Walnut 4% Strawberry 17%

Copyright © 2011 Pearson Canada Inc.

Chocolate 28%

19


Instructor’s Solutions Manual - Chapter 2

Develop Your Skills 2.4 16. Automobile sales are seasonal and cyclical, although this may not be as much the case as it once was. Sales tend to be higher when new models become available, and generally, auto sales are lower at year-end, when many people are focused on the holiday season. For these reasons, monthly sales data would be appropriate. Annual data would hide the month-to-month variations in sales. 17. Whatever your data source (as long as the data are accurate) you should see that over this period, the price of $1US in Canadian dollars was on an increasing trend from January 2000 until the beginning of 2002, with the highest exchange value of 1.599618 (monthly average) in January of 2002. From then until near the end of 2007, the exchange value of the US dollar in terms of Canadian dollars was on a declining trend, reaching 0.968 in November of 2007. The rate then stabilized around par, beginning to increase in the latter part of 2008, ending with a monthly average rate of 1.2343619 in December. The graph below shows the trends.

Copyright © 2011 Pearson Canada Inc.

20


Instructor’s Solutions Manual - Chapter 2

18. The Bank of Canada Bank Rate was at 4.5% in January 2007, and stayed there until July of 2007, when it increased to 4.75%. The rate declined to 4.5% in December of 2007, and continued to decline to 3.25% in April of 2008. The rate held steady at 3.25% until October of 2008, when it declined to 2.5%, further falling to 1.75% in December of 2008. Usually a line graph would be used for such a long time series. However, in this case, the movements in the Bank Rate are infrequent, and small, so the graph is not too cluttered. The advantage in using a bar chart is that it highlights the change in rates from one period to the next. 19. Your commentary should describe the data for the company you chose. Here is a checklist to help you: - Be sure to note the start and end dates for the data. - Comment on at least a couple of specific values in the data set (e.g., the high and low for the period). - Keep your language objective and descriptive. Do not leap to any conclusions about why the data might look the way they do.

Copyright © 2011 Pearson Canada Inc.

21


Instructor’s Solutions Manual - Chapter 2

20.

Computer Price Index, Canada, Consumers (2002 = 100) 140 Computer Price Index

120 100 80 60 40 20 Jan‐08

Jul‐07

Jan‐07

Jul‐06

Jan‐06

Jul‐05

Jan‐05

Jul‐04

Jan‐04

Jul‐03

Jan‐03

Jul‐02

Jan‐02

0

Your graph should look something like this. Be sure that it is labelled completely and correctly. For example, it is important to indicate the base year for any price index. Your commentary should note that this price index has declined significantly since 2002, with the index hitting a low of 21.91 in May of 2008. The decline in the price index was most pronounced at the beginning of the period, in 2002. The rate of the decline in this price index has slowed somewhat at the end of the period (late 2007 and early 2008).

Copyright © 2011 Pearson Canada Inc.

22


Instructor’s Solutions Manual - Chapter 2

Develop Your Skills 2.5 21.

Monthly Spending on Restaurant Meals

Spending on Restaurant Meals and Income $250 $200 $150 $100 $50 $‐ $‐

$1,000

$2,000

$3,000

$4,000

$5,000

Monthly Income

There appears to be a slight positive relationship between monthly income and monthly spending on restaurant meals, that is, the higher the monthly income, the greater the monthly spending on restaurant meals. However, there is a great deal of variability in the spending on restaurant meals, and the relationship is weak.

Copyright © 2011 Pearson Canada Inc.

23


Instructor’s Solutions Manual - Chapter 2

22.

Jack's Cookies, Daily Sales 70

Quantity Sold

65 60 55 50 45 $0.40

$0.50

$0.60

$0.70

$0.80

$0.90

$1.00

$1.10

Price per Cookie

Note this graph shows the data with the explanatory variable (price) on the x-axis, and the response variable (quantity sold) on the y-axis, which matches convention. However, if you have taken economics, you might recognize this as a demand curve. For historical reasons, a demand curve is normally graphed with price on the y-axis. So, it is also acceptable to graph these data as follows:

Jack's Cookies, Daily Sales $1.10

Price per Cookie

$1.00 $0.90 $0.80 $0.70 $0.60 $0.50 $0.40 45

50

55

60

65

70

Quantity Sold

Notice that in both graphs, the axes do not begin at (0, 0). It is reasonable to scale the axes as shown, but this should always be clearly indicated. The data for daily sales of Jack’s Cookies show that the quantity sold and the price are negatively related, that is, the higher the price, the lower the quantity sold. This conforms to the Law of Demand.

Copyright © 2011 Pearson Canada Inc.

24


Instructor’s Solutions Manual - Chapter 2

23.

Survey of Drugstore Customers Amount of Most Recent Purchase

$45 $40 $35 $30 $25 $20 $15 $10 $5 $‐ $20,000 $30,000 $40,000 $50,000 $60,000 $70,000 $80,000 Customer Annual Income

There does not appear to be a strong relationship between the customer’s income and the amount of the most recent purchase. Note the scale on the x-axis does not start at zero. If you think about it, this should not come as a surprise. While we might have expected those customers with greater incomes to have higher purchases, this effect is less likely to appear for a single purchase. There may be more of a positive relationship between annual income and total annual drugstore purchases. 24.

Semester Average Mark (%)

Hours of Work and Semester Marks, Random Sample of Students 100 80 60 40 20 0 0

100

200

300

400

Total Hours of Paid Work During Semester

There appears to be a negative relationship between the total hours of paid work during the semester, and the semester average mark, that is, the greater the hours of work, the lower the semester average mark.

Copyright © 2011 Pearson Canada Inc.

25


Instructor’s Solutions Manual - Chapter 2

25. Exhibit 2.70c is not correct, because the explanatory variable is years of service, and it should be graphed on the x-axis. Exhibit 2.70b is probably not correct, because it depicts a negative relationship, that is, those with more years of service earn lower salaries. Exhibit 2.70a is the only possible choice, as it shows higher salaries associated with longer years of service. Develop Your Skills 2.6 26. The first obvious problem with this graph is the 3-D aspect. It makes it hard to read the height of the bars. It is not clear if the bars for “good” and “fair” are the same height. The graph, with the 3-D aspect, is an appropriate way to represent these data. The graph would be improved if the 3-D aspect removed, as shown below.

Speed of Service Ratings, Survey of Drugstore Customers 20

Number of Customers

18 16 14 12 10 8 6 4 2 0 Excellent

Good

Fair

Poor

Another possible improvement would be to calculate relative frequencies for each rating category.

Copyright © 2011 Pearson Canada Inc.

26


Instructor’s Solutions Manual - Chapter 2

27. The pictograph looks as follows:

The year 2000 dollar is worth just under half of the value of the 1980 loonie, but the total area of the year 2000 loonie is only about a quarter of the area of the year 1980 loonie, so the pictograph misleads the viewer. When the loonie is shrunk, it shrinks not only in height (as a bar in a bar graph would), it is also shrunk in width, so the image is not distorted. This decreases the area disproportionately. 28. This is a good graph. The labels and titles are clear, and the graph can be understood without reference to anything else. We can see that there appears to be a positive relationship between the total monthly sales for Hendrick Software salespeople and the number of sales contacts during the previous month. 29. This graph cannot be interesting, because we have no clue what it is about. We can see that the distribution is skewed, but that’s all. With no title, and no meaningful labels on the axes, the graph is useless.

Copyright © 2011 Pearson Canada Inc.

27


Instructor’s Solutions Manual - Chapter 2

30. There are quite a few categories in this data set, and a bar graph would be preferred. Also, the title is not correct. The survey was of favourite flavours of ice cream (not people). The labels on the pie slices contain the code for the flavour, which is unnecessary and only serves to clutter up the graph. There is also a spelling mistake in one of the labels: “maple walnut”. A better graph is shown below. It has the advantage of sorting the flavours from most favourite to least favourite.

Survey of Favourite Ice Cream Flavours, Random Sample of People Walking Around Kempenfelt Bay Number of People

12 10 8 6 4 2

Copyright © 2011 Pearson Canada Inc.

Other

Maple Walnut

Vanilla

Pralines and Cream

Chocolate Chip

Chocolate

Strawberry

0

28


Instructor’s Solutions Manual - Chapter 2

Chapter Review Exercises 1a. These data are qualitative, unranked, and cross-sectional. 1b. These data are qualitative, unranked, and cross-sectional. 1c. These data are quantitative, discrete, and cross-sectional. 1d. These are ranked qualitative data. 1e. These are time-series continuous quantitative data. 2a. A double bar graph could show males and females along the x-axis, with two bars above, one for those with fitness club membership, one bar for those without. Alternatively, categories of fitness club membership could show along the x-axis, with bars for males and females above each category. 2b. A bar graph could be organized with the four store locations along the x-axis, and bars above, each one corresponding to the type of payment. Alternatively, the payment types could show along the x-axis, with four bars above each, one for each store location. 2c. If the total number of pedestrians is recorded, there are only two data points, the number of people who passed by each location. A graph would not really add much to a simple table displaying these numbers, with a proper title and headings. 2d. A double bar graph could be used, with the ratings ("barely edible" to "absolutely delicious") showing along the x-axis, and two bars above each rating, one for each chef. 2e. It is likely that there is interest in the relationship between sales and advertising. A scatter diagram would be appropriate, with advertising along the x-axis, and sales on the y-axis. 3.

These graphs are meant to be amusing and entertaining. Quirky images and bright colours make them attractive, but they are not good examples of graphs to summarize data.

Copyright © 2011 Pearson Canada Inc.

29


Instructor’s Solutions Manual - Chapter 2

4.

The stem and leaf display is shown below. The order of the leaves in your display may be different, if you went through the data by rows instead of columns.

0 1 2 3 4 5 6

9 9 2 1 3 8 4

8 2 1 2 0 0

8 2 0 0

9 3 0

The data set is skewed to the right. Those who are under 30 years old are the largest age groups in the sample. There are only two people in their 40’s and two in their 50's, and only one in the 60’s. 5.

It appears that there is a positive correlation between months of experience in the company, and salary, but the correlation does not appear to be particularly strong. In fact, the two observations with the highest salaries give the appearance of a positive correlation, and without them, there is no obvious relationship.

6.

Since these are quantitative data, histograms are required. Since the data are quite different in range, it is a challenge to decide what class width to use. The graphs below show a class width of $10, which is probably too wide for the data for purchases by males, but allows comparison with the purchases by females. Note that it might have been wise to use relative frequencies, rather than frequencies, to make this comparison, since the data sets have different sizes. However, the sample sizes differ by only one, so it is not crucial to do this here.

Copyright © 2011 Pearson Canada Inc.

30


Instructor’s Solutions Manual - Chapter 2

Music Store Purchases by Females 9 8 Number of Purchases

7 6 5 4 3 2 1 0

Value of Purchase

Music Store Purchases by Males 9 8 Number of Purchases

7 6 5 4 3 2 1 0

Value of Purchase

The histograms show that there is more variability in the music store purchases by females. As well, there are more purchases of higher value for females than males. The purchases by males are in the $10-$50 range, while the purchases by females are in the $10-$60 range.

Copyright © 2011 Pearson Canada Inc.

31


Instructor’s Solutions Manual - Chapter 2

7.

There are two possible graphical displays, a bar chart or a pie chart. Both are shown below. The pie chart has been formatted for black and white printout.

Ratings from a 360 Degree Review for a Trainee Number of Ratings

5 4 3 2 1 Best Possible Performance

Very Good, Very Little Improvement Required

Good

Acceptable

Poor, With Major Improvement Required

Worst Possible Performance

0

Ratings from a 360 Degree Review for a Trainee Best Possible Worst Possible Performance Performance 6% 7% Very Good, Very Little Improvement Required 20%

Good 27%

Copyright © 2011 Pearson Canada Inc.

Poor, With Major Improvement Required 27%

Acceptable 13%

32


Instructor’s Solutions Manual - Chapter 2

Whichever graphical display is used, it is apparent that the trainee’s ratings are not consistent. About 67% of raters indicated that the trainee's performance was acceptable or better. However 27% suggested that major improvement was required, and 6% rated the trainee's performance as the worst possible. Certainly, there seems to be a wide range of opinions about this trainee. 8.

An appropriate graphical display is shown below.

Employee Ratings of Previous and Current Presidents Number of Ratings

4 Ratings for the Previous President's Performance Ratings for the Current President's Performance

3 2 1 0 1

2

3

4

5

6

7

8

9

10

1= Worst Performance, 10= Best Possible Performance

The performance ratings for the new president are generally lower than for the previous president. However, there seems to be great variability in the ratings for both presidents.

Copyright © 2011 Pearson Canada Inc.

33


Instructor’s Solutions Manual - Chapter 2

9.

In this case, since the number of students in each sample is the same, it is appropriate to compare the number of students directly. An appropriate graph is shown below. The graph shows that the B.C. students were much more likely to rate this university as “excellent” than the Ontario students, with Ontario students much more likely to rate it as “poor”. The Ontario and B.C. students have different opinions about this university.

Ratings of a Canadian University by Ontario and BC Students 9

Ontario Student Ratings

Number of Students

8 7

BC Student Ratings

6 5 4 3 2 1 0 Excellent

Copyright © 2011 Pearson Canada Inc.

Good

Fair

Poor

34


Instructor’s Solutions Manual - Chapter 2

10. A graph to summarize the data is shown below.

Payments by Type at Four Store Locations 60

Percentage of Payments

50 40 Cash/Debit Card

30

Credit Card Cheque

20 10 0 Store A

Store B

Store C

Store D

In this case, since the total number of payments is different at the stores, percentage of payments is displayed on the graph, so that the values are directly comparable. The graph shows that the percentage of payments by cash or debit card is highest at Store A, at 40% of payments, and lowest at Store C, accounting for only 20% of payments. The percentage of payments made by credit card is 30% at both Stores A and C, and is 40% at Stores B and D. Cheques account for 50% of the payments at Store C, which is higher than at any other store. The other three stores have a similar percentage of payments by cheque, from 27.5% to 35%.

Copyright © 2011 Pearson Canada Inc.

35


Instructor’s Solutions Manual - Chapter 2

11. Since the samples are different sizes, relative frequencies must be used to make the comparison.

Percentage of Students in Program

Origins of Students in Two College Programs 60% From Local Area

50% 40%

Not From Local Area

30% 20% 10% 0% Business

Technology

Nursing

All three program areas draw a greater percentage of students from outside the local area, although the tendency is strongest for the Business program (about 57% of students not from the local area) and weakest for Nursing (about 51% of students not from the local area).

Copyright © 2011 Pearson Canada Inc.

36


Instructor’s Solutions Manual - Chapter 2

12. The use of the glass with a swizzle stick does not make the graph more interesting, it just makes it more difficult to read. It is quite difficult to judge the level of operating revenues from the pictures—is it the top of the glass or the top of the swizzle stick that we should read? A bar graph (or a line graph) would be a better choice to display these data, as shown below.

Soft Drink Company, Net Income $7,000

Net Income ($ Millions)

$6,000 $5,000 $4,000 $3,000 $2,000 $1,000 $0 2004

Copyright © 2011 Pearson Canada Inc.

2005

2006

2007

2008

37


Instructor’s Solutions Manual - Chapter 2

13. Two histograms, properly set up for comparison, are shown below.

Marks for a Random Sample of Students in Ms. Nice's Statistics Class 12

Number of Marks

10 8 6 4 2 0

Final Grade (%)

Marks for a Random Sample of Students in Mr. Mean's Statistics Class 12

Number of Marks

10 8 6 4 2 0

Final Grade (%)

Notice that the graphs are set up with the same x- and y-axis scales, for direct comparison. They are also similarly sized, so that it is possible to make a direct visual comparison. Class widths of 10 were used, because these are comfortable for marks data, and they allow us to make a distinction between passing and failing grades (assuming 50 is a pass).

Copyright © 2011 Pearson Canada Inc.

38


Instructor’s Solutions Manual - Chapter 2

The marks of the students from Mr. Mean’s class are generally higher and less variable than the marks of the students from Ms. Nice’s class. Half of the students from Ms. Nice’s class failed the course, while only two of the students from Mr. Mean’s class failed. 14. The two histograms are shown below.

Daily Pedestrian Traffic at Location 1 14

Number of Days

12 10 8 6 4 2 0

Number of Pedestrians

Daily Pedestrian Traffic at Location 2 14

Number of Days

12 10 8 6 4 2 0

Number of Pedestrians

(Note that the histograms are set up with matching x- and y-axes, and are sized similarly, for ease of comparison. Because the two locations were surveyed for the same number of days, we can compare the numbers directly.)

Copyright © 2011 Pearson Canada Inc.

39


Instructor’s Solutions Manual - Chapter 2

The histograms clearly show that daily pedestrian traffic is more variable at Location 2 than at Location 1. At Location 1, the daily traffic is in the 75-150 range, while at Location 2, it is in the 45-195 range. Generally, it appears the daily traffic at Location 1 is less than at Location 2. For both locations, the histograms are reasonably symmetric. 15. An appropriate histogram is shown below.

Downtown Automotive, Random Sample of Daily Sales 12

Number of Days

10 8 6 4 2 0 Daily Sales

For Downtown Automotive, daily sales are usually above $1,000, with sales falling into the $1,000 to < $1,500 class on 10 of the 29 days in the sample. The distribution is somewhat skewed to the right, that is, there are a few days when sales are above $2,000. Daily sales range from $690 to $2,878.

Copyright © 2011 Pearson Canada Inc.

40


Instructor’s Solutions Manual - Chapter 2

16. The appropriate graph is a line graph, such as the one shown below. It covers the 10year period ending in May 2008. You will have more recent data available.

Average Retail Prices of Regular Unleaded Gasoline at Self‐Service Filling Stations in Montreal 160

Cents per Litre

140 120 100 80 60 40 20 Dec‐07

Jun‐07

Jun‐06

Dec‐06

Jun‐05

Dec‐05

Dec‐04

Jun‐04

Dec‐03

Jun‐03

Dec‐02

Jun‐02

Jun‐01

Dec‐01

Dec‐00

Jun‐00

Dec‐99

Jun‐99

Dec‐98

Jun‐98

0

Over the 10-year period from June 1998 to May 2008, retail gas prices have been rising. The lowest price over the period was 52.2¢ per litre, in February of 1999, and the highest price was $1.36 per litre, in May of 2008. Prices rose fairly rapidly over the end of 1999 and the beginning of 2000, and then stayed fairly steady until June of 2001. At that point, prices fell, from 84.5¢ per litre in May of 2001 to 61.9¢ in November of 2001. They then began to climb again, reaching a high of $1.185 in September of 2005. Retail gas prices in Montreal showed great variability in the range between 88.5¢ per litre to $1.145 per litre through 2006 and 2007, with a sharp increase in April and May of 2008.

Copyright © 2011 Pearson Canada Inc.

41


Instructor’s Solutions Manual - Chapter 2

17. In this case, while there may be an association between the two variables, the causality link would not be strong. It would not be correct to say that a high mark in Business Math caused a high mark in Statistics, because there is very little overlap between the content of the two courses. However, a student with good study habits and good class attendance might do better in both courses. In this case, while the Business Math mark is not really the explanatory variable, since this course came first, we will put it on the x-axis. A graph of the data is shown below.

Marks from First and Second Year for a Random Sample of Students 120

Statistics Mark (%)

100 80 60 40 20 0 0

20

40

60

80

100

120

Business Math Mark (%)

There does appear to be a positive relationship between the two marks.

Copyright © 2011 Pearson Canada Inc.

42


Instructor’s Solutions Manual - Chapter 2

18. A graph of the data is shown below.

Woodbon Furniture Company $140,000 $120,000

Annual Sales

$100,000 $80,000 $60,000 $40,000 $20,000 $0 $0

$1,000

$2,000

$3,000

$4,000

Annual Advertising Expenditure

It appears there is a positive correlation between advertising and sales, that is, when advertising expenditure is higher, annual sales are also higher. 19. (Choosing an appropriate class width for comparison takes some thought. $10,000 is probably too wide (resulting in only 4 classes), and $5,000 is probably too narrow. A class width of $7,500 was used for the two histograms shown on the next page. Because the samples are of different size, relative frequencies should be used for comparison.)

Copyright © 2011 Pearson Canada Inc.

43


Instructor’s Solutions Manual - Chapter 2

Percentage of Male Customers

Survey of Drugstore Customers, Annual Incomes of Males 50% 45% 40% 35% 30% 25% 20% 15% 10% 5% 0%

Annual Income

Percentage of Female Customers

Survey of Drugstore Customers, Annual Incomes of Females 50% 45% 40% 35% 30% 25% 20% 15% 10% 5% 0%

Annual Income

Annual incomes for female drugstore customers are generally in the $37,500 to < $45,000 class, which accounts for over 48% of female customers' incomes. Some incomes of female customers are higher, but this is unusual in the sample, so the distribution of female customers' incomes is skewed to the right. In contrast, the incomes of male drugstore customers are more variable. Incomes between $30,000 and < $60,000 account for over 86% of male customers' incomes, with incomes spread fairly evenly throughout this range. In general, greater percentages of male customers' incomes are in the higher classes.

Copyright © 2011 Pearson Canada Inc.

44


Instructor’s Solutions Manual - Chapter 2

Percentage of Male or Female Customers Surveyed

20. The appropriate graph is shown below. Note that the graph shows percentages of males and females, because of the different sample sizes.

Drugstore Customer Survey, Speed of Service Ratings 60% 50%

Percentage of Males

40%

Percentage of Females

30% 20% 10% 0% Excellent

Good

Fair

Poor

Approximately the same small percentage of male and female customers rated the speed of service at the drugstore as excellent (about 5% of male customers and about 7% of female customers). The largest group of female customers (about 55%) rated the speed of service as "fair", and the largest group of male customers (about 57%) rated the speed of service as "good". It appears that male and female customers rate the speed of service very differently at the drugstore.

Copyright © 2011 Pearson Canada Inc.

45


Instructor’s Solutions Manual - Chapter 2

21. Two histograms are shown below. Note that samples are the same size, so relative frequencies are not required.

Flight Delays Before Airport Upgrades 16 14

Number of Flights

12 10 8 6 4 2 0

Flight Delay in Minutes

Flight Delays After Airport Upgrades 16 14

Number of Flights

12 10 8 6 4 2 0

Flight Delay in Minutes

The histograms seem to indicate that flight delays have changed after the airport upgrade. Before the upgrade, flight delays were mainly in the 10 to < 40 minute range. Only two delays were less than 10 minutes, and three were more than 40 minutes (but less than 50 minutes). After the upgrades, there were seven delays less than 10 minutes, so a greater number of flights had shorter delays. As well, there was only one flight delayed more than 40 minutes. However, the number of flight delays

Copyright © 2011 Pearson Canada Inc.

46


Instructor’s Solutions Manual - Chapter 2

of 10 - < 20 minutes has been reduced from 7 to 3 after the upgrades, while the number of delays of 20 - < 30 minutes has increased from 13 to 14. The greater number of flights with delays less than 10 minutes indicates some reduction in delays, but results appear mixed.

22. The two graphs are shown below.

Rating by Students of College Experience 350

Number of Students

300 250 200

Excellent

150

Good

100

Fair

50

Poor

0 Business Studies

Computer Studies

Engineering Technology Studies

Percentage of Students in Program

Rating by Students of College Experience 90% 80% 70% 60% 50% 40% 30% 20% 10% 0%

Excellent Good Fair Poor Business Studies

Copyright © 2011 Pearson Canada Inc.

Computer Studies

Engineering Technology Studies

47


Instructor’s Solutions Manual - Chapter 2

Because the programs have different numbers of students, the first graph using student numbers distorts the comparison. For example, it appears as if Business Studies and Engineering Technology Studies students choose a rating of "good" equally. However, the second graph reveals that a greater percentage of Engineering Technology Studies students rate their college experience as good. While the relative sizes of the ratings for each individual program remain the same, comparisons across programs are not valid unless relative frequencies are used.

Copyright © 2011 Pearson Canada Inc.

48


Instructor’s Solutions Manual - Chapter 2

23. The three histograms are shown below. Quarterly Operating Profits 35

Canadian Oil and Gas Extraction and Support Activities, I 1988 to III 2008

Number of Quarters

30 25 20 15 10 5 0 Millions of Dollars

Quarterly Operating Profits Canadian Oil and Gas Extraction and Support Activities, I 1988 to III 2008 45 40

Number of Quarters

35 30 25 20 15 10 5 0

Millions of Dollars

Number of Quarters

Quarterly Operating Profits 50 45 40 35 30 25 20 15 10 5 0

Canadian Oil and Gas Extraction and Support Activities, I 1988 to III 2008

Millions of Dollars

Copyright © 2011 Pearson Canada Inc.

49


Instructor’s Solutions Manual - Chapter 2

All three histograms show the same general shape, that is, the distribution is rightskewed. In most quarters, operating profits in the oil and gas sector were below $1.5 billion, but there were much higher profits in some quarters. In the first histogram the classes may be too narrow, as there are very low frequencies in many of the classes. However, this histogram provides more information about the many quarters when operating profits were low, as there is a breakdown for below $1 billion, and from $1 billion to < $2 billion. This information is hidden in the histogram with the widest classes. It can be a challenge to decide on appropriate class widths when the distribution is very skewed. In the histogram with the widest classes, a lot of data is contained in the first class (half of the data points are there), and so these classes may be a bit wide. However, any one of these histograms would be acceptable. The particular choice depends on the focus of the analysis. 24. Because you will have more up-to-date data, we cannot provide the histograms for this question. However, you should use the histograms and commentary in the text as guidelines. Be sure to use the same class widths for your comparison, and size your histograms similarly. Choose a class width that works for both data sets (you will probably be able to use $2 billion as the class width, as in the text). Remember that your commentary should simply describe the data sets. Do not get carried away with speculation about why the data look the way they do.

Copyright © 2011 Pearson Canada Inc.

50


Instructor’s Solutions Manual - Chapter 3

Chapter 3 Solutions Develop Your Skills 3.1 1. Σy = 2 + 4 + 6 + 8 = 20 2.

3.

Σy2 = 22 + 42 + 62 + 82 = 120 (Σy)2 = (2 + 4 + 6 + 8) = 202 = 400 The answers are different because of the different order of operations y 20  5 n 4 x 16  4 n 4

4.

( x  4)  (1  4)  (3  4)  (5  4)  (7  4)  3  (1)  1  3  0 ( y  5)  (2  5)  (4  5)  (6  5)  (8  5)  3  (1)  1  3  0

5. a. b.

Consider the data set: 34, 67, 2, 31, 89, 35. For this data set, calculate: x  34  67  2  31  89  35  258 x 2  34 2  67 2  2 2  312  89 2  35 2  15756 x 258   43 n 6

c.

2  x  x 

2  258  15756 

2

d.

n 1

n

6

5

15756  11094  5

4662  932 .4  30.535 5

Develop Your Skills 3.2 6. The mean age is 41.2, the median age is 35.5, and the mode of the ages is 30. In this case, because the data set is severely skewed to the right (as we saw when we created the histogram of ages in Develop Your Skills 2.2, Exercise 6), the median is the better measure of central tendency.

7.

The mean income is $47,868.10, and the median income is $44,925. This data set is skewed to the right (as we saw when we created the histogram of incomes in Develop Your Skills 2.2, Exercise 7). As a result the unusually high incomes have pulled the mean to the right of the median. The median is the better measure of central tendency.

Copyright © 2011 Pearson Canada Inc.

51


Instructor’s Solutions Manual - Chapter 3

8.

From the stem and leaf display we constructed in Develop Your Skills 2.2, Exercise 8, we can see that this data set is slightly skewed to the right, but not much. This is reflected in the calculations of mean and median (when in doubt, calculate both!). The mean of this data set is 26.12, and the median is 26. Either would be acceptable as a measure of central tendency, but the mean is preferred, because its calculation depends on the value of every single data point in the data set.

9.

Because the mean and the median are almost equal, we expect the distribution to be symmetric.

10. Because the quarterly operating profits of the oil and gas sector are highly skewed to the right, the median of $1.816 billion is the appropriate measure of central tendency. Although the distribution of operating profits for the manufacturing sector is not as skewed, so we might have considered using the mean as a measure of central tendency, we must use the median so that we are comparing the same measure for both data sets. The median quarterly operating profit for the manufacturing sector is $8.909 billion. Generally, the quarterly operating profits are much higher for the manufacturing sector than for the oil and gas sector. Develop Your Skills 3.3 11. Since the age data are skewed to the right, the IQR is the best measure of variability. Using Excel calculations, we find: Q1 = 31 Q3 = 42 IQR = 11 The Empirical Rule could not be applied here, as the data are not symmetric and bellshaped.

12. Since the data are skewed to the right, the IQR is the best measure of variability. Using Excel calculations we find: Q1 = $40,350 Q3 = $55,400 IQR = $15,050 13. Since this data set is fairly symmetric with no obvious outliers, the standard deviation is the preferred measure of variability.

s

x 2 

x 2

n 1

n

Copyright © 2011 Pearson Canada Inc.

653 2 25  18381  17056 .36  1324 .64  55.19333  7.429 24 24 24

18381  

52


Instructor’s Solutions Manual - Chapter 3

14. Because the distribution is reasonably symmetric and bell-shaped, the Empirical Rule can be applied. You must create a histogram to check this. Shown below is one possible histogram for the data set.

Number of Days

Daily Customer Counts, Downtown Automotive 10 9 8 7 6 5 4 3 2 1 0

Number of Customers

15. The mean number of daily customers at Downtown Automotive is 26.12 (calculated for Develop Your Skills 3.2, Exercise 8). The standard deviation is 7.43 (the answer to Exercise 13 above). The Empirical Rule says that about 95% of the data points will lie within 2 standard deviations of the mean. x  2s  26.12 + 2(7.43) = 40.98 x  2s  26.12 - 2(7.43) = 11.26 If this sample is representative of the population, then 95% of the daily customer counts will be between 11.26 and 40.98. Since the data set is (more or less) symmetric, this means about 2½% of the data will lie below 11.26, and about 2½% will lie above 40.98. About 97.5% of the time, the maximum number of customers Doug would need to plan for is 41. Develop Your Skills 3.4 16. The scatter diagram showed some (not much) evidence of a positive relationship between household income and monthly spending on restaurant meals. Since the relationship appears to be linear, the Pearson r is the appropriate measure of association. Excel calculates it as 0.42. This is positive and less than 0.5, as we would expect, since the relationship is not very strong.

Copyright © 2011 Pearson Canada Inc.

53


Instructor’s Solutions Manual - Chapter 3

17. The only choice is b (-0.88). Choices a and c are incorrect, because they are positive and the relationship is clearly negative. Choice d is not correct, because the negative relationship is obviously fairly strong. 18. The Spearman rank correlation coefficient must be used here, since the data are ranked. The Spearman r (calculated with Excel) is 0.61. This indicates a positive relationship between the recruiter’s ranking and the supervisor’s ranking, but the relationship is not particularly strong. 19. These are quantitative data, and the graph created for Develop Your Skills 2.5, Exercise 24 shows a linear relationship. The Pearson r is the correct measure of association. Excel calculates it at -0.67 (note that you must check for linearity of the relationship before you calculate the Pearson r). There is a negative relationship between the two variables. The greater the number of hours of paid employment during the semester, the lower the semester average mark. 20. Exhibit 3.44b is the graph that corresponds to the negative correlation coefficient of -0.90. This is obvious, since it is the only graph of the three showing a negative relationship. Exhibits 3.44a and c share the same correlation coefficient of 0.73. This is interesting because the correlation probably “looks” stronger in Exhibit 3.44a. However, notice that these two graphs depict exactly the same data, but with the xand y-axes reversed. Realize that you cannot reliably “eyeball” the strength of a relationship. The correlation coefficient allows us to make much more precise comparisons. Chapter Review Exercises 1. The mean mark is quite a bit higher than the median mark. This suggests that the distribution of marks is skewed to the right. It is likely that there are a few unusually high marks in the distribution.

2.

The mean weekly sales for both businesses are similar, although the mean sales at the haircutting salon are a bit lower than at the day spa. However, the mean sales at the haircutting salon are much less variable than at the day spa. This would result in a greater number of weeks with higher sales for the haircutting salon, and as a result, it would be a better purchase (all other things being equal).

3.

The mean age is 26.05, and the median age is 20.5. This is as expected. Because the distribution of ages is skewed to the right, the mean is greater than the median. There are several modes in the data set: 8, 9, 12, 20. Clearly, the three lower modes are not good indications of central tendency in this data set. The standard deviation is 17.1. Calculation of the interquartile range (manual method) is as follows: The location of Q1 is 5.25, and its value is 12. The location of Q3 is 15.75, and its value is 38, so the IQR is 26.

4.

Because the Pearson r is higher for Don's data set, the correlation between test marks and calories consumed (for Don) will be higher than the correlation between test

Copyright © 2011 Pearson Canada Inc.

54


Instructor’s Solutions Manual - Chapter 3

marks and hours spent studying (for Jane). However, there is no obvious reason why eating more calories would result in higher test marks. There is a logical connection between hours spent studying and test marks, so this cause and effect relationship is stronger. 5.

First, remember that with sample data, we cannot absolutely prove anything. As well, although the correlation coefficient is low, this does not mean that there is no relationship between incomes and purchases. As discussed in the answer to Develop Your Skills Exercise 23 in Chapter 2, the lack of relationship between an individual purchase and annual income does not preclude the existence of a relationship between annual purchases and annual income.

6.

Both histograms showed some right-skewness, particularly the purchases by females. However, the mean and median purchases for both groups are similar. The mean purchase by females is $30.86, and the median purchase is $29.50. The mean purchase for males is $28.90, and the median purchase is $28.38. Because the means and medians are so close, we will use the mean as the measure of central tendency. On average, the purchases of males are slightly higher than the purchases of females. Because we used the mean for the measure of central tendency, we will use the standard deviation as the measure of variability. The standard deviation for purchases by females is $10.97, while the standard deviation for purchases by males is $5.39. As we saw in the histograms we created in Chapter 2 (Chapter Review Exercise 6), there is less variability in purchases by males than purchases by females.

7.

Because a histogram of the data is symmetric and bell-shaped, we can apply the Empirical Rule. If the sample is representative of the population, then we can expect that about 68% of the data lie within one standard deviation of the mean, that is, between 170 cm and 184.4 cm, with 32% divided between the two tails of the distribution. This means that about 16% of young men aged 18-24 would be shorter than 170 cm. Almost all of the heights would be within three standard deviations of the mean, that is, between 155.6 cm and 198.8 cm. Therefore, there would not be many young men aged 18-24 who were taller than 199 cm.

8.

This is a small data set, and so it is not possible to create a histogram to assess the shape of each location’s sales distribution. However, if you order the two data sets, in both cases, there are more observations on the high end of the range than elsewhere, suggesting some skewness. Therefore the interquartile range is probably the best measure of variability. If you do the calculations by hand, the results are as follows. Both data sets have 7 data points. Q1 location is the 0.25(n+1) = 0.25(8) = 2nd place Q3 location is the 0.75(8) = 6th place

Copyright © 2011 Pearson Canada Inc.

55


Instructor’s Solutions Manual - Chapter 3

Red Deer Vernon Q1 109.55 112.30 Q3 122.48 122.01 IQR 12.93 9.71 The Red Deer location’s sales are more variable than the Vernon location’s sales over the period. If you do the calculations with Excel, the numerical results are different but the conclusion is the same.

Q1 Q3 IQR 9.

Red Deer Vernon 112.23 114.795 122.42 121.675 10.19 6.88

The mean price of the inkjet printer cartridges is $26.93. The median price is $25.95.

10. The mean weight of the honey in the jars is 497.3 grams. While this is below 500 grams, it is not much below. Without knowing more about the variability of the weights of honey in the jars, we cannot make a conclusion about whether the jars are being consistently underfilled. When you master the techniques of Chapter 7, you will be able to decide.

Copyright © 2011 Pearson Canada Inc.

56


Instructor’s Solutions Manual - Chapter 3

11. This data set is reasonably symmetric and bell-shaped. The mean weekly sales for the sample of stores trying out the new marketing approach are $5101.07, with a standard deviation of $325.60. Applying the Empirical Rule, almost all of the sales would be between $4124.28 and $6077.85.

Weekly Sales for Stores with New Marketing Approach

Number of Stores

5 4 3 2 1 0

Weekly Sales

12. First we must assess the shape of the distribution. The histogram below shows a reasonably symmetric and bell-shaped data set.

Annual Days Off (Other Than Vacation) for a Random Sample of Employees

Number of Employees

12 10 8 6 4 2 0 Days Off

Copyright © 2011 Pearson Canada Inc.

57


Instructor’s Solutions Manual - Chapter 3

The mean days off is 6.04, with a standard deviation of 1.29. The sample mean is below the average days off in the past, so there is a reason to hope that there has been an improvement. However, we need to do a formal hypothesis test (covered in Chapter 7) to determine whether there is sufficient evidence to conclude that average days off among all employees has actually decreased. As well, even if we conclude that there has been a decrease in average days off, we cannot necessarily conclude that the wellness program is the cause. Other factors may account for the difference, such as a change in the workforce. 13. Since there are many tied values, it is a bit of a challenge to do the ranking process. The results are as follows. Ratings by Customers Rank Ratings by Bosses Rank 2 2 3 4 3 3 2 2 1 2

4 4 8 10 8 8 4 4 1 4

3 4 2 1 2 2 4 1 3 2

7.5 9.5 4.5 1.5 4.5 4.5 9.5 1.5 7.5 4.5

The Spearman rank correlation coefficient is -0.574. This indicates a negative correlation between the ratings by customers and the ratings by bosses, that is, the ratings tend to be higher by customers when the ratings by bosses are lower. However, the correlation coefficient indicates that this relationship is weak. 14. Since the data sets are fairly symmetric, the mean and the standard deviation are the appropriate measures of central tendency and variability. The results are shown below. Location 1 Location 2 Mean 108.4 124.4 Standard Deviation 17.3 29.6 Mean daily pedestrian traffic is higher at Location 2, at 124.4, compared with 108.4 at Location 1. The daily pedestrian traffic at Location 2 (standard deviation is 29.6) is also more variable than at Location 1 (standard deviation of 17.3).

Copyright © 2011 Pearson Canada Inc.

58


Instructor’s Solutions Manual - Chapter 3

15. Because the data are quantitative and appear to be linearly related, the Pearson r is the appropriate measure of association. The Pearson r is 0.958, indicating a high correlation between the mark in Business Math and in Statistics. 16. Because the data are quantitative and appear to be linearly related, the Pearson r is the appropriate measure of association. The Pearson r is 0.941, indicating a high correlation between annual advertising expenditure and annual sales. 17. The customer incomes are skewed to the right, with a few incomes much higher than the rest in the data set. Therefore, the median and the interquartile range are the appropriate measures. The median income of the drugstore customers is $44,925. The interquartile range is 15,050 (Excel) or 15,762.5 (by hand).

Copyright © 2011 Pearson Canada Inc.

59


Instructor’s Solutions Manual - Chapter 3

18. Before we can decide on appropriate numerical measures to compare the data sets, we must examine the shapes of the distributions, by creating histograms, as shown below. Note that these histograms are not appropriate for comparison of the distributions—they are just for deciding on the appropriate measures.

Kate's Clients' RRSP Holdings 60

Number of Clients

50 40 30 20 10 0

RRSP Holding

Wally's Clients' RRSP Holdings 40 35

Number of Clients

30 25 20 15 10 5 0

RRSP Holding

Copyright © 2011 Pearson Canada Inc.

60


Instructor’s Solutions Manual - Chapter 3

Since both data sets are reasonably symmetric, we can use the mean and the standard deviation to compare them. The results are shown in the table below. Kate’s Clients’ Wally’s Clients’ RRSP Holdings RRSP Holdings Mean $111,021.86 $101,092.89 Standard Deviation $ 32,050.79 $ 40,192.47 The mean holdings of Kate’s clients’ RRSPs are higher, at $111,021.86, than the mean holdings of Wally’s clients RRSPs, at $101,092.89. The variability of the RRSP holdings of Kate’s clients is less than for Wally’s clients (standard deviation of $32,050.79, compared with $40,192.47). 19. First, we must examine the shape of the distribution. A histogram shows a reasonably symmetric data set (see below).

Contents of a Sample of Soup Cans 14 12 Number of Cans

10 8 6 4 2 0

Contents in Millilitres

The mean measurement is 540.4 mL, with a standard deviation of 5.17 mL. The maximum measurement in the sample is 551 mL, and this does not give any cause for concern that the cans contain more than 556 mL. As well, if we apply the Empirical Rule, we note that almost all of the measurements would be between 524.9 mL and 555.9 mL, which again does not give any cause for concern that the cans contain more than 556 mL. A measurement of 530 mL is about two standard deviations below the mean. The Empirical Rule says that about 95% of the data will lie within two standard deviations of the mean, with the remaining 5% split between the two tails of the distribution. If this can be applied to the population data, then about 2½ % of the cans would contain less than 530 mL.

Copyright © 2011 Pearson Canada Inc.

61


Instructor’s Solutions Manual - Chapter 3

20. Once again, we must check the distribution of the data set to see if the Empirical Rule applies.

Contents of a Sample of Soup Cans 12

Number of Cans

10 8 6 4 2 0

Contents in Millilitres

Since the distribution is approximately bell-shaped and symmetric, we can apply the Empirical Rule. The mean measurement is 543.63 mL, with a standard deviation of 6.44 mL. There is one can of soup in the sample that contains more than 556 mL. Applying the Empirical Rule, we note that 95% of the soup cans would contain between 530.7 mL and 556.5 mL. This leaves 2½% of the soup cans with more than 556 mL, and 2½% of the soup cans with less than 530 mL. 21. Once again, the answers will depend on the most up-to-date data available when you are answering this question. (As a guide, the retail sector data that matches the data in the text for the manufacturing sector are discussed. Because of revisions, the more recent data sets may not exactly match these data. For example, the retail series was significantly changed between the time of the original download in May of 2009, and a subsequent download in November 2009. It is a challenge to come up with a class width for comparison for the two sectors, since quarterly operating profits are much smaller for the retail sector than for the manufacturing sector. The compromise choice of $1 billion is really too narrow for the manufacturing data, and not wide enough for the retail sector data. However, these histograms give a starting point for the analysis.)

Copyright © 2011 Pearson Canada Inc.

62


Instructor’s Solutions Manual - Chapter 3

Quarterly Operating Profits Canadian Retail Sector, I 1988 to III 2008 45 40 Number of Quarters

35 30 25 20 15 10 5 0

Millions of Dollars

Quarterly Operating Profits Canadian Manufacturing, I 1988 to III 2008

Number of Quarters

12 10 8 6 4 2 0

Millions of Dollars

The histograms show that quarterly operating profits are smaller for the Canadian retail sector than for the manufacturing sector. For over half the period, quarterly operating profits for the retail sector were < $2 billion. In over 60% of the quarters in the period under study, the quarterly operating profits of the Canadian manufacturing sector were $8 billion or more. The distributions of quarterly operating profits also differ. The distribution for the retail sector profits is skewed to the right, with profits above $3 billion in a few quarters. The distribution for the manufacturing sector profits is skewed to the left, with a few quarters where operating profits were unusually low (below $5 billion).

Copyright © 2011 Pearson Canada Inc.

63


Instructor’s Solutions Manual - Chapter 3

Median quarterly operating profits for the manufacturing sector were $8.909 billion, much greater than the median quarterly operating profits for the retail sector, at $1.827 billion. Quarterly operating profits for the manufacturing sector were much more variable over the period, with an interquartile range of $4.648 billion, compared with only $1.1 billion for the retail sector.

Quarterly Operating Profits of the Manufacturing Sector ($ Millions)

There appears to be some slight positive correlation between the quarterly operating profits of the two sectors, but it is not strong. The scatter diagram below illustrates.

Quarterly Operating Profits for Two Canadian Sectors, I 1988 to III 2008 $16,000 $14,000 $12,000 $10,000 $8,000 $6,000 $4,000 $2,000 $0 $0

$1,000

$2,000

$3,000

$4,000

$5,000

Quarterly Operating Profits of the Retail Sector ($ Millions)

The Pearson r is 0.42, confirming the impression from the scatter diagram of a weak positive relationship. When quarterly operating profits of the retail sector are higher, the quarterly operating profits of the manufacturing sector tend to be higher, but the correlation is weak. (Your comparison, with more up-to-date data, should contain all of the elements shown in this answer.)

Copyright © 2011 Pearson Canada Inc.

64


Instructor’s Solutions Manual - Chapter 4

Chapter 4 Solutions Develop Your Skills 4.1 1a. Sample space: 246 employees commute more than 40 km by car 350-246=104 commute  40 km by car P(randomly selected employee commutes > 40 km by car) = 246/350 = 0.7029 1b. Sample space: 150 employees arrange rides with others 350-150=200 ride alone P(randomly selected employee arranges rides with others) = 150/350 = 0.4286 Note that the sample space need not be more complicated than necessary. A full description of the sample space, for both worker characteristics (commuting distance and arranging rides) could look as follows. Commuting Characteristics of Car Part Manufacturing Plant Arrange Rides With Others Ride Alone Totals Commute > 40 km 246-135=111 135 246 104-65=39 200-135=65 350-246=104 Commute  40 km 150 350-150=200 350 Totals 2.

Sample space: We are interested only in managers, so the sample space is as follows: Only High-School Education: 0 Grad Degree Or Post-Grad Studies: 10 up To 4 Years Of Post-Secondary Education: 37 – 10 – 0 = 27 P(randomly selected manager has up to 4 years of post-secondary education)=27/37=0.7297

3.

Sample space: Professional Employees: 372 Managers: 37 (Note that "managerial" and "professional" are separate job classifications, so managers are not included in the count of professional workers.) Clerical Employees: 520-372-37 = 111 P(randomly selected employee is professional) = 372/520 = 0.7154 P(randomly selected employee is clerical) = 111/520 = 0.2135

Copyright © 2011 Pearson Canada Inc.

65


Instructor’s Solutions Manual - Chapter 4

4.

Sample space: Shoppers Doing A Quick Trip: 62 Shoppers Doing A Major Stock-Up: 13 Shoppers Doing A Fill-In Shop: 100-62-13=25 P(randomly selected customer is doing a fill-in shop) = 25/100 = 0.25

5.

Sample space: Loved Previous Math Courses: 56 Worked Very Hard In Previous Math Courses, But Did Not Enjoy Them, Or Thought Previous Math Courses Were Far Too Difficult, Or Equated Previous Math Courses With Sticking Needles In The Eyes: 225-56=169 P(randomly selected student did not love his/her previous math courses)=169/225=0.7511

Develop Your Skills 4.2 6. These probabilities are given: P(pays with credit card) =P(CC) = 0.80 P(buys something other than gas) P(OTG) = 0.25 P(pays with credit card and buys something other than gas) =P(CC & OTG)= 0.20 Want to know P(pays with credit card GIVEN buys something other than gas) = P(CC  OTG) P(CC & OTG) 0.20    0.80 P(OTG) 0.25

Also want to know P(buys something other than gas GIVEN pays with a credit card) =P(OTG  CC) P(CC & OTG) 0.20    0.25 P(CC) 0.80

To check for independence, compare P(CC  OTG) with P(CC). The two probabilities are equal. So, paying with a credit card and buying something other than gas are independent (not related).

Copyright © 2011 Pearson Canada Inc.

66


Instructor’s Solutions Manual - Chapter 4

7.

P(honey-nut flavour GIVEN family size box) P(honeynut & family size) 180    0.5714 P(family size) 315 To check for independence, compare this with 315  180 = 0.55 P(honey-nut flavour)  900 Since the two probabilities are NOT equal, the events are NOT independent, that is, the size of the box and the flavour are related, in terms of sales.

8.

P(female given employed in agriculture industry) 

96.5 96.5   0.2951 (96.5  230.5) 327

P(employed in public administration, given male) 

454  0.0503 9,021.3

P(employed in public administration) 

454  471.7 925.7   0.0541 (9,021.3  8,104.5) 17,125.8

Since P(employed in public administration, given male) ≠ P(employed in public administration), we can say that sex of the employee and industry of employment were not independent in Canada in 2008. However, the difference in probabilities is not that great (both are around 5%). However, we can see that P(employed in manufacturing) 

1,410.4  559.9 1,970.3   0.1150 (9,021.3  8104.5) 17,125.8

which is not equal to P(employed in manufacturing given male) 1,410.4   0.1563 9,021.3 This is more convincing evidence that sex of the employee was not independent of the industry of employment in Canada in 2008.

Copyright © 2011 Pearson Canada Inc.

67


Instructor’s Solutions Manual - Chapter 4

9.

It is easiest to proceed if we first compute row and column totals for the table.

Accounts Receivable for a Roofing Company amount age $5,000 $5,000 - <$10,000 $10,000 total < 30 days 12 15 10 37 30 - <60 days 7 11 2 20 60 days and over 3 4 1 8 total 22 30 13 65

For the accounts receivable at this roofing company: P(< 30 days ) = 37/65 = 0.5692 P( < 30 days  $5000) = 12/22 = 0.5455 Since the two probabilities are not equal, account age and amount are not independent. 10. First calculate the row and column totals to make the probability calculations easier. Customer Survey for a Dry Cleaning Company Service Will Use Will Not Use Total Rating Services Again Services Again Poor 174 986 1,160 Fair 232 928 1,160 Good 2,436 174 2,610 Excellent 754 116 870 Total 3,596 2,204 5,800 P(customer will use dry cleaning company’s services again) 3596   0.62 5800 P(customer will not use the dry cleaning company’s services again, given a rating of “good” or “excellent”) 174  116   0.0833 2610  870 This is interesting, because about 8% of customers are not planning on using the dry cleaning company’s services again, even though they rate the service as good or excellent. Clearly something other than the service is keeping these customers away.

Copyright © 2011 Pearson Canada Inc.

68


Instructor’s Solutions Manual - Chapter 4

To test for independence, we could compare P(customer will not use the dry cleaning company’s services again) with the conditional probability we just calculated above. P(customer will not use the dry cleaning company’s services again) =1 – P(customers will use dry cleaning company’s services again) =1 – 0.620 = 0.38  0.0833 Since these probabilities are not equal, the events are not independent. There is a relationship between the rating of the service and the tendency to use it again. Overall, 38% of customers don’t plan to use the service again. However, only about 8% of those who rated the service positively don’t plan to use it again. Develop Your Skills 4.3 11. We are told that the two employees live in different parts of the city, and so presumably could not be held up by the same traffic problems. Assume that each employee’s lateness is independent of the other’s lateness. Then P(both are late) =P(Jane is late and Oscar is late) =0.02 • 0.04 = 0.0008 The probability is low that both Jane and Oscar will be late for work.

12. We are told the friends are very different, and will assume that any one of them getting a job in the financial services industry is independent of the others getting such a job. Label the friends “1”, “2” and “3”. a.

P(all three of them succeed) =P(1 does and 2 does and 3 does) =0.4 • 0.5 • 0.35 = 0.07

b.

P(none of them succeeds) =P(1 does not and 2 does not and 3 does not) =(1-0.4)•(1-0.5)•(1-0.35) =0.6•0.5•0.65=0.195

c.

P(at least one of them succeeds) =1-P(none of them succeeds) =1-0.195= 0.805 Using the complement rule here is a life-saver!

Copyright © 2011 Pearson Canada Inc.

69


Instructor’s Solutions Manual - Chapter 4

13. P(game aimed at 15-25 year olds succeeding) = 0.34 P(accounting program for small business succeeding) = 0.12 P(payroll system for government organizations succeeding) = 0.10 We are told to assume that the events are independent. P(all three succeed) = 0.34 • 0.12 • 0.10 = 0.00408 For calculation of at least two out of three succeeding, we need to think about what this means, in terms of the sample space. At least two out of three succeeding means exactly two out of the three succeeding, or all three succeeding. A tree diagram might be helpful to picture this. We need to calculate and add the probabilities for the cases shown in bolded letters on the right-hand side of the tree diagram.

0.12

0.88

SSS

0.90

F

SSF

0.10

S

SFS

F

SFF F 0.90 accounting payroll program system FSS 0.10 S

game

0.66

S

S

S 0.34

0.10

0.12

S 0.90

F

FSF

0.10

S

FFS

0.90

F

FFF

F 0.88

F

Once we have the cases identified, it is just a matter of arithmetic. P(at least two out of three succeed) =(0.34 • 0.12 • 0.10)+(0.34 • 0.12 • 0.90)+(0.34 • 0.88 • 0.10)+(0.66 • 0.12 • 0.10) =0.00408+0.03672+0.02992+0.00792=0.07864

Copyright © 2011 Pearson Canada Inc.

70


Instructor’s Solutions Manual - Chapter 4

14. A fully-labelled tree diagram for the GeorgeConn customer data is shown below.

S

P(R and S)=0.3

P(NR)=1/4

N

P(R and N)=0.1

P(SU)=4/6

S

P(U and S)=0.4

N

P(U and N)=0.2

R

P(S and R)=0.3

P(US)=4/7

U

P(S and U)=0.4

P(RN)=1/3

R

P(N and R)=0.1

U

P(N and U)=0.2

P(SR)=3/4 R

P(R)=4/10

P(U)=6/10

U

P(NU)=2/6

Another way to set up the tree diagram:

P(RS)=3/7 S

P(S)=7/10

P(N)=3/10

N

P(UN)=2/3

Copyright © 2011 Pearson Canada Inc.

71


Instructor’s Solutions Manual - Chapter 4

15. P(hourly worker or only high school education) =P(hourly worker) + P(only high school education) – P(hourly worker and only high school education) = (790+265+2)/1345 + (790+7+1+0)/1345 – 790/1345 = (790 + 265 + 2 + 7 + 1)/1345 = 1065/1345 = 0.7918 Chapter Review Exercises

1.

P(account paid early) 119   0.1587 750 P(account paid on time) 320   0.4267 750 P(account paid late) 200   0.2667 750 P(account uncollectible) 111   0.1480 750

2.

The probability calculations may seem easier if you organize the information into a table, as follows.

Business Diploma No Business Diploma Total Men Women Total

30 25 55

a.

P(employee is a man) 60   0.60 100

b.

P(employee is a man with a Business diploma)

Copyright © 2011 Pearson Canada Inc.

30 15 45

60 40 100

72


Instructor’s Solutions Manual - Chapter 4

30  0.30 100

c.

P(employee is a woman) 40   0.40 100

d.

P(employee is a woman with a Business diploma) 25   0.25 100

e.

P(employee has a Business diploma) 55   0.55 100

f.

P(employee is a man without a Business diploma) 30   0.30 100

g.

P(employee is without a Business diploma) 45   0.45 100

h.

P(employee is a woman without a Business diploma) 15   0.15 100

i.

P(employee is a woman or employee with a Business diploma) 40 55 25     0.70 100 100 100

j.

P(employee is a man or employee with a Business diploma) 60 55 30     0.85 100 100 100

k.

P(employee is a woman or employee without a Business diploma) 40 45 15     0.70 100 100 100

l.

P(employee is a man or employee without a Business diploma) 60 45 30     0.75 100 100 100

3.

P(employee has Business diploma given she is a woman)

Copyright © 2011 Pearson Canada Inc.

73


Instructor’s Solutions Manual - Chapter 4

25  0.625 40

To test to see if gender and possession of a Business diploma are related for GeorgeConn employees, we can compare the probability above to P(employee has a Business diploma) 

55  0.55 100

P(employee has Business diploma given she is a woman)) ≠ P(employee has a Business diploma), so gender and possession of a Business diploma are related for GeorgeConn employees. Female employees are more likely to have a Business diploma. 4.

P(caller directly connected) = 0.80 P(caller forced to wait) = 0.20 P(caller connected directly for three different calls) = P(caller connected on first day AND connected on second day AND connected on third day) = 0.8 • 0.8• 0.8 = 0.512 P(caller forced to wait for three different calls) = P(caller forced to wait on first day AND forced to wait on second day AND forced to wait on third day) = 0.2 • 0.2• 0.2 = 0.008

5.

Start by totalling the rows and columns of the table. This will speed up the probability calculations.

a.

P(primary skill is bookkeeping) 30   0.30 100

b.

P(employee has less than one year of experience) 50   0.50 100

c.

P(primary skill is reception)

Copyright © 2011 Pearson Canada Inc.

74


Instructor’s Solutions Manual - Chapter 4

25  0.25 100

d.

P(employee has one to two years of experience) 23   0.23 100

e.

P(primary skill is document management) 45   0.45 100

f.

P(employee has more than two years of experience) 27   0.27 100

6.

Begin by totalling rows and columns. Survey of Restaurant Customers Opinion About Food Satisfied with Service Not Satisfied with Service Totals Excellent 0.36 0.06 0.42 Good 0.18 0.07 0.25 Fair 0.10 0.08 0.18 Poor 0.05 0.10 0.15 Totals 0.69 0.31

a.

P(customer is satisfied with service and rates the food as poor) = 0.05

b.

P(customer is not satisfied with service) = 0.31

c.

P(not satisfied with service given food rated as poor) 0.10   0.6667 0.15

d.

To test for independence, we could compare the probability in part c with P(not satisfied with service) =0.31 The two probabilities are not equal, so the service rating and the food rating are related (that is, NOT independent). People who rate the food as poor are more likely to be dissatisfied with the service.

Copyright © 2011 Pearson Canada Inc.

75


Instructor’s Solutions Manual - Chapter 4

7.

P(salesperson will exceed targets two years in a row) = P(salesperson exceeds target this year and exceeds target next year) = P(exceeds target this year) • P(exceeds target next year  exceeds target this year) = 0.78 • 0.15 = 0.0117

8.

Begin by summing rows and columns of the table. Follow-up Survey of Customers Who Bought Netbook Computers Satisfied Not Satisfied Totals 1 GB of RAM or Less 0.30 0.05 0.35 More than 1 GB of RAM 0.50 0.15 0.65 Totals 0.80 0.20

a.

P(satisfied with his/her purchase) = 0.30 + 0.50 = 0.80

b.

P(satisfied with his/her purchase more than 1 GB of RAM) = 0.50/(0.50 + 0.15) = 0.7692

c.

The amount of RAM affects whether or not the purchaser was satisfied with his/her purchase. P(satisfied  more than 1 GB of RAM) = 0.7692 ≠ P(satisfied) = 0.80. Customers who bought netbooks with more than 1 GB of RAM were less likely to be satisfied with their purchases. This may seem odd, because generally more RAM means a better computer. However, it is possible that customers buying machines with more RAM had higher performance expectations that could not be met with slower processors.

Copyright © 2011 Pearson Canada Inc.

76


Instructor’s Solutions Manual - Chapter 4

9.

A tree diagram is helpful: P(P)=0.75

P

P(P)=0.75

P(F)=0.25

P(PF)=0.90

P

F

P(FF)=0.10

F

P(F and P) =0.25 • 0.90 =0.225 P(F and F) =0.25 • 0.10 =0.025

P(pass with no more than two attempts) = P(pass the first time) + P(fail the first time and pass the second time) = 0.75 + 0.225 = 0.975 10. Again, begin by summing rows and columns in the table. Customers of an Insurance Company Single Married Divorced Totals Male 25 125 30 180 Female 50 50 20 120 Totals 75 175 50 a.

P(female or married) = P(female) + P(married) – P( female and married) = (50+50+20)/300 + (125 + 50)/300 – 50/300 = (50 + 50 + 20 + 125 + 50 - 50)/300 = 245/300 = 0.8167

b.

P(married  male) = P(married and male)/P(male) = (125/300) / (25 + 125 + 30)/300 = 125/ (180) =0.6944

c.

P(married) = (125 + 50)/300 = 0.5833 ≠ P(married  male) = 0.6944 Since the probabilities are not equal, gender and marital status are not independent.

Copyright © 2011 Pearson Canada Inc.

77


Instructor’s Solutions Manual - Chapter 4

11. R: market rises RC: market does not rise P: newsletter predicts rise PC: newsletter predicts market will not rise

P(PR)=0.70

P

P(R and P) =0.60 • 0.70 =0.42

R

P(R)=0.60

C

P(R )=0.40

P(R and PC) =0.60 • 0.30 =0.18

P(PCR)=0.30

PC

P(PRC)=0.30

P

P(RC and P) =0.40 • 0.30 =0.12

PC

P(RC and PC) =0.40 • 0.70 =0.28

RC P(PCRC)=0.70

P(correct prediction) = P(market rises and newsletter predicts rise) + P(market does not rise and newsletter predicts market will not rise) = 0.42 + 0.28 (from tree diagram) = 0.70 12. P(all of the students selected by the company are female) = 6/10 • 5/9 • 4/8 = 120/720 = 0.1667 P(all of the students selected by the company are male) = 4/10 • 3/9 • 2/8 = 24/720 = 0.0333

Copyright © 2011 Pearson Canada Inc.

78


Instructor’s Solutions Manual - Chapter 4

13. If we can identify one situation where gender and tendency to use the health facilities are related, we can say that gender and the tendency to use health facilities are related. Compare P(used the facilities) with P(used the facilities  male) P(used the facilities) = 210/350 = 0.6 P(used the facilities  male) = 65/170 = 0.3824 Since these two probabilities are not equal, gender and tendency to use the health and fitness facilities are not independent (that is, they are related). 14.

P(MU)=65/210

M

U

P(U and M) =65/350 =0.1857

P(U)= 210/350 P(FU)=145/210

P(MD)=105/140

F

M

P(D and M) =105/350 =0.3

F

P(D and F) =35/350 =0.1

P(D)= 140/350 D

P(FD)=35/140

P(U and F) =145/350 =0.4143

The joint probabilities are the same, as we would expect them to be. 15. One of the ways to test for independence (or lack of it) is as follows. P(purchased the product) 228   0.76 300 P(purchased the product given saw the TV ad) 152 152   = 0.8085 152  36 188

Copyright © 2011 Pearson Canada Inc.

79


Instructor’s Solutions Manual - Chapter 4

These two probabilities are not equal, so purchasing behaviour is related to seeing the TV ad. Those who saw the ad were more likely to purchase the product. 16.

A: Brenda moves to Alberta AC: Brenda does not move to Alberta B: Brenda is offered the job at Canada’s largest bank BC: Brenda is not offered the job at Canada’s largest bank

P(AB)=0.7

A

P(B and A) =0.25 • 0.7 =0.175

B

P(B)=0.25

P(ACB)=0.3

AC

P(ABC)=0.35

A

P(BC)=0.75 BC

P(ACBC)=0.65

C

A

P(B and AC) =0.25 • 0.3 =0.075 P(BC and A) =0.75 • 0.35 =0.2625 P(BC and AC) =0.75 • 0.65 =0.4875

From the tree diagram: P(Brenda will be offered the job and not move to Alberta) = 0.075

Copyright © 2011 Pearson Canada Inc.

80


Instructor’s Solutions Manual - Chapter 4

17. I: Canadian adult has taken instruction in canoeing IC: Canadian adult has not taken instruction in canoeing CT: Canadian adult is going on a canoe trip this summer CTC: Canadian adult is not going on a canoe trip this summer

P(CT I)=0.46

CT

P(I and CT) =0.03 • 0.46 =0.0138

I

P(I)=0.03

P(I and CTC) C =0.03 • 0.54 P(CTC I)=0.54 CT =0.0162 P(CT IC)=0.20

P(IC)=0.97

CT

P(IC and CT) =0.97 • 0.20 =0.194

CTC

P(IC and CTC) =0.97 • 0.80 =0.776

IC P(CTC IC)=0.80

18. P(a randomly-selected Canadian adult is going on a canoe trip this summer, and has taken some canoeing instruction) = 0.0138 P(a randomly-selected Canadian adult is going on a canoe trip this summer, and has not taken any canoeing instruction) = 0.194 19. P(a randomly-selected customer from one of these stores uses a cash/debit card or a credit card for payment) = (150 + 180)/500 = 0.66 20. If we can identify one situation where payment method and store location are related, we can say that they are related in general. One approach is to compare P(cheque) with P(cheque  Store A) P(cheque) = 170/500 = 0.34 P(cheque  Store A) = 30/100 = 0.30 Since these two probabilities are not equal, payment method and store location are not independent.

Copyright © 2011 Pearson Canada Inc.

81


Instructor’s Solutions Manual - Chapter 4

21. For people who visit the facility: P(buy a membership) = 0.40 P(buy a membership and sign up for fitness classes) = 0.30 P(fitness classes  bought a membership) =0.30/0.40 = 0.75 22. Buying a membership and signing up for fitness classes are NOT mutually exclusive. We are told that P(buy a membership and sign up for fitness classes) = 0.30 ≠ 0. A person can do both, so the events are not mutually exclusive. We cannot assess independence without more information. For example if we knew P(fitness classes), we could compare that with P(fitness classes  bought a membership). 23. If we can identify one situation where gender and type of alcoholic drink are related, we can say that they are related in general. However, there is no case where gender and type of alcoholic drink are related. P(wine) = (36 + 54)/(42 + 63 + 36 + 54 + 22 + 33)=90/250=0.36 P(wine  female) = 54/(63 + 54 + 33) = 0.36 P(wine  male) = 36/(42 + 36 + 22) = 0.36 So we can see that P(wine) = P(wine  female) = P(wine  male). Similarly, P(beer) = P(beer  female) = P(beer  male). As well, P(other alcoholic drinks) = P(other alcoholic drinks  female) = P(other alcoholic drinks  male). We cannot identify a situation where gender and type of alcoholic drink are not independent, so we conclude that gender and type of alcoholic drink ordered are independent in this sample. 24. 500 circuit boards, of which 30 are defective, 470 are not defective P(all three are defective) =30/500 • 29/499 • 28/498 =24,360/124,251,000 = 0.000196054 P(defective board on 1st selection) = 30/500 = 0.06 P(all 3 boards defective, assuming independence) = 0.06 • 0.06 • 0.06 = 0.000216 The probabilities agree, to four decimal places. 25. In this case, we are stuck. We have only one probability (25%), but we do not have independent events. Once the first randomly-selected Canadian is asked about RRSP plans, he/she is removed from further consideration. Depending whether this person plans to make an RRSP contribution over the next year, this will affect the 25% probability of making a contribution. However, it will not affect it very much,

Copyright © 2011 Pearson Canada Inc.

82


Instructor’s Solutions Manual - Chapter 4

because there are many millions of Canadians. So, although the events are not really independent, we can still use the probability as if they were. P(all four intend to contribute to their RRSPs over the next year) = 0.25 • 0.25 • 0.25 • 0.25 = 0.0039 26. This looks like a long and complicated question, but it isn't, as long as the information is organized properly. You can use the Sort tool (under the Data tab) in Excel to help organize the data as you require it (the use of the Sort tool was described in Chapter 1). You might also explore the use of the Filter tool (use Excel's Help function if you can't see how it works.) There are 24 employees in total. a.

P(employee has low experience) 13   0.5417 24 P(employee has high experience) 11   0.4583 24 P(employee has a specialty in spreadsheet software) 11   0.4583 24 P(employee has a specialty in presentation software) 4   0.1667 24

b.

There are 11 employees who specialize in spreadsheet software. P(first employee selected has high experience given specialization in spreadsheet software) 4   0.3636 11 P(second employee selected has high experience given specialization in spreadsheet software given first employee had high experience given specialization in spreadsheet software) 3   0.30 10 P(both employees selected have high experience, given specialization in spreadsheet software) 4 3 12    =0.1091 11 10 110

Copyright © 2011 Pearson Canada Inc.

83


Instructor’s Solutions Manual - Chapter 4

P(at least one of the two employees has high experience) = 1 – P(none of the employees has high experience) 42 68 7 6  1     1   0.6182 110 110  11 10  c.

The joint probability table is shown below.

Database Software Low High Totals

Presentation Software 0 1 1

Spreadsheet Software

Word Processing Software

4 0 4

7 4 11

Totals 2 6 8

13 11 24

P(high experience, given specialty is word processing) 6   0.75 8 P(low experience, given specialty is spreadsheet) 7   0.6364 11 d.

The joint probability table is shown below.

Database Software F M Totals

Presentation Software 0 1 1

Spreadsheet Software 3 1 4

5 6 11

Word Processing Software

Totals

4 4 8

12 12 24

P(word processing specialization given female) 4   0.3333 12

Copyright © 2011 Pearson Canada Inc.

84


Instructor’s Solutions Manual - Chapter 4

P(male given spreadsheet specialization) 6   0.5455 11 e.

Because the tree diagram has three stages, it takes up an entire page (see the next page). Notice that the end-stage probabilities could have been calculated directly from the table of information about the employees. It is useful to explore the structure of the sample space with the tree diagram. M: Male F: Female D: Specializes in Database Software P: Specializes in Presentation Software S: Specializes in Spreadsheet Software W: Specializes in Word Processing Software L: Low Experience H: High Experience

Copyright © 2011 Pearson Canada Inc.

85


Instructor’s Solutions Manual - Chapter 4

0/1 D

1/1

L

0

H

1/24

L

1/24

H

0

L

3/24

H

3/24

L

1/24

H

3/24

L

0

H

0

L

3/24

H

0

L

4/24

H

1/24

L

1/24

H

3/24

1/12 1/1 P

1/12 M

6/12

3/6 S

12/24

0/1

3/6

4/12 1/4 W

3/4

0/0 D

0/0

0/12 12/24

3/3 3/12 F

P

5/12

0/3

4/5 S

1/5

4/12 1/4 W

Copyright © 2011 Pearson Canada Inc.

3/4

86


Instructor’s Solutions Manual - Chapter 4

f.

This tree diagram is just another way of representing the sample space. The endstage probabilities match those in part e.

D

1/24

P

0/24

S

3/24

W

3/24

D

0/24

P

0/24

S

1/24

W

3/24

D

0/24

P

1/24

S

3/24

W

1/24

D

0/24

P

3/24

S

4/24

W

1/24

1/7 0/7 3/7

M 7/11

3/7 H 4/11

11/24

0/4 0/4 F

1/4 3/4

0/5 1/5 13/24

M

3/5

5/13 1/5

L 8/13

0/8 3/8 F

4/8 1/8

Copyright © 2011 Pearson Canada Inc.

87


Instructor’s Solutions Manual - Chapter 5

Chapter 5 Solutions Develop Your Skills 5.1 1. a. Discrete. The number of passengers on a flight from Toronto to Paris is a count. b. Continuous. The time it takes you to drive to work in the morning could take on any one of an infinite number of possible values, within some range (shortest possible trip to longest possible trip). c. Discrete. The number of cars who arrive at the local car dealership for an express oil change service on Wednesday is a count. d. Continuous. The time it takes to cut a customer’s lawn could take on any one of an infinite number of possible values, within some range (smallest easiest lawn to largest most difficult lawn). e. Discrete. The number of soft drinks a student buys during one week is a count. f. Continuous. The kilometres driven on one tank of gas could take on any one of an infinite number of possible values, within some range (shortest possible distance to longest possible distance). 2. x 0 1 2 3 P(x) 0.195 0.43 0.305 0.07 P(x=3)=0.07 from previous calculations P(x=0)=0.195 from previous calculations P(x=1) =P(1 does, 2 does not, 3 does not)+P(1 does not, 2 does, 3 does not)+P(1 does not, 2 does not, 3 does) =(0.4 • 0.5 • 0.65) + (0.6 • 0.5 • 0.65) + (0.6 • 0.5 • 0.35) =0.13 + 0.195 + 0.105 = 0.43 P(x=2) = 1 – P(x = 1 or 1 or 3) = 1 – 0.195 – 0.43 – 0.07 = 0.305

Copyright © 2011 Pearson Canada Inc.

88


Instructor’s Solutions Manual - Chapter 5

3.

0.12

S

S 0.88

0.34 game

0.66

0.12

S

P(SSS)=0.00408

0.90

F

P(SSF)=0.03672

0.10

S

P(SFS)=0.02992

F

F P(SFF)=0.26928 0.90 accounting payroll program system 0.10 S P(FSS)=0.00792 S F P(FSF)=0.07128 0.90

F 0.88

0.10

0.10

S

P(FFS)=0.05808

0.90

F

P(FFF)=0.52272

F

P(x=0) = 0.52272 P(x=3) = 0.00408 P(x=1) = 0.26928 + 0.07128 + 0.05808 = 0.39864 P(x=2) = 1- 0.52272 - 0.00408 - 0.39864 = 0.07456 x 0 1 2 3 P(x) 0.52272 0.39864 0.07456 0.00408 The expected number of successes is  = 0(0.52272) + 1(0.39864) + 2(0.07456) + 3(0.00408) = 0.56

Copyright © 2011 Pearson Canada Inc.

89


Instructor’s Solutions Manual - Chapter 5

4.

D: defective circuit board found OK: circuit board not defective

D

P(D and D) =(30 • 29)/(5000 • 4999)

P(OK D) =4970/4999

OK

P(D and OK) =(30 • 4970)/(5000 • 4999)

P(D OK) =30/4999

D

P(OK and D) =(4970 • 30)/(5000 • 4999)

P(D D) =29/4999

D P(D)=30/5000

P(OK)=4970/5000

OK P(OK OK) OK

P(OK and OK) =(4970 • 4969)/(5000 • 4999)

=4969/4999

x 0 1 2 P(x) 0.9880348 0.0119304 0.0000348

Copyright © 2011 Pearson Canada Inc.

90


Instructor’s Solutions Manual - Chapter 5

5. No. of customers who order the daily special at a restaurant, out of the next 6 customers x 0 1 2 3 4 5 6 P(x) 0.03 0.05 0.28 0.45 0.12 0.04 0.03 P(x=6) = 1 – 0.03 – 0.05 – 0.28 – 0.45 – 0.12 – 0.04 = 0.03   x  P ( x )  0  0.3  1  0.05  2  0.28  3  0.45  4  0.12  5  0.04  6  0.03  2.82   x 2 P ( x )   2  (0 2  0.03  12  0.05  2 2  0.28  3 2  0.45  4 2  0.12  5 2  0.04  6 2  0.03)  2.82 2  9.22  7.9524  1.2676  1.1259

Develop Your Skills 5.2 6. In this case, sampling is without replacement, but we assume the college has thousands of students. The sample size is probably much less than 5% of the population, so the binomial distribution can still be used to approximate the probabilities. According to the newspaper, p = 0.8 n = 10 P(x  4, n=10, p=0.8) = 0.0064. This is a very unlikely result, if the newspaper’s claim about 80% support is true. The evidence from the sample casts doubt on the newspaper’s claim. 7.

P(pass) =P(x13, n=25, p=1/5) =1 - P(x12) =1- 1 = 0 (using the tables) With Excel, we get the slightly more accurate result of 0.000369048. Whichever, it would be basically impossible to pass this test by guessing. Does this result change any ideas you might have had that multiple “guess” tests are easy?

Copyright © 2011 Pearson Canada Inc.

91


Instructor’s Solutions Manual - Chapter 5

8.

In this case, sampling is without replacement, but we assume the tire plant produces thousands and thousands of tires. The sample size is probably much less than 5% of the population, so the binomial distribution can still be used to approximate the probabilities. n = 20 p = 0.05 P(x=1) = P(x  1) – P(x  0) = 0.736 – 0.358 = 0.378, using the tables Using Excel, P(x=1) = 0.3774 Using the formula:  20  P( x  1)   0.051 0.9519 1 20!  0.051 0.9519 1!19!  20(0.05)(0.377353603)  0.3774

9.

In this case, it is likely that the respondents to the poll on losing weight would not be a random sample, but rather a subset of the population of visitors to the site. Therefore, we should not apply the probability from this sample to all visitors to the site.

10. In this case, sampling is without replacement, but we assume there are thousands of managers in the population. The sample size from the poll is probably much less than 5% of the population, so the binomial distribution can still be used to approximate the probabilities. n = 30 p = 0.342 P(x 10) = 0.544782141, from Excel

Develop Your Skills 5.3 11.  = 5,000,  = 367 P( x  6000) 6000  5000    P z   367    P( z  2.72)  1  0.9967  0.0033 Only 0.33% of the bulbs will last more than 6000 hours. With Excel, we find P(x  6000) = 0.0032.

Copyright © 2011 Pearson Canada Inc.

92


Turn static files into dynamic content formats.

Create a flipbook