Paper For Above instruction
Question A:
You will create a set of 5 points that are very close together and calculate the standard deviation. Then, add a sixth point that is far from the original five, and recalculate the standard deviation. The impact of adding the distant point is that the standard deviation increases. Specifically, when the points are close, the standard deviation reflects minimal variability around the mean because the data points are similar with little dispersion. Once a distant point is introduced, the variability among the data points increases significantly, which in turn raises the standard deviation because the points are more spread out from the mean.
Question B:
You will create two data sets, each with 8 points and a mean close to 10. The first data set should have a standard deviation around 1, and the second around 4. To achieve the larger standard deviation, you introduce more variability among the data points, spreading them further from the mean. In contrast, the dataset with a smaller standard deviation has data points clustered more tightly around the mean.
Next, you will clear previous data from Question 1 and input the data set: 50, 50, 50, 50, 50. The standard deviation in this case is zero because all data points are identical, meaning there is no dispersion or variability among them. Standard deviation measures the average distance of each data point from the mean; if all points are the same, this distance is zero for each point, resulting in a standard deviation of zero.
Then, you will analyze three new data sets:
Data set 1: 0, 0, 0, 100, 100, 100
Data set 2: 0, 20, 40, 60, 80, 100
Data set 3: 0, 40, 45, 55, 60, 100
Each data set has a median of 50, but the spread varies. The relationship observed is that the greater the spread of data points (how far apart they are from the median), the larger the standard deviation. This connection exists because standard deviation quantifies how much data points deviate from the mean; more spread out data leads to higher dispersion, which increases the standard deviation.
For the last two questions, you will use the Project 1 Data Set.
Question 4:
Explain what an outlier is. Then, determine if there are any outliers in the Project 1 Data Set. If none, state that there are no outliers.
Question 5:
Identify four states with temperatures that seem most questionable or unrealistic. Justify your choices by providing the state names and temperatures, explaining why these are questionable based on their values relative to typical temperature ranges or patterns.
References
Everitt, B. S. (2002). The Cambridge Dictionary of Statistics. Cambridge University Press.
Weiss, N. A. (2012). Introductory Statistics. Pearson Education.
Carver, R. (2018). Standard Deviation and Variance. In Statistical Concepts. Routledge.
Ghasemi, A., & Zahediasl, S. (2012). Normality testing for statistical analysis: a guide for non-statisticians. International Journal of Endocrinology and Metabolism, 10(2), 486–489.
Hollander, M., & Wolfe, D. A. (1999). Nonparametric Statistical Methods. Wiley-Interscience.
Field, A. (2013). Discovering Statistics Using IBM SPSS Statistics. Sage Publications.
Olejnik, S., & Algina, J. (2003). Generalized Eta and Omega Squared Statistics for Factorial ANOVA. Psychological Methods, 8(4), 434–447.
Lowry, R. (2014). Discovering Statistics Using R. Sage Publications.
Wilcox, R. R. (2012). Introduction to Robust Estimation and Hypothesis Testing. Academic Press.
Wikipedia contributors. (2023). Outlier. In Wikipedia. https://en.wikipedia.org/wiki/Outlier