• Home
  • Textbooks
  • Statistics Informed Decisions Using Data
  • Inference on Categorical Data

Statistics Informed Decisions Using Data

Michael Sullivan III

Chapter 12

Inference on Categorical Data - all with Video Answers

Educators


Section 1

Goodness-of-Fit Test

01:08

Problem 1

True or False: The shape of the chi-square distribution depends on the degrees of freedom.

Harsh Gadhiya
Harsh Gadhiya
Numerade Educator
01:26

Problem 2

A ____________ test is an inferential procedure used to determine whether a frequency distribution follows a specific distribution.

Harsh Gadhiya
Harsh Gadhiya
Numerade Educator
01:25

Problem 3

Suppose there are $n$ independent trials of an experiment with $k>3$ mutually exclusive outcomes, where $p_{i}$ represents the probability of observing the ith outcome. The _______ _________ for each possible outcome are given by $E_{i}=$ _______.

Harsh Gadhiya
Harsh Gadhiya
Numerade Educator
00:53

Problem 4

What are the two requirements that must be satisfied to perform a goodness-of-fit test?

Harsh Gadhiya
Harsh Gadhiya
Numerade Educator
01:45

Problem 5

Determine the expected counts for each outcome.
$$\begin{array}{llll}n=500 & & & \\\hline p_{i} & 0.2 & 0.1 & 0.45 & 0.25 \\\hline \text { Expected counts } & & & & \\\hline\end{array}$$

Harsh Gadhiya
Harsh Gadhiya
Numerade Educator
01:27

Problem 6

Determine the expected counts for each outcome.
$$\begin{array}{llll}n=700 & & & \\\hline p_{i} & 0.15 & 0.3 & 0.35 & 0.20 \\\hline \text { Expected counts } & & & & \\\hline\end{array}$$

Harsh Gadhiya
Harsh Gadhiya
Numerade Educator
09:18

Problem 7

$H_{0}: p_{\mathrm{A}}=p_{\mathrm{B}}=p_{\mathrm{C}}=p_{\mathrm{D}}=\frac{1}{4}$
$H_{1}:$ At least one of the proportions is different from the others.
$$\begin{array}{lllll}\text { Outcome } & \mathbf{A} & \mathbf{B} & \mathbf{C} & \mathbf{D} \\\hline \text { Observed } & 30 & 20 & 28 & 22 \\\hline \text { Expected } & 25 & 25 & 25 & 25\end{array}$$

Harsh Gadhiya
Harsh Gadhiya
Numerade Educator
10:59

Problem 8

$H_{0}: p_{\mathrm{A}}=p_{\mathrm{B}}=p_{\mathrm{C}}=p_{\mathrm{D}}=p_{\mathrm{E}}=\frac{1}{5}$
$H_{1}:$ At least one of the proportions is different from the others.
$$\begin{array}{lccccc}\text { Outcome } & \mathbf{A} & \mathbf{B} & \mathbf{C} & \mathbf{D} & \mathbf{E} \\\hline \text { Observed } & 38 & 45 & 41 & 33 & 43 \\\hline \text { Expected } & 40 & 40 & 40 & 40 & 40\end{array}$$

Harsh Gadhiya
Harsh Gadhiya
Numerade Educator
11:14

Problem 9

$H_{0}:$ The random variable $X$ is binomial with $n=4, p=0.8$
$H_{1}:$ The random variable $X$ is not binomial with $n=4$ $p=0.8$
$$\begin{array}{llllll}\boldsymbol{X} & \mathbf{0} & \mathbf{1} & \mathbf{2} & \mathbf{3} & \mathbf{4} \\\hline \text { Observed } & 1 & 38 & 132 & 440 & 389 \\\hline \text { Expected } & 1.6 & 25.6 & 153.6 & 409.6 & 409.6\end{array}$$

Harsh Gadhiya
Harsh Gadhiya
Numerade Educator
11:24

Problem 10

$H_{0}:$ The random variable $X$ is binomial with $n=4, p=0.3$
$H_{1}:$ The random variable $X$ is not binomial with $n=4$ $p=0.3$
$$\begin{array}{lccccc}\boldsymbol{X} & \mathbf{0} & \mathbf{1} & \mathbf{2} & \mathbf{3} & \mathbf{4} \\\hline \text { Observed } & 260 & 400 & 280 & 50 & 10 \\\hline \text { Expected } & 240.1 & 411.6 & 264.6 & 75.6 & 8.1\end{array}$$

Harsh Gadhiya
Harsh Gadhiya
Numerade Educator
16:39

Problem 11

According to the manufacturer of M\&Ms, $13 \%$ of the plain $\mathrm{M} \& \mathrm{Ms}$ in a bag should be brown, $14 \%$ yellow, $13 \%$ red, $24 \%$ blue, $20 \%$ orange, and $16 \%$ green. A student randomly selected a bag of plain $\mathrm{M} \& \mathrm{Ms}$. He counted the number of $\mathrm{M} \& \mathrm{Ms}$ that were each color and obtained the results shown in the table. Test whether plain M\&Ms follow the distribution stated by M\&M/Mars at the $\alpha=0.05$ level of significance.
$$\begin{array}{lc}\text { Color } & \text { Frequency } \\\hline \text { Brown } & 61 \\\hline \text { Yellow } & 64 \\\hline \text { Red } & 54 \\\hline \text { Blue } & 61 \\\hline \text { Orange } & 96 \\\hline \text { Green } & 64\end{array}$$

Harsh Gadhiya
Harsh Gadhiya
Numerade Educator
16:24

Problem 12

According to the manufacturer of M\&Ms,12\% of the peanut M\&Ms in a bag should be brown, $15 \%$ yellow, $12 \%$ red, $23 \%$ blue, $23 \%$ orange, and $15 \%$ green. A student randomly selected a bag of peanut $\mathrm{M} \& \mathrm{Ms}$. He counted the number of $\mathrm{M} \& \mathrm{Ms}$ that were each color and obtained the results shown in the table. Test whether peanut M\&Ms follow the distribution stated by M\&M/Mars at the $\alpha=0.05$ level of significance.
$$\begin{array}{lc}\text { Color } & \text { Frequency } \\\hline \text { Brown } & 53 \\\hline \text { Yellow } & 66 \\\hline \text { Red } & 38 \\\hline \text { Blue } & 96 \\\hline \text { Orange } & 88 \\\hline \text { Green } & 59\end{array}$$

Harsh Gadhiya
Harsh Gadhiya
Numerade Educator
20:03

Problem 13

Our number system consists of the digits $0,1,2,3,4,5,6,7,8,$ and $9 .$ The first significant digit in any number must be $1,2,3,4,5,6,7,8,$ or 9 because we do not write numbers such as 12 as 012. Although we may think that each digit appears with equal frequency so that each digit has a $\frac{1}{9}$ probability of being the first significant digit, this is not true. In $1881,$ Simon Newcomb discovered that first digits do not occur with equal frequency. This same result was discovered again in 1938 by physicist Frank Benford. After studying much data, he was able to assign probabilities of occurrence to the first digit in a number as shown.
$$\begin{array}{lccccc}\text { Digit } & 1 & 2 & 3 & 4 & 5 \\\hline \text { Probability } & 0.301 & 0.176 & 0.125 & 0.097 & 0.079 \\\hline \text { Digit } & 6 & 7 & 8 & 9 & \\\hline \text { Probability } & 0.067 & 0.058 & 0.051 & 0.046 & \\\hline\end{array}$$
The probability distribution is now known as Benford's Law and plays a major role in identifying fraudulent data on tax returns and accounting books. For example, the following distribution represents the first digits in 200 allegedly fraudulent checks written to a bogus company by an employee attempting to embezzle funds from his employer.
$$\begin{array}{lrrrrrrrrr}\text { First digit } & 1 & 2 & 3 & 4 & 5 & 6 & 7 & 8 & 9 \\\hline \text { Frequency } & 36 & 32 & 28 & 26 & 23 & 17 & 15 & 16 & 7 \\\hline\end{array}$$
(a) Because these data are meant to prove that someone is guilty of fraud, what would be an appropriate level of significance when performing a goodness-of-fit test?
(b) Using the level of significance chosen in part (a), test whether the first digits in the allegedly fraudulent checks obey Benford's Law.
(c) Based on the results of part (b), do you think that the emplovee is guilty of embezzlement?

Harsh Gadhiya
Harsh Gadhiya
Numerade Educator
16:37

Problem 14

Refer to Problem 13. The following distribution lists the first digit of the surface area (in square miles) of 335 rivers. Is there evidence at the $\alpha=0.05$ level of significance to support the belief that the distribution follows Benford's Law?
$$\begin{array}{lrrrrrrrrr}\text { First digit } & 1 & 2 & 3 & 4 & 5 & 6 & 7 & 8 & 9 \\\hline \text { Frequency } & 104 & 55 & 36 & 38 & 24 & 29 & 18 & 14 & 17 \\\hline\end{array}$$

Harsh Gadhiya
Harsh Gadhiya
Numerade Educator
17:13

Problem 15

The National Highway Traffic Safety Administration publishes reports about motorcycle fatalities and helmet use. The distribution shows the proportion of fatalities by location of injury for motorcycle accidents.
$$\begin{array}{lccccc}\begin{array}{l}\text { Location } \\\text { of injury }\end{array} & \begin{array}{c}\text { Multiple } \\\text { Locations }\end{array} & \text { Head } & \text { Neck } & \begin{array}{c}\text { Abdomen/ } \\\text { Thorax }\end{array} & \begin{array}{c}\text { Lumbar/Spine } \\0.03\end{array} \\\hline \text { Proportion } & 0.57 & 0.31 & 0.03 & 0.06 & 0.03\end{array}$$
The following data show the location of injury and fatalities for 2068 riders not wearing a helmet.
$$\begin{array}{lccccc}\text { Location } & \text { Multiple } & & & & \text { Abdomen/ } \\\text { of injury } & \text { Locations }& \text { Head } & \text { Neck } & \text { Thorax } & \text {Lumbar/Spine } \\\hline \text { Number } & 1036 & 864 & 38 & 83 & 47\end{array}$$
(a) Does the distribution of fatal injuries for riders not wearing a helmet follow the distribution for all riders? Use the $\alpha=0.05$ level of significance.
(b) Compare the observed and expected counts for each category. What does this information tell you?

Harsh Gadhiya
Harsh Gadhiya
Numerade Educator
17:05

Problem 16

Nationally, the distribution of weapons used in robberies is as shown in the table.
$$\begin{array}{lcccc}\text { Weapon } & \text { Gun } & \text { Knife } & \text { Strong-arm } & \text { Other } \\\hline \text { Proportion } & 0.42 & 0.09 & 0.40 & 0.09 \\\hline\end{array}$$
The data on the following page represent the weapon of choice in 1652 robberies on school property.
$$\begin{array}{llccc}\text { Weapon } & \text { Gun } & \text { Knife } & \text { Strong-arm } & \text { Other } \\\hline \text { Proportion } & 329 & 122 & 857 & 344\end{array}$$
(a) Does the distribution of weapon choice in robberies in schools follow the national distribution? Use the $\alpha=0.05$ level of significance.
(b) Compare the observed and expected counts for each category. What does this information tell you?

Harsh Gadhiya
Harsh Gadhiya
Numerade Educator
06:30

Problem 17

Does the location of your seat in a classroom play a role in attendance or grade? To answer this question, Professors Katherine Perkins and Carl Wieman randomly assigned 400 students* in a general education physics course to one of four groups. Source: Perkins, Katherine K. and Wieman, Carl $\mathbf{E},$ "The Surprising Impact of Seat Location on Student Performance" The Physics Teacher, Vol. 43, Jan. 2005. The 100 students in group 1 sat 0 to 4 meters from the front of the class, the 100 students in group 2 sat 4 to 6.5 meters from the front, the 100 students in group 3 sat 6.5 to 9 meters from the front, and the 100 students in group 4 sat 9 to 12 meters from the front.
(a) For the first half of the semester, the attendance for the whole class averaged $83 \% .$ So, if there is no effect due to seat location, we would expect $83 \%$ of students in each group to attend. The data show the attendance history for each group. How many students in each group attended, on average? Is there a significant difference among the groups in attendance patterns? Use the $\alpha=0.05$ level of significance.
$$\begin{array}{lllll}\text { Group } & 1 & 2 & 3 & 4 \\\hline \text { Attendance } & 0.84 & 0.84 & 0.84 & 0.81\end{array}$$
(b) For the second half of the semester, the groups were rotated so that group 1 students moved to the back of class and group 4 students moved to the front. The same switch took place between groups 2 and $3 .$ The attendance for the second half of the semester averaged $80 \% .$ The data show the attendance records for the original groups (group 1 is now in back group 2 is 6.5 to 9 meters from the front, and so on). How many students in each group attended, on average? Is there a significant difference in attendance patterns? Use the $\alpha=0.05$ level of significance. Do you find anything curious about these data?
$$\begin{array}{lllll}\text { Group } & 1 & 2 & 3 & 4 \\\hline \text { Attendance } & 0.84 & 0.81 & 0.78 & 0.76\end{array}$$
(c) At the end of the semester, the proportion of students in the top $20 \%$ of the class was determined. Of the students in group $1,25 \%$ were in the top $20 \% ;$ of the students in group $2,20 \%$ were in the top $20 \% ;$ of the students in group $3,15 \%$ were in the top $20 \% ;$ of the students in group 4 $19 \%$ were in the top $20 \% .$ How many students would we expect to be in the top $20 \%$ of the class if seat location plays no role in grades? Is there a significant difference in the number of students in the top $20 \%$ of the class by group?
(d) In earlier sections, we discussed results that were statistically significant, but did not have any practical significance. Discuss the practical significance of these results. In other words, given the choice, would you prefer sitting in the front or back?

Saeeda Aman
Saeeda Aman
Numerade Educator
02:31

Problem 18

On January $1,2004,$ it became mandatory for all police departments in Illinois to record data pertaining to race from every traffic stop. The village of Mundelein, Illinois, has been collecting data since $2000 .$ Rather than using census data to determine the racial distribution of the village, they thought it better to use data based on who is using the roads in Mundelein. So they collected data on at-fault drivers involved in car accidents in the village (the implicit assumption here is that race is independent of fault in a crash) and obtained the following distribution of race for users of roads in Mundelein.
$$\begin{array}{lllll} &{\text { African }} \\\text { Race } & \text { White } & \text { American } & \text { Hispanic } & \text { Asian } \\\hline \text { Proportion } & 0.719 & 0.028 & 0.207 & 0.046 \\\hline\end{array}$$
The following data represent the races of all 9868 drivers who were stopped for a moving violation in the village of Mundelein in a recent year.
$$\begin{array}{lllll} & {\text { African }} \\\text { Race } & \text { White } & \text { American } & \text { Hispanic } & \text { Asian } \\\hline \text { Proportion } & 7079 & 273 & 2025 & 491 \\\hline\end{array}$$
(a) Does the distribution of race in traffic stops reflect the distribution of drivers in Mundelein? In other words, is there any evidence of racial profiling in Mundelein? Use the $\alpha=0.05$ level of significance.
(b) Compare observed and expected counts for each category. What does this information tell you?

Sheryl Ezze
Sheryl Ezze
Numerade Educator
03:18

Problem 19

In $2008,$ Malcolm Gladwell released his book Outliers. In the book, Gladwell claims that more hockey players are born in January through March than in October through December. The following data show the number of players selected in the 2010 National Hockey League draft according to their birth month. Is there evidence to suggest that hockey players' birthdates are not uniformly distributed throughout the year? Use the $\alpha=0.05$ level of significance.
$$\begin{array}{lc}\text { Birth Month } & \text { Frequency } \\\hline \text { January-March } & 63 \\\hline \text { April-June } & 56 \\\hline \text { July-September } & 28 \\\hline \text { October-December } & 34 \\\hline\end{array}$$

Saeeda Aman
Saeeda Aman
Numerade Educator
02:35

Problem 20

A researcher wanted to determine whether bicycle deaths were uniformly distributed over the days of the week. She randomly selected 200 deaths that involved a bicycle, recorded the day of the week on which the death occurred, and obtained the following results (the data are based on information obtained from the Insurance Institute for Highway Safety).
$$\begin{array}{lc|lc}\begin{array}{l}\text { Day of } \\\text { the Week }\end{array} & \text { Frequency } & \begin{array}{l}\text { Day of } \\\text { the Week }\end{array} & \text { Frequency } \\\hline \text { Sunday } & 16 & \text { Thursday } & 34 \\\hline \text { Monday } & 35 & \text { Friday } & 41 \\\hline \text { Tuesday } & 16 & \text { Saturday } & 30 \\\hline \text { Wednesday } & 28 & &\end{array}$$
Is there reason to believe that bicycle fatalities occur with equal frequency with respect to day of the week at the $\alpha=0.05$ level of significance?

Saeeda Aman
Saeeda Aman
Numerade Educator
03:18

Problem 21

A researcher wanted to determine whether pedestrian deaths were uniformly distributed over the days of the week. She randomly selected 300 pedestrian deaths, recorded the day of the week on which the death occurred, and obtained the following results (the data are based on information obtained from the Insurance Institute for Highway Safety).
$$\begin{array}{lc|lc}\begin{array}{l}\text { Day of } \\\text { the Week }\end{array} & \text { Frequency } & \begin{array}{l}\text { Day of } \\\text { the Week }\end{array} & \text { Frequency } \\\hline \text { Sunday } & 39 & \text { Thursday } & 41 \\\hline \text { Monday } & 40 & \text { Friday } & 49 \\\hline \text { Tuesday } & 30 & \text { Saturday } & 61 \\\hline \text { Wednesday } & 40 & &\end{array}$$
Test the belief that the day of the week on which a fatality happens involving a pedestrian occurs with equal frequency at the $\alpha=0.05$ level of significance.

Saeeda Aman
Saeeda Aman
Numerade Educator
03:20

Problem 22

A player in a craps game suspects that one of the dice being used in the game is loaded. A loaded die is one in which all the possibilities $(1,2,3,4,5, \text { and } 6)$ are not equally likely. The player throws the die 400 times, records the outcome after each throw, and obtains the following results:
$$\begin{array}{cc|cc}\text { Outcome } & \text { Frequency } & \text { Outcome } & \text { Frequency } \\\hline 1 & 62 & 4 & 62 \\\hline 2 & 76 & 5 & 57 \\\hline 3 & 76 & 6 & 67\end{array}$$
She obtains a sample of 25 home-schooled children within her district that yields the following data:
$$\begin{array}{cc}
\text { Grade } & \text { Frequency } \\
\hline \mathrm{K} & 6 \\\hline 1-3 & 9 \\\hline 4-5 & 3 \\\hline 6-8 & 4 \\\hline 9-12 & 3\end{array}$$
(a) Because of the low cell counts, combine cells into three categories $\mathrm{K}-3,4-8,$ and $9-12$
(b) Is the grade distribution of home-schooled children different in her district from the national grade distribution at the $\alpha=0.05$ level of significance?

Sheryl Ezze
Sheryl Ezze
Numerade Educator
04:17

Problem 23

A player in a craps game suspects that one of the dice being used in the game is loaded. A loaded die is one in which all the possibilities $(1,2,3,4,5, \text { and } 6)$ are not equally likely. The player throws the die 400 times, records the outcome after each throw, and obtains the following results:
$$\begin{array}{cc|cc}\text { Outcome } & \text { Frequency } & \text { Outcome } & \text { Frequency } \\\hline 1 & 62 & 4 & 62 \\\hline 2 & 76 & 5 & 57 \\\hline 3 & 76 & 6 & 67\end{array}$$
She obtains a sample of 25 home-schooled children within her district that yields the following data:
$$\begin{array}{cc}
\text { Grade } & \text { Frequency } \\
\hline \mathrm{K} & 6 \\\hline 1-3 & 9 \\\hline 4-5 & 3 \\\hline 6-8 & 4 \\\hline 9-12 & 3\end{array}$$
(a) Because of the low cell counts, combine cells into three categories $\mathrm{K}-3,4-8,$ and $9-12$
(b) Is the grade distribution of home-schooled children different in her district from the national grade distribution at the $\alpha=0.05$ level of significance?

Saeeda Aman
Saeeda Aman
Numerade Educator
02:44

Problem 24

The Fibonacci sequence is a famous sequence of numbers whose elements commonly occur in nature. The terms in the Fibonacci sequence are $1,1,2,3,5,8,13,21, \ldots .$ The ratio of consecutive terms approaches the golden ratio, $\Phi=\frac{1+\sqrt{5}}{2} .$ The distribution of the first digit of the first 85 terms in the Fibonacci sequence, is as shown in the following table:
$$\begin{array}{lccccccccc}\text { Digit } & 1 & 2 & 3 & 4 & 5 & 6 & 7 & 8 & 9 \\\hline \text { Frequency } & 25 & 16 & 11 & 7 & 7 & 5 & 4 & 6 & 4\end{array}$$
Is there evidence to support the belief that the first digit of the Fibonacci numbers follows the Benford distribution (shown in Problem 13 ) at the $\alpha=0.05$ level of significance? Note: Although it is technically necessary to consolidate cells, we will not.

Sheryl Ezze
Sheryl Ezze
Numerade Educator
02:07

Problem 25

In Example 2 we found that there was not enough evidence to support the conclusion that the distribution of the population in the United States was shifting. The following table shows the results of a survey of $n=5000$ households.
$$\begin{array}{lc}\text { Region } & \text { Frequency } \\\hline \text { Northeast } & 900 \\\hline \text { Midwest } & 1089 \\hline \text { South } & 1846 \\\hline \text { West } & 1165 \\\hline\end{array}$$
(a) Refer to the distribution of households in 2000 from Example 2 . Determine the expected number of households in each region under the assumption the population distribution has not changed since 2000 .
(b) Determine the proportion of households in each region based on the results of the survey. Compare these proportions to those in Table 2 from Example $2 .$ What do you notice?
(c) Does the sample evidence above suggest that the distribution of residents in the United States has changed since $2000 ?$
(d) Discuss the role sample size can play in determining whether the statement in the null hypothesis is rejected.

Sheryl Ezze
Sheryl Ezze
Numerade Educator
02:14

Problem 26

Statistical software and graphing calculators with advanced statistical features use random-number generators to create random numbers conforming to a specified distribution.
(a) Use a random-number generator to create a list of 500 trials of a binomial experiment with $n=5$ and $p=0.2$
(b) What proportion of the numbers generated should be 0? 1? 2?3? 4? 5?
(c) Test whether the random-number generator is generating random outcomes of a binomial experiment with $n=5$ and $p=0.2$ by performing a chi-square goodness-of-fit test at the $\alpha=0.01$ level of significance.

Carson Merrill
Carson Merrill
Numerade Educator
02:42

Problem 27

According to the U.S. Census Bureau, $7.1 \%$ of all babies born are of low birth weight $(<5 \mathrm{lb}, 8 \mathrm{oz})$ An obstetrician wanted to know whether mothers between the ages of 35 and 39 years give birth to a higher percentage of lowbirth-weight babies. She randomly selected 240 births for which the mother was 35 to 39 years old and found 22 low-birth-weight babies.
(a) If the proportion of low-birth-weight babies for mothers in this age group is 0.071, compute the expected number of lowbirth-weight births to 35- to 39-year-old mothers. What is the expected number of births to mothers 35 to 39 years old that are not low birth weight?
(b) Answer the obstetrician's question at the $\alpha=0.05$ level of significance using the chi-square goodness-of-fit test.
(c) Answer the question by using the approach presented in Section 10.2

Sheryl Ezze
Sheryl Ezze
Numerade Educator
02:37

Problem 28

In $2000,25.8 \%$ of Americans 15 years of age or older lived alone, according to the Census Bureau. A sociologist, who believes that this percentage is greater today, conducts a random sample of 400 Americans 15 years of age or older and finds that 164 are living alone.
(a) If the proportion of Americans aged 15 years or older living alone is $0.258,$ compute the following expected numbers: Americans 15 years of age or older who live alone; Americans
15 years of age or older who do not live alone.
(b) Test the sociologist's belief at the $\alpha=0.05$ level of significance using the goodness-of-fit test.
(c) Test the belief by using the approach presented in Section 10.2

Sheryl Ezze
Sheryl Ezze
Numerade Educator
03:21

Problem 29

In Thomas Pynchon's book Gravity Rainbow, the characters discuss whether the Poisson probabilistic model can be used to describe the locations that Germany's feared V-2 rocket would land in. They divided London into $0.25-\mathrm{km}^{2}$ regions. They then counted the number of rockets that landed in each region, with the following results:
$$\begin{array}{lcccccccc}\text { Number of rocket hits } & 0 & 1 & 2 & 3 & 4 & 5 & 6 & 7 \\\hline \begin{array}{l}\text { Observed number } \\\text { of regions }\end{array} & 229 & 211 & 93 & 35 & 7 & 0 & 0 & 1 \\\hline\end{array}$$
(a) Estimate the mean number of rocket hits in a region by computing $\mu=\sum x P(x) .$ Round your answer to four decimal places.
(b) Explain why the requirements for conducting a goodness-offit test are not satisfied.
(c) After consolidating the table, we obtain the following distribution for rocket hits. Using the Poisson probability model, $P(x)=\frac{\mu^{x}}{x !} e^{-\mu},$ where $\mu$ is the mean from part (a), we can obtain the probability distribution for the number of rocket hits. Find the probability of hits in a region. Then find the probability of 1 hit, 2 hits, 3 hits, and 4 or more hits.
$$\begin{array}{lccccc}\text { Number of } & & & & & \\
\text { rocket hits } & 0 & 1 & 2 & 3 & 4 \text { or more } \\\hline \text { Observed number } & & & & & \\\text { of regions } & 229 & 211 & 93 & 35 & 8\end{array}$$
(d) A total of $n=576$ rockets was fired. Determine the expected number of regions hit by computing "expected number of regions" $=n p,$ where $p$ is the probability of observing that particular number of hits in the region.
(e) Conduct a goodness-of-fit test for the distribution using the $\alpha=0.05$ level of significance. Do the rockets appear to be modeled by a Poisson random variable?

Saeeda Aman
Saeeda Aman
Numerade Educator
04:47

Problem 30

On February 2,1894 Frank Raphael Weldon wrote a letter to Francis Galton that included the results of 26,306 rolls of 12 dice. Weldon recorded the results such that a roll of a 5 or 6 resulted in a success, while a roll of $1,2,3,$ or 4 was a failure. The number of successes in each roll of the 12 dice are recorded in the table.
$$\begin{array}{|c|c|cc|}\hline \begin{array}{c}
\text { Number of } \\\text { Successes }\end{array} & \text { Frequency } & \begin{array}{c}\text { Number of } \\\text { Successes }\end{array} & \text { Frequency } \\\hline 0 & 185 & 7 & 1331 \\\hline 1 & 1149 & 8 & 403 \\\hline 2 & 3265 & 9 & 105 \\\hline 3 & 5475 & 10 & 14 \\\hline 4 & 6114 & 11 & 4 \\\hline 5 & 5194 & 12 & 0 \\\hline 6 & 3067 & & \\\hline\end{array}$$
(a) What is the probability of rolling a 5 or 6 when throwing a six-sided fair die?
(b) Treating the probability determined in part (a) as the probability of success, compute the theoretical probability of $0,1,2, \ldots, 12$ successes in throwing 12 dice.
(c) Use the probabilities found in part (b) to determine the expected frequency in observing $0,1,2, \ldots, 12$ successes after throwing the 12 dice 26,306 times.
(d) Conduct a goodness-of-fit test to determine if the number of successes follows a binomial probability distribution. Note: Combine 11 and 12 into a single bin.

Shu Naito
Shu Naito
Numerade Educator
View

Problem 31

Why is goodness of fit a good choice for the title of the procedures used in this section?

Rashmi Sinha
Rashmi Sinha
Numerade Educator
View

Problem 32

Explain why chi-square goodness-of-fit tests are always right tailed.

Rashmi Sinha
Rashmi Sinha
Numerade Educator
00:46

Problem 33

If the expected count of a category is less than $1,$ what can be done to the categories so that a goodness-of-fit test can still be performed?

Prashant Bana
Prashant Bana
Numerade Educator