• Home
  • Textbooks
  • Statistics
  • Multiple Regression and Model Building

Statistics

James T. McClave, Terry T. Sincich

Chapter 12

Multiple Regression and Model Building - all with Video Answers

Educators


Chapter Questions

View

Problem 1

Write a first-order model relating $E(y)$ to
a. two quantitative independent variables
b. four quantitative independent variables
c. five quantitative independent variables

Shu Naito
Shu Naito
Numerade Educator
01:43

Problem 2

List the four assumptions about the random error $\varepsilon$ required for a multiple-regression analysis.

Brandon Cleary
Brandon Cleary
Numerade Educator
01:50

Problem 3

Outline the six steps in a multiple-regression analysis.

Lucas Finney
Lucas Finney
Numerade Educator
01:43

Problem 4

What are the caveats to conducting $t$ -tests on all of the individual $\beta$ parameters in a multiple-regression model?

Lucas Finney
Lucas Finney
Numerade Educator
02:25

Problem 5

How should you test the overall adequacy of a multipleregression model?

Neel Faucher
Neel Faucher
Numerade Educator
02:39

Problem 6

MINITAB was used to fit the model $y=\beta_{0}+$ $\beta_{1} x_{1}+\beta_{2} x_{2}+\varepsilon$ to $n=20$ data points, and the printout on p. 697 was obtained.
a. What are the sample estimates of $\beta_{0}, \beta_{1},$ and $\beta_{2}$ ?
b. What is the least squares prediction equation?
c. Find $\mathrm{SSE}, \mathrm{MSE},$ and $s$. Interpret the standard deviation in the context of the problem.
d. Test $H_{0}: \beta_{1}=0$ against $H_{\mathrm{a}}: \beta_{1} \neq 0 .$ Use $\alpha=.05$.
e. Use a $95 \%$ confidence interval to estimate $\beta_{2}$.
f. Find $R^{2}$ and $R_{a}^{2}$ and interpret these values.
g. Use the two formulas given in this section to calculate the test statistic for the null hypothesis $H_{0}: \beta_{1}=\beta_{2}=0$.

Compare your results with the test statistic shown on the printout.
h. Find the observed significance level of the test you conducted in part $\mathrm{g}$. Interpret the value.

Dominador Tan
Dominador Tan
Numerade Educator
02:35

Problem 7

Suppose you fit the model
$$
y=\beta_{0}+\beta_{1} x_{1}+\beta_{2} x_{2}+\beta_{3} x_{3}+\varepsilon
$$
to $n=30$ data points and obtain the following result:
$$
\hat{y}=2.2-2.8 x_{1}+1.9 x_{2}+.97 x_{3}
$$
The estimated standard errors of $\hat{\beta}_{2}$ and $\hat{\beta}_{3}$ are 1.06 and .27 respectively.
a. Test the null hypothesis $H_{0}: \beta_{2}=0$ against the alternative hypothesis $H_{\mathrm{a}}: \beta_{2} \neq 0 .$ Use $\alpha=.05$.
b. Test the null hypothesis $H_{0}: \beta_{3}=0$ against the alternative hypothesis $H_{\mathrm{a}}: \beta_{3} \neq 0 .$ Use $\alpha=.05$.
c. The null hypothesis $H_{0}: \beta_{2}=0$ is not rejected. In contrast, the null hypothesis $H_{0}: \beta_{3}=0$ is rejected. Explain how this can happen even though $\hat{\beta}_{2}>\hat{\hat{\beta}}_{3}$ ?

Nick Johnson
Nick Johnson
Numerade Educator
02:57

Problem 8

Suppose you fit the first-order multiple-regression model
$$
y=\beta_{0}+\beta_{1} x_{1}+\beta_{2} x_{2}+\varepsilon
$$
to $n=25$ data points and obtain the prediction equation
$$
\hat{y}=6.4+3.1 x_{1}+.92 x_{2}
$$
The estimated standard deviations of the sampling distributions of $\hat{\beta}_{1}$ and $\hat{\beta}_{2}$ are 2.3 and $.27,$ respectively.
a. Test $H_{0}: \beta_{1}=0$ against $H_{\mathrm{a}}: \beta_{1}>0 .$ Use $\alpha=.05$.
b. Test $H_{0}: \beta_{2}=0$ against $H_{\mathrm{a}}: \beta_{2} \neq 0 .$ Use $\alpha=.05$.
c. Find a $90 \%$ confidence interval for $\beta_{1}$. Interpret the interval.
d. Find a $99 \%$ confidence interval for $\beta_{2}$. Interpret the interval.

James Kiss
James Kiss
Numerade Educator
00:35

Problem 9

How is the number of degrees of freedom available for estimating $\sigma^{2}$ (the variance of $\varepsilon$ ) related to the number of independent variables in a regression model?

Christopher Stanley
Christopher Stanley
Numerade Educator
View

Problem 10

Consider the following first-order model equation in three quantitative independent variables:
$$
E(y)=2+4 x_{1}-2 x_{2}-5 x_{3}
$$
a. Graph the relationship between $y$ and $x_{1}$ for $x_{2}=-2$ and $x_{3}=2$.
b. Repeat part a for $x_{2}=3$ and $x_{3}=3$.
c. How do the graphed lines in parts a and b relate to each other? What is the slope of each line?
d. If a linear model is first order in three independent variables, what type of geometric relationship will you obtain when you graph $E(y)$ as a function of one of the independent variables for various combinations of values of the other independent variables?

Shu Naito
Shu Naito
Numerade Educator
02:35

Problem 11

Suppose you fit the first-order model
$$
y=\beta_{0}+\beta_{1} x_{1}+\beta_{2} x_{2}+\beta_{3} x_{3}+\beta_{4} x_{4}+\beta_{5} x_{5}+\varepsilon
$$
to $n=27$ data points and obtain
$$
\mathrm{SSE}=.33 \quad R^{2}=.92
$$
a. Do the values of $\operatorname{SSE}$ and $R^{2}$ suggest that the model provides a good fit to the data? Explain.
b. Is the model of any use in predicting $y ?$ Test the null hypothesis $H_{0}: \beta_{1}=\beta_{2}=\cdots=\beta_{5}=0$ against the alternative hypothesis: At least one of the parameters $\beta_{1}, \beta_{2}, \cdots, \beta_{5}$ is nonzero. Use $\alpha=.10$.

Nick Johnson
Nick Johnson
Numerade Educator
01:04

Problem 12

If the analysis-of-variance $F$ -test leads to the conclusion that at least one of the model parameters is nonzero, can you conclude that the model is the best predictor for the dependent variable $y ?$ Can you conclude that all of the terms in the model are important in predicting $y ?$ What is the appropriate conclusion?

Prashant Bana
Prashant Bana
Numerade Educator
05:38

Problem 13

In poker, making bad decisions due to negative emotions is known as tilting. A study in the Journal of Gambling Studies (Mar. 2014) investigated the factors that affect the severity of tilting for online poker players. A survey of 214 online poker players produced data on the dependent variable, severity of tilting $(y),$ measured on a 30 -point scale (where higher values indicate a higher severity of tilting). Two independent variables measured were poker experience $\left(x_{1},\right.$ measured on a 30 -point scale) and perceived effect of experience on tilting $\left(x_{2},\right.$ measured on a 28 -point scale). The researchers fit the interaction model, $E(y)=\beta_{0}+\beta_{1} x_{1}+\beta_{2} x_{2}+\beta_{3} x_{1} x_{2} .$ The results are shown below ( $p$ -values in parentheses).

a. Evaluate the overall adequacy of the model. Use $\alpha=.01$ to test statistical significance.
b. The researchers hypothesize that the rate of change of severity of tilting $(y)$ with perceived effect of experience on tilting $\left(x_{2}\right)$ depends on poker experience $\left(x_{1}\right) .$ Do you agree? Test using $\alpha=.01$.

Jameson Kuper
Jameson Kuper
Numerade Educator
01:48

Problem 14

Aluminum scraps that are recycled into alloys are classified into three categories: softdrink cans, pots and pans, and automobile crank chambers. A study of how these three materials affect the metal elements present in aluminum alloys was published in Advances in Applied Physics (Vol. 1, 2013). Data on 126 production runs at an aluminum plant were used to model the percentage $(y)$ of various elements (e.g., silver, boron, iron) that make up the aluminum alloy. Three independent variables were used in the model: $x_{1}=$ proportion of aluminum scraps from cans,
$x_{2}=$ proportion of aluminum scraps from pots/pans, and $x_{3}=$ proportion of aluminum scraps from crank chambers. The first-order model, $E(y)=\beta_{0}+\beta_{1} x_{1}+\beta_{2} x_{2}+\beta_{3} x_{3},$
was fit to the data for several elements. The estimates of the model parameters $(p$ -values in parentheses) for silver and iron are shown in the table above.
a. Is the overall model statistically useful (at $\alpha=.05)$ for predicting the percentage of silver in the alloy? If so, give a practical interpretation of $R^{2}$.
b. Is the overall model statistically useful (at $\alpha=.05$ ) for predicting the percentage of iron in the alloy? If so, give a practical interpretation of $R^{2}$.
c. Based on the parameter estimates, sketch the relationship between percentage of silver $(y)$ and proportion of aluminum scraps from cans $\left(x_{1}\right)$. Conduct a test to determine if this relationship is statistically significant at $\alpha=.05$.
d. Based on the parameter estimates, sketch the relationship between percentage of iron $(y)$ and proportion of aluminum scraps from cans $\left(x_{1}\right)$. Conduct a test to determine if this relationship is statistically significant at $\alpha=.05$.

Hast Aggarwal
Hast Aggarwal
Numerade Educator
07:04

Problem 15

Creativity and Innovation Management (Feb. 2008) published an article on identifying the social network characteristics of lead users of children's computer games. Data were collected for $n=326$ children and the following variables measured: lead-user rating $(y,$ measured on a 5 -point scale $),$ gender $\left(x_{1}=1\right.$ if female, 0 if male), age $\left(x_{2},\right.$ years $),$ degree of centrality $\left(x_{3},\right.$ measured as the number of direct ties to other peers in the network), and betweenness centrality $\left(x_{4},\right.$ measured as the number of shortest paths between peers). A first-order model for $y$ was fit to the data, yielding the following least squares prediction equation:
$$
\hat{y}=3.58+.01 x_{1}-.06 x_{2}-.01 x_{3}+.42 x_{4}
$$
a. Give two properties of the errors of prediction that result from using the method of least squares to obtain the parameter estimates.
b. Give a practical interpretation of the estimate of $\beta_{4}$ in the model.
c. A test of $H_{0}: \beta_{4}=0$ resulted in a $p$ -value of .002. Make the appropriate conclusion at $\alpha=.05$.

Heather Duong
Heather Duong
Numerade Educator
02:22

Problem 16

Refer to the Journal of Adolescence (Apr. 2010) study of adolescents' disclosure of their dating and romantic relationships, Exercise 8.43 (p. 419). Data collected for a sample of 222 high school students were used to determine the level of disclosure of the date's identity to an adolescent's mother (measured on a 5-point scale, where $1=$ "never tell," $2=$ "rarely tell," $3=$ "sometimes tell," $4=$ "almost always tell," and $5=$ "always tell"). Multiple regression was used to model level of disclosure $(y)$ to several independent variables, including gender $\left(x_{1}=1\right.$ if female, 0 if male $),$ age $\left(x_{2},\right.$ years $),$ dating experience $\left(x_{3},\right.$ years $),$ and level of trust in parents $\left(x_{4}, 5\right.$ -point scale $)$
a. Give the equation of a first-order model for $y$ as a function of the four independent variables.
b. The coefficient of determination for the model, part a, was reported as $R^{2}=.24$. Give a practical interpretation of this value.
c. Give the null hypothesis for testing the overall adequacy of the model, part a.
d. The test, part $\mathbf{c},$ resulted in $F=56.60$ with $p$ -value $<.001$. Interpret this result using $\alpha=.05$.
e. The estimate of the beta coefficient for age $\left(x_{2}\right)$ was reported as $-.09 .$ Give a practical interpretation of this value.

Lucas Finney
Lucas Finney
Numerade Educator
05:01

Problem 17

Refer to the Marine Mammal Science (April 2010) study of whales entangled in fishing gear, Exercise 10.32 (p. 552). Data collected for a sample of 207 entanglements in the East Sea of Korea were used to model the length $(y)$ of an entangled whale (in meters). Two independent variables used to predict whale length were water depth of the entanglement $\left(x_{1},\right.$ in meters) and distance of the entanglement from land $\left(x_{2},\right.$ in miles).
a. Give the equation of a first-order model for length $(y)$ as a function of the two independent variables.
b. The researchers theorize that the length of an entangled whale will increase linearly as the water depth increases, for entanglements that are a fixed distance from land. Explain how to use the model, part $\mathbf{a},$ to test this theory.
c. The $p$ -value for testing $H_{0}: \beta_{2}=0$ in the model, part a, was reported as .013. Interpret this result using $\alpha=.05$.

Lucas Finney
Lucas Finney
Numerade Educator
05:22

Problem 18

Can the population of an urban area be estimated without taking a census? Geographical Analysis (Jan. 2007) demonstrated the use of satellite image maps in estimating urban population. A portion of Columbus, Ohio, was partitioned into $n=125$ census block groups, and satellite imagery was obtained. For each census block, the following variables were measured: population density $(y),$ proportion of block with low-density residential areas $\left(x_{1}\right),$ and proportion of block with high-density residential areas $\left(x_{2}\right) .$ A first-order model for $y$ was fitted to the data and produced the following results:
$$
\hat{y}=-.0304+2.006 x_{1}+5.006 x_{2}, R^{2}=.686
$$
a. Give a practical interpretation of each $\beta$ -estimate in the model.
b. Give a practical interpretation of the coefficient of determination, $R^{2}$
c. State $H_{0}$ and $H_{\mathrm{a}}$ for a test of the overall adequacy of the model.
d. Refer to part c. Compute the value of the test statistic.
e. Refer to parts $\mathbf{c}$ and $\mathbf{d}$. Make the appropriate conclusion at $\alpha=.01$.

Jameson Kuper
Jameson Kuper
Numerade Educator
01:55

Problem 19

Consider a multipleregression model for predicting the total number of runs scored by a Major League Baseball (MLB) team during a season. Using data on number of walks $\left(x_{1}\right),$ singles $\left(x_{2}\right),$ doubles $\left(x_{3}\right),$ triples $\left(x_{4}\right),$ home runs $\left(x_{5}\right),$ stolen bases $\left(x_{6}\right),$ times caught stealing $\left(x_{7}\right),$ strike outs $\left(x_{8}\right),$ and ground outs $\left(x_{9}\right)$ for each of the 30 teams during the 2014 MLB season, a 1st-order model for total number of runs scored (y) was fit. The results are shown in the accompanying Minitab printout.

a. Write the least squares prediction equation for $y=$ total number of runs scored by a team during the 2014 season.
b. Give practical interpretations of the beta estimates.
c. Conduct a test of $H_{0}: \beta_{7}=0$ against $H_{\mathrm{a}}: \beta_{7}<0 \mathrm{at}$ $\alpha=.05 .$ Interpret the results.
d. Form a $95 \%$ confidence interval for $\beta_{5} .$ Interpret the results.
e. Predict the number of runs scored in 2014 by your favorite Major League Baseball team. How close is the predicted value to the actual number of runs scored by your team? (Note: You can find data on your favorite team on the Internet at www.majorleaguebaseball.com.)

Dominador Tan
Dominador Tan
Numerade Educator
View

Problem 20

Refer to the IEEE International Conference on Web Intelligence and Intelligent Agent Technology (2010) study on using the volume of chatter on Twitter.com to forecast movie box office revenue, Exercise 11.36 (p. 631 ). Recall that opening weekend box office revenue data (in millions of dollars) were collected for a sample of 24 recent movies. In addition to each movie's tweet rate, i.e., the average number of tweets referring to the movie per hour one week prior to the movie's release, the researchers also computed the ratio of positive to negative tweets (called the $P N$ -ratio ).
a. Give the equation of a first-order model relating revenue $(y)$ to both tweet rate $\left(x_{1}\right)$ and $\mathrm{PN}$ -ratio $\left(x_{2}\right)$.
b. Which $\beta$ in the model, part a, represents the change in revenue $(y)$ for every 1 -tweet increase in the tweet rate $\left(x_{1}\right),$ holding PN-ratio $\left(x_{2}\right)$ constant?
c. Which $\beta$ in the model, part a, represents the change in revenue $(y)$ for every 1 -unit increase in the PN-ratio $\left(x_{2}\right),$ holding tweet rate $\left(x_{1}\right)$ constant?
d. The following coefficients were reported: $R^{2}=.945$ and $R_{a}^{2}=.940 .$ Give a practical interpretation for both $R^{2}$ and $R_{a}^{2}$
e. Conduct a test of the null hypothesis, $H_{0}: \beta_{1}=\beta_{2}=0$. Use $\alpha=.05$.
f. The researchers reported the $p$ -values for testing $H_{0}$ :
$\beta_{1}=0$ and $H_{0}: \beta_{2}=0$ as both less than $.0001 .$ Interpret these results (use $\alpha=.01$ ).

Victor Salazar
Victor Salazar
Numerade Educator
05:53

Problem 21

Children with attentiondeficit/hyperactivity disorder (ADHD) were monitored to evaluate their risk for substance (e.g., alcohol, tobacco, illegal drug) use (Journal of Abnormal Psychology, Aug. 2003). The following data were collected on 142 adolescents diagnosed with ADHD:
$y=$ frequency of marijuana use the past six months $x_{1}=$ severity of inattention ( 5 -point scale) $x_{2}=$ severity of impulsivity-hyperactivity $(5$ -point scale $)$ $\begin{aligned} x_{3}=& \text { level of oppositional-defiant and conduct disorder } \\ &(5 \text { -point scale }) \end{aligned}$
a. Write the equation of a first-order model for $E(y)$.
b. The coefficient of determination for the model is $R^{2}=.08 .$ Interpret this value.
c. The global $F$ -test for the model yielded a $p$ -value less than .01. Interpret this result.
d. The $t$ -test for $H_{0}: \beta_{1}=0$ resulted in a $p$ -value less than .01. Interpret this result.
e. The $t$ -test for $H_{0}: \beta_{2}=0$ resulted in a $p$ -value greater than .05. Interpret this result.
f. The $t$ -test for $H_{0}: \beta_{3}=0$ resulted in a $p$ -value greater than .05. Interpret this result.

Beth Stone
Beth Stone
Numerade Educator
02:51

Problem 22

The relationship between the novelty of a vacation destination and vacationing golfers' demographics was investigated in the $A$ nnals of Tourism Research (Apr. 2002). Data were obtained from a mail survey of 393 golf vacationers to a large U.S. coastal resort. Several measures of novelty level (on a numerical scale) were obtained for each vacationer, including "change from routine," "thrill," "boredom-alleviation," and "surprise." The researcher employed four independent variables in a regression model: $x_{1}=$ number of rounds of golf per year, $x_{2}=$ total number of golf vacations taken, $x_{3}=$ number of years the respondent played golf, and $x_{4}=$ average golf score.
a. Give the hypothesized equation of a first-order model for the novelty measure $y=$ change from routine.
b. A test of $H_{0}: \beta_{3}=0$ versus $H_{\mathrm{a}}: \beta_{3}<0$ yielded a $p$ -value of $.005 .$ Interpret this result if $\alpha=.01$.
c. The estimate of $\beta_{3}$ was found to be negative. On the basis of this result (and the result of part $\mathbf{b}$ ), the researcher concluded that "those who have played golf for more years are less apt to seek change from their normal routine in their golf vacations." Do you agree with this statement? Explain.
d. The regression results for the three other dependent measures of novelty are summarized in the accompanying table. Give the null hypothesis for testing the overall adequacy of each first-order regression model.

e. Give the rejection region for the test mentioned in part d. Use $\alpha=.01$.
f. Use the test statistics reported in the table and the rejection region from part e to conduct the test for each of the dependent measures of novelty.
g. Verify that the $p$ -values in the table support the conclusions you drew in part $\mathbf{f}$
h. Interpret the values of $R^{2}$ reported in the table.

Sheryl Ezze
Sheryl Ezze
Numerade Educator
03:40

Problem 23

Environmental Science \& Technology (Jan. 2005) reported on a study of the reliability of a commercial kit designed to test for arsenic in groundwater. The field kit was used to test a sample of 328 groundwater wells in Bangladesh. In addition to the arsenic level (in micrograms per liter), the latitude (degrees), longitude (degrees), and depth (feet) of each well were measured. The first and last five observations of the data are listed in the following table.

a. Write a first-order model for arsenic level $(y)$ as a function of latitude, longitude, and depth.
b. Use the method of least squares to fit the model to the data.
c. Give practical interpretations of the $\beta$ estimates.
d. Find the standard deviation $s$ of the model, and interpret its value.
e. Find and interpret the values of $R^{2}$ and $R_{a}^{2}$.
f. Conduct a test of overall model utility at $\alpha=.05$.
g. On the basis of the results you obtained in parts $\mathbf{d}-\mathbf{f}$, would you recommend using the model to predict arsenic level $(y) ?$ Explain.

Robin Corrigan
Robin Corrigan
Numerade Educator
01:24

Problem 24

Cardanol, an agricultural by-product of cashew nut shells, is a cheap and abundantly available renewable resource. In Industrial \& Engineering Chemistry Research (May 2013), researchers investigated the use of cardanol as an additive for natural rubber. Cardanol was grafted onto pieces of natural rubber latex and the chemical properties examined. One property of interest is the grafting efficiency (measured as a percentage) of the chemical process. The researchers manipulated several independent variables in the chemical process: $x_{1}=$ initiator concentration (parts per hundred resin), $x_{2}=$ cardanol concentration (parts per hundred resin), $x_{3}=$ reaction temperature (degrees Celsius), and $x_{4}=$ reaction time (hours). Values of these variables, as well as the dependent variable $y=$ grafting efficiency, were recorded for a sample of $n=9$ chemical runs. The data are provided in the accompanying table. A MINITAB analysis of the first-order model, $E\left(y_{1}\right)=\beta_{0}+\beta_{1} x_{1}+\beta_{2} x_{2}+\beta_{3} x_{3}+\beta_{4} x_{4},$ is also shown.
a. Conduct a test of overall model adequacy. Use $\alpha=.10$.
b. Interpret practically the value of $R_{a}^{2}$
c. Interpret practically the value of $s$.
d. Find and interpret a $90 \%$ confidence interval for $\beta_{3}$.
e. Conduct a test of $H_{0}: \beta_{4}=0 .$ What do you conclude?

Tyler Moulton
Tyler Moulton
Numerade Educator
View

Problem 25

How much influence do the media, especially reality television programs, have on one's decision to undergo cosmetic surgery? This was the question of interest to psychologists who published an article in Body Image: An International Journal of Research (Mar. 2010). In the study, 170 college students answered questions about their impression of reality TV shows featuring cosmetic surgery, level of self-esteem, satisfaction with their own body, and desire to have cosmetic surgery to alter their body. The variables analyzed in the study were measured as follows:
DESIRE-scale ranging from 5 to $25,$ where the higher the value, the greater the interest in having cosmetic surgery; GENDER-1 if male, 0 if female; SELFESTM-scale ranging from 4 to 40 , where the higher the value, the greater the level of self-esteem; BODYSAT-scale ranging from 1 to 9 , where the higher the value, the greater the satisfaction with one's own body; and IMPREAL-scale ranging from 1 to 7 , where the higher the value, the more one believes reality television shows featuring cosmetic surgery are realistic. The data for the study (simulated based on statistics reported in the journal article) are saved in the IMAGE file. Selected observations are listed in the table on p. 701 . The psychologists used multiple regression to model desire to have cosmetic surgery $(y)$ as a function of gender $\left(x_{1}\right),$ self-esteem $\left(x_{2}\right),$ body satisfaction $\left(x_{3}\right),$ and impression of reality $\mathrm{TV}\left(x_{4}\right)$
a. Fit the first-order model, $E(y)=\beta_{0}+\beta_{1} x_{1}+\beta_{2} x_{2}+$ $\beta_{3} x_{3}+\beta_{4} x_{4},$ to the data. Give the least squares prediction equation.
b. Interpret the $\beta$ -estimates in the words of the problem.
c. Is the overall model statistically useful for predicting desire to have cosmetic surgery? Test using $\alpha=.01$.
d. Which statistic, $R^{2}$ or $R_{a}^{2}$, is the preferred measure of model fit? Practically interpret the value of this statistic.
e. Conduct a test to determine whether desire to have cosmetic surgery decreases linearly as level of body satisfaction increases. Use $\alpha=.05$.
f. Find a $95 \%$ confidence interval for $\beta_{4}$. Practically interpret the result.

Victor Salazar
Victor Salazar
Numerade Educator
View

Problem 26

$\boldsymbol{R}^{2}$ and model fit. Because the coefficient of determination, $R^{2},$ always increases when a new independent variable is added to a model, it is tempting to include many variables in the model in order to force $R^{2}$ to be near
1. However, doing so reduces the number of degrees of freedom available for estimating $\sigma^{2},$ which adversely affects our ability to make reliable inferences. Suppose you want to use 20 psychological and sociological factors to predict a student's standardized test score. You fit the model
$$
y=\beta_{0}+\beta_{1} x_{1}+\cdots+\beta_{20} x_{20}+\varepsilon
$$
where $y=$ test score and $x_{1}, x_{2}, \ldots, x_{20}$ are the psychological and sociological factors. Only 22 years of data $(n=22)$ are used to fit the model, and you obtain $R^{2}=.95 .$ Test to see whether this impressive-looking $R^{2}$ is large enough for you to infer that the model is useful-that is, that at least one term in the model is important in predicting test scores $\alpha=.01$

Victor Salazar
Victor Salazar
Numerade Educator
03:29

Problem 27

Cooling method for gas turbines. Refer to the Journal of Engineering for Gas Turbines and Power (Jan. 2005$)$ study of a high-pressure inlet fogging method for a gas turbine engine, presented in Exercise 8.46 (p. 420). Recall that the heat rate (kilojoules per kilowatt per hour) was measured for each in a sample of 67 gas turbines augmented with high-pressure inlet fogging. In addition, several other variables were measured, including cycle speed (revolutions per minute), inlet temperature $\left({ }^{\circ} \mathrm{C}\right),$ exhaust gas temperature $\left({ }^{\circ} \mathrm{C}\right),$ cycle pressure ratio, and air mass flow rate (kilograms per second). The first and last five observations of the data are listed in the following table.

a. Write a first-order model for heat rate $(y)$ as a function of speed, inlet temperature, exhaust temperature, cycle pressure ratio, and air mass flow rate.
b. Use the method of least squares to fit the model to the data.
c. Give practical interpretations of the $\beta$ estimates.
d. Find the standard deviation $s$ of the model, and interpret its value.
e. Find $R_{a}^{2}$ and interpret its value.
f. Is the overall model statistically useful in predicting heat rate $(y) ?$ Test, using $\alpha=.01$.

Lucas Finney
Lucas Finney
Numerade Educator
01:55

Problem 28

In industry cooling applications (e.g., cooling of nuclear reactors), a process called subcooled flow boiling is often employed. Subcooled flow boiling is susceptible to small bubbles that occur near the heated surface. The characteristics of these bubbles were investigated in Heat Transfer Engineering (Vol. 34,2013 ). A series of experiments was conducted to measure two important bubble behaviors: bubble diameter (millimeters) and bubble density (liters per meters squared). The mass flux (kilograms per meters squared per second) and heat flux (megawatts per meters squared) were varied for each experiment. The data obtained at a set pressure are listed in the following table.
a. Consider the multiple regression model $E\left(y_{1}\right)=\beta_{0}+$ $\beta_{1} x_{1}+\beta_{2} x_{2},$ where $y_{1}=$ bubble diameter, $x_{1}=$ mass
flux, and $x_{2}=$ heat flux. Use statistical software to fit the model to the data and test the overall adequacy of the model.
b. Consider the multiple regression model $E\left(y_{2}\right)=\beta_{0}+$ $\beta_{1} x_{1}+\beta_{2} x_{2},$ where $y_{2}=$ bubble density, $x_{1}=$ mass flux, and $x_{2}=$ heat flux. Use statistical software to fit the model to the data and test the overall adequacy of the model.
c. Which of the two dependent variables, diameter $\left(y_{1}\right)$ or density $\left(y_{2}\right),$ is better predicted by mass flux $\left(x_{1}\right)$ and heat flux $\left(x_{2}\right) ?$

Dominador Tan
Dominador Tan
Numerade Educator
06:52

Problem 29

Researchers at Montana State University have written a tutorial on an empirical method for analyzing before and after highway crash data (Montana Department of Transportation, Research Report, May 2004 ). The initial step in the methodology is to develop a safety performance function (SPF) - a mathematical model that estimates the probability of occurrence of a crash for a given segment of roadway. Using data on over 100 segments of roadway, the researchers fit the model $E(y)=\beta_{0}+\beta_{1} x_{1}+\beta_{2} x_{2},$ where $y=$ number of crashes per three years, $x_{1}=$ roadway length (miles), and $x_{2}=$ average annual daily traffic (number of vehicles ) $=$ AADT. The results are shown in the following tables.

a. Give the least squares prediction equation for the interstate highway model.
b. Give practical interpretations of the $\beta$ estimates you made in part a.
c. Refer to part a. Find a $95 \%$ confidence interval for $\beta_{1}$ and interpret the result.
d. Refer to part a. Find a $95 \%$ confidence interval for $\beta_{2}$ and interpret the result.
e. Repeat parts a-d for the non-interstate-highway model.

Jameson Kuper
Jameson Kuper
Numerade Educator
02:03

Problem 30

Bordeaux wine sold at auction. The uncertainty of the weather during the growing season, the phenomenon that wine tastes better with age, and the fact that some vineyards produce better wines than others encourage speculation concerning the value of a case of wine produced by a certain vineyard during a certain year (or of a certain vintage). The publishers of a newsletter titled Liquid Assets: The International Guide to Fine Wine used a multiple regression approach to predicting the London auction price of red Bordeaux wine. The natural logarithm of the price $y$ (in dollars) of a case containing a dozen bottles of red wine was modeled as a function of weather during the growing season and age of vintage. Consider the multiple regression results for hypothetical data collected for 30 vintages (years) shown at the bottom of the page.
a. Conduct a $t$ -test for each of the beta parameters in the model. Interpret the results.
b. When the natural logarithm of $y$ is used as a dependent variable, the antilogarithm of a beta coefficient minus 1 (i.e., $e^{\beta}-1$ ) represents the percentage change in $y$ for every one-unit increase in the associated $x$ value.* Use this information to interpret each of the beta estimates.
c. Interpret the values of $R^{2}$ and $s$ for the model. Do you recommend using the model to predict red Bordeaux wine prices? Explain.

Dominador Tan
Dominador Tan
Numerade Educator
01:26

Problem 31

Explain why we use $\hat{y}$ as an estimate of $E(y)$ and to predict $y .$

Kaylee Mcclellan
Kaylee Mcclellan
Numerade Educator
00:53

Problem 32

Which interval will be narrower, a $95 \%$ confidence interval for $E(y)$ or a $95 \%$ prediction interval for $y ?$ (Assume that the values of the $x$ 's are the same for both intervals.)

Himanshu Garg
Himanshu Garg
Numerade Educator
07:36

Problem 33

Refer to the Creativity and Innovation Management (Feb. 2008) study of lead users of children's computer games, Exercise 12.15 (p. 698 ). Recall that the researchers modeled lead-user rating $(y,$ measured on a 5 -point scale) as a function of gender $\left(x_{1}=1\right.$ if female, 0 if male ), age $\left(x_{2},\right.$ years $),$ degree of centrality $\left(x_{3},\right.$ measured as the number of direct ties to other peers in the network), and betweenness centrality $\left(x_{4},\right.$ measured as the number of shortest paths between peers). The least squares prediction equation was $\hat{y}=3.58+.01 x_{1}-.06 x_{2}-.01 x_{3}+.42 x_{4}$
a. Compute the predicted lead-user rating of a 10 -year-old female child with 5 direct ties to other peers in her social network and with 2 shortest paths between peers.
b. Compute an estimate for the mean lead-user rating of all 8-year-old male children with 10 direct ties to other peers and with 4 shortest paths between peers.

Trent Speier
Trent Speier
Numerade Educator
02:11

Problem 34

Refer to the Chance (Fall 2000 ) study of runs scored in Major League Baseball games, Exercise 12.19 (p. 698 ). Multiple regression was used to model total number of runs scored $(y)$ of a team during the season as a function of number of walks $\left(x_{1}\right)$, number of singles $\left(x_{2}\right),$ number of doubles $\left(x_{3}\right),$ number of triples $\left(x_{4}\right),$ number of home runs $\left(x_{5}\right),$ number of stolen bases $\left(x_{6}\right),$ number of times caught stealing $\left(x_{7}\right),$ number of strikeouts $\left(x_{8}\right),$ and total number of outs $\left(x_{9}\right) .$ Using the $\beta$ -estimates given in Exercise $12.19,$ predict the number of runs scored by your favorite Major League Baseball team last year. How close is the predicted value to the actual number of runs scored by your team? [Note: You can find data on your favorite team on the Internet at www.majorleaguebaseball.com.]

Lucas Finney
Lucas Finney
Numerade Educator
03:00

Problem 35

Refer to the Body Image:
An International Journal of Research (Mar. 2010) study of the impact of reality TV shows on one's desire to undergo cosmetic surgery, Exercise 12.25 (p. 700). Recall that psychologists used multiple regression to model desire to have cosmetic surgery $(y)$ as a function of gender $\left(x_{1}\right),$ self-esteem $\left(x_{2}\right),$ body satisfaction $\left(x_{3}\right),$ and impression of reality $\mathrm{TV}$
$\left(x_{4}\right) .$ The SAS printout below shows a confidence interval for $E(y)$ for each of the first five students in the study.
a. Interpret the confidence interval for $E(y)$ for student $1 .$
b. Interpret the confidence interval for $E(y)$ for student $4 .$

Marc Lauzon
Marc Lauzon
Numerade Educator
05:50

Problem 36

The Usability Professionals' Association (UPA) supports people who research, design, and evaluate the user experience of products and services. The UPA conducted a salary survey of its members (UPA Salary Survey, Aug. 18,2009 ). One of the report's authors, Jeff Sauro, investigated how much having a PhD affects salaries in this profession and discussed his analysis on the blog www.measuringusability.com. Sauro fit a first-order multiple regression model for salary $(y,$ in dollars) as a function of years of experience $\left(x_{1}\right), \mathrm{PhD}$ status $\left(x_{2}=1\right.$ if $\mathrm{PhD}, 0$ if not $)$, and manager status $\left(x_{3}=1\right.$ if manager, 0 if not). The following prediction equation was obtained:
$$
\hat{y}=52,484+2,941 x_{1}+16,880 x_{2}+11,108 x_{3}
$$
a. Predict the salary of a UPA member with 10 years of experience who does not have a $\mathrm{PhD}$ but is a manager.
b. Predict the salary of a UPA member with 10 years of experience who does have a $\mathrm{PhD}$ but is not a manager.
c. Why is a $95 \%$ prediction interval preferred over the predicted values given in parts a and $\mathbf{b}$ ?

James Kiss
James Kiss
Numerade Educator
03:29

Problem 37

Refer to the Journal of Engineering for Gas Turbines and Power (Jan. 2005) study of a high-pressure inlet fogging method for a gas turbine engine, presented in Exercise 12.27 (p. 701). Recall that you fitted a first-order model for heat rate $(y)$ as a function of speed $\left(x_{1}\right)$, inlet temperature $\left(x_{2}\right),$ exhaust temperature $\left(x_{3}\right),$ cycle pressure ratio $\left(x_{4}\right),$ and air mass flow rate $\left(x_{5}\right)$. A MINITAB printout with both a $95 \%$ confidence interval for $E(y)$ and a prediction interval for $y,$ for selected values of the $x$ 's, is shown above.
a. Interpret the $95 \%$ prediction interval for $y$ in the words of the problem.
b. Interpret the $95 \%$ confidence interval for $E(y)$ in the words of the problem.
c. Will the confidence interval for $E(y)$ always be narrower than the prediction interval for $y ?$ Explain.

Lucas Finney
Lucas Finney
Numerade Educator
02:25

Problem 38

An article published in Geography (July 1980) used multiple regression to predict annual rainfall levels in California. Data on the average annual precipitation $(y),$ altitude $\left(x_{1}\right),$ latitude $\left(x_{2}\right),$ and distance from the Pacific coast $\left(x_{3}\right)$ for 30 meteorological stations scattered throughout California are saved in the CALRAIN file. (Selected observations are listed in the table below.) Consider the first-order model $y=\beta_{0}+\beta_{1} x_{1}+\beta_{2} x_{2}+\beta_{3} x_{3}+\varepsilon$
a. Fit the model to the data and give the least squares prediction equation.
b. Is there evidence that the model is useful in predicting annual precipitation $y ?$ Test, using $\alpha=.05$.
c. Find a $95 \%$ prediction interval for $y$ for the Giant Forest meteorological station (station 9$)$. Interpret the interval.

Lucas Finney
Lucas Finney
Numerade Educator
04:27

Problem 39

The Engineering Project Organizational Journal (Vol. 3,2013$)$ published the results of an exploratory study designed to gain a better understanding of how the emotional intelligence of individual team members relates directly to the performance of their team during an engineering project. Undergraduate students enrolled in the course Introduction to the Building Industry participated in the study. All students completed an emotional intelligence test and received an interpersonal score, stress management score, and mood score. Students were grouped into $n=23$ teams and assigned a group project. However, each student received an individual project score. These scores were averaged to obtain the dependent variable in the analysis: mean project score $(y)$. Three independent variables were determined for each team: range of interpersonal scores $\left(x_{1}\right),$ range of stress management scores $\left(x_{2}\right),$ and range of mood scores $\left(x_{3}\right) .$ Data (simulated from information provided in the article) are listed in the table.

Mohan Jain
Mohan Jain
Numerade Educator
View

Problem 40

In a production facility, an accurate estimate of hours needed to complete a task is crucial to management in making such decisions as hiring the proper number of workers, quoting an accurate deadline for a client, or performing cost analyses regarding budgets. A manufacturer of boiler drums wants to use regression to predict the number of hours needed to erect the drums in future projects. To accomplish this task, data on 36 boilers were collected. In addition to hours $(y),$ the variables measured were boiler capacity $\left(x_{1}=\mathrm{lb} / \mathrm{hr}\right),$ boiler design pressure $\left(x_{2}=\right.$ pounds per square inch, or psi), boiler type $\left(x_{3}=1\right.$ if industry field erected, 0 if utility field erected), and drum type $\left(x_{4}=1\right.$ if steam, 0 if mud). The data are saved in the BOILERS file. (Selected observations are shown in the table below.)
a. Fit the model $E(y)=\beta_{0}+\beta_{1} x_{1}+\beta_{2} x_{2}+\beta_{3} x_{3}+\beta_{4} x_{4}$
to the data and give the prediction equation.
b. Conduct a test for the global utility of the model. Use $\alpha=.01$
c. Find a $95 \%$ confidence interval for $E(y)$ when $x_{1}=150,000, x_{2}=500, x_{3}=1,$ and $x_{4}=0 .$ Interpret
the result.
d. What type of interval would you use if you want to estimate the average number of hours required to erect all industrial mud boilers with a capacity of $150,000 \mathrm{lb} / \mathrm{hr}$ and a design pressure of 500 psi?

Shu Naito
Shu Naito
Numerade Educator
03:40

Problem 41

Arsenic in groundwater. Refer to the Environmental Science \& Technology (Jan. 2005) study of the reliability of a commercial kit designed to test for arsenic in groundwater, presented in Exercise 12.23 (p. 700). You fit a first-order model for arsenic level $(y)$ as a function of latitude, longitude, and depth. On the basis of the model statistics, the researchers concluded that the arsenic level is highest at a low latitude, high longitude, and low depth. Do you agree? If so, find a $95 \%$ prediction interval for arsenic level for the lowest latitude, highest longitude, and lowest depth that are within the range of the sample data. Interpret the result.

Robin Corrigan
Robin Corrigan
Numerade Educator
01:06

Problem 42

If two variables $x_{1}$ and $x_{2}$ do not interact, how would you describe their effect on the mean response $E(y) ?$

Tyler Moulton
Tyler Moulton
Numerade Educator
01:48

Problem 43

Write an interaction model relating the mean value of $y$, $E(y),$ to
a. two quantitative independent variables
b. three quantitative independent variables [Hint: Include all possible two-way cross-product terms.]

Prashant Bana
Prashant Bana
Numerade Educator
View

Problem 44

Suppose the true relationship between $E(y)$ and the quantitative independent variables $x_{1}$ and $x_{2}$ is given by the model below.
$$
E(y)=-5+x_{1}+3 x_{2}-3 x_{1} x_{2}
$$
to $n=32$ data points and obtain the results below:
$$
\mathrm{SS}_{y y}=500 \quad \mathrm{SSE}=21 \quad \hat{\beta}_{3}=10 \quad s_{\hat{\beta}_{3}}=4
$$
a. Find $R^{2}$ and interpret its value.
b. Is the model adequate for predicting $y ?$ Test at $\alpha=.05$
c. Use a graph to explain the contribution of the $x_{1} x_{2}$ term to the model.
d. Is there evidence that $x_{1}$ and $x_{2}$ interact? Test at $\alpha=.05$

Victor Salazar
Victor Salazar
Numerade Educator
02:39

Problem 46

MINITAB was used to fit the model
$$
y=\beta_{0}+\beta_{1} x_{1}+\beta_{2} x_{2}+\beta_{3} x_{1} x_{2}+\varepsilon
$$
to $n=15$ data points. The resulting printout is shown below.
a. What is the prediction equation for the response surface?
b. Describe the geometric form of the response surface of part a.
c. Plot the prediction equation for the case when $x_{2}=1$. Do this twice more on the same graph for the cases
$$
\text { when } x_{2}=3 \text { and } x_{2}=5
$$
d. Explain what it means to say that $x_{1}$ and $x_{2}$ interact. Explain why the graph you plotted in part c suggests that $x_{1}$ and $x_{2}$ interact.
e. Specify the null and alternative hypotheses you would use to test whether $x_{1}$ and $x_{2}$ interact.
f. Conduct the hypothesis test of part e, using $\alpha=.01$.

Dominador Tan
Dominador Tan
Numerade Educator
03:39

Problem 47

By law, food servers at restaurants are not entitled to minimum wages because they are tipped by customers. Can food servers increase their tips by complimenting the customers they are waiting on? To answer this question, researchers collected data on the customer tipping behavior for a sample of 348 dining parties and reported their findings in the Journal of Applied Social Psychology (Vol. 40,2010 ). Tip size $(y,$ measured as a percentage of the total food bill) was modeled as a function of size of the dining party $\left(x_{1}\right)$ and whether the server complimented the customers' choice of menu items $\left(x_{2}\right)$. One theory states that the effect of size of the dining party on tip size is independent of whether the server compliments the customers' menu choices. A second theory hypothesizes that the effect of size of the dining party on tip size will be greater when the server compliments the customers' menu choices as opposed to when the server refrains from complimenting menu choices.
a. Write a model for $E(y)$ as a function of $x_{1}$ and $x_{2}$ that corresponds to Theory 1 .
b. Write a model for $E(y)$ as a function of $x_{1}$ and $x_{2}$ that corresponds to Theory 2 .
c. The researchers summarized the results of their analysis with the following graph. Based on the graph, which of the two models would you expect to fit the data better? Explain.

Bon Zapata
Bon Zapata
Numerade Educator
04:26

Problem 48

The impact of global warming and subsequent snow melt on the flight season (migration dates) of butterflies inhabiting the High Arctic was investigated in Current Zoology (Apr. 2014). Annual data for a certain species of butterfly were collected for a 14-year period. The day of the onset of the flight season $(y)$ was modeled as a function of the timing of the snowmelt $\left(x_{1},\right.$ the first date when less than $10 \mathrm{~cm}$ of snow was measured) and average July temperature $\left(x_{2},\right.$ in degrees Celsius). The interaction model, $E(y)=\beta_{0}+\beta_{1} x_{1}+\beta_{2} x_{2}+\beta_{3} x_{1} x_{2},$ was
fit to the data with the following results:
$$
\begin{array}{c}
\hat{y}=-144.2+2.05 x_{1}+35.4 x_{2}-.21 x_{1} x_{2}, \quad R^{2}=73 \\
F \text { -test } p \text { -value }=.0032
\end{array}
$$
a. Is the overall model statistically useful for predicting day of the onset of the flight season $(y)$ ? Test using $\alpha=.01$
b. Give a practical interpretation of $R^{2}$.
c. A test of $H_{0}: \beta_{3}=0$ was rejected at $\alpha=.01$. Should the researchers conclude that timing of snowmelt and average July temperature interact to predict the day of the onset of the flight season?
d. Give an estimate of the change in the day of the onset of the flight season $(y)$ for every 1 -day increase in the timing of the snowmelt $\left(x_{1}\right)$ for years with an average July temperature of $x_{2}=30$ degrees Celsius.

James Kiss
James Kiss
Numerade Educator
00:54

Problem 49

Refer to the Marine Mammal Science (Apr. 2010) study of whales entangled in fishing gear, Exercise 12.17 (p. 698 ). Recall that the length (y) of an entangled whale (in meters) was modeled as a function of water depth of the entanglement $\left(x_{1},\right.$ in meters $)$ and distance of the entanglement from land $\left(x_{2},\right.$ in miles).
a. Give the equation of an interaction model for length
(y) as a function of the two independent variables.
b. The researchers theorize that the length of an entangled whale will increase linearly as the water depth increases. In terms of the parameters in the model, part a, write the slope of the line relating length $(y)$ to water depth $\left(x_{1}\right)$ for a distance of $x_{2}=10$ miles from land.
c. Repeat part $\mathbf{b}$ for a distance of $x_{2}=25$ miles from land.

Nick Auwerda
Nick Auwerda
Numerade Educator
03:01

Problem 50

Retail interest is defined by marketers as the level of interest a consumer has in a given retail store. Marketing professors at the University of Tennessee at Chattanooga and the University of Alabama investigated the role of retailer interest in consumers' shopping behavior (Journal of Retailing, Summer 2006 ). Using survey data collected on $n=375$ consumers, the professors developed an interaction model for $y=$ willingness of the consumer to shop at a retailer's store in the future (called "repatronage intentions") as a function of $x_{1}=$ consumer satisfaction and $x_{2}=$ retailer interest. The regression results are shown below.

a. Is the overall model statistically useful in predicting $y ?$ Test, using $\alpha=.05$.
b. Conduct a test for interaction at $\alpha=.05$.
c. Use the $\beta$ -estimates to sketch the estimated relationship between repatronage intentions $(y)$ and satisfaction $\left(x_{1}\right)$ when retailer interest is $x_{2}=1($ a low value $)$.
d. Repeat part $\mathbf{c}$ for the case when retailer interest is $x_{2}=7$ (a high value).
e. Put the two lines you sketched in parts $\mathbf{c}$ and $\mathbf{d}$ on the same graph to illustrate the nature of the interaction.

Kari Hasz
Kari Hasz
Numerade Educator
01:15

Problem 51

A study of how eye movement behavior can distort one's judgment of the location of an object was published in Advances in Cognitive Psychology (Vol. 6,2010 ). The researchers had volunteers fixate their eyes on a cross in the middle of a computer screen. A probe was then spatially extended near the cross and each volunteer was asked to judge the location of the probe. Saccadic (i.e., fast voluntary) eye movement was monitored during each session. The researchers used spatial position of the probe $\left(x_{1},\right.$ measured in degrees $)$ and position of the cross $\left(x_{2},\right.$ degrees) to predict the amplitude $(y)$ of saccadic eye movement. The following model was fit to the data:
$$
E(y)=\beta_{0}+\beta_{1} x_{1}+\beta_{2} x_{2}+\beta_{3} x_{1} x_{2}
$$
a. The model yielded $R^{2}=.994$. Interpret this result.
b. In the words of the problem, what does it mean to say that " $x_{1}$ and $x_{2}$ interact"?
c. The least squares prediction equation was determined as: $\hat{y}=.91+.70 x_{1}-.06 x_{2}-.03 x_{1} x_{2} .$ Illustrate interaction by graphing the relationship between predicted amplitude $(\hat{y})$ and cross position $\left(x_{2}\right)$ for probe positions $x_{1}=3.5$ and $x_{1}=6.5$.

Wendi Zhao
Wendi Zhao
Numerade Educator
02:02

Problem 52

While waiting in a long line for service (e.g., to use an ATM or at the post office), at some point you may decide to leave the line. The Journal of Consumer Research (Nov. 2003) published a study of consumer behavior while waiting in a line. College students (sample size $n=148$ ) were asked to imagine that they were waiting in line at a post office to mail a package and that the estimated waiting time was 10 minutes or less. After a 10-minute wait, students were asked about their level of negative feelings (annoyed, anxious) on a scale of 1 (strongly disagree) to 9 (strongly agree). Before answering, however, the students were informed about how many people were ahead of them and behind them in the line. The researchers used regression to relate negative feelings score $(y)$ to number ahead in line $\left(x_{1}\right)$ and number behind in line $\left(x_{2}\right)$.
a. The researchers fit an interaction model to the data. Write the hypothesized equation of this model.
b. In the words of the problem, explain what it means to say that " $x_{1}$ and $x_{2}$ interact to affect $y$."
c. A $t$ -test for the interaction $\beta$ resulted in a $p$ -value greater than $.25 .$ Interpret this result.
d. From their analysis, the researchers concluded that "the greater the number of people ahead, the higher [is] the negative feeling score" and "the greater the number of people behind, the lower [is] the negative feeling score." Use this information to determine the signs of $\hat{\beta}_{1}$ and $\hat{\beta}_{2}$ in the model.

Sheryl Ezze
Sheryl Ezze
Numerade Educator
09:00

Problem 53

Refer to the IEEE International Conference on Web Intelligence and Intelligent Agent Technology (2010) study on using the volume of chatter on Twitter.com to forecast movie box office revenue, Exercise $12.20(\mathrm{p} .699) .$ The researchers modeled a movie's opening weekend box office revenue $(y)$ as a function of tweet rate $\left(x_{1}\right)$ and ratio of positive to negative tweets $\left(x_{2}\right)$ using a first-order model.
a. Write the equation of an interaction model for $E(y)$ as a function of $x_{1}$ and $x_{2}$.
b. In terms of the $\beta$ 's in the model, part a, what is the change in revenue $(y)$ for every 1 -tweet increase in the tweet rate $\left(x_{1}\right),$ holding PN-ratio $\left(x_{2}\right)$ constant at a value of $2.5 ?$
c. In terms of the $\beta$ 's in the model, part a, what is the change in revenue $(y)$ for every 1 -tweet increase in the tweet rate $\left(x_{1}\right),$ holding PN-ratio $\left(x_{2}\right)$ constant at a value of $5.0 ?$
d. In terms of the $\beta$ 's in the model, part a, what is the change in revenue $(y)$ for every 1 -unit increase in the PN-ratio $\left(x_{2}\right),$ holding tweet rate $\left(x_{1}\right)$ constant at a value of $100 ?$
e. Give the null hypothesis for testing whether tweet rate $\left(x_{1}\right)$ and PN-ratio $\left(x_{2}\right)$ interact to affect revenue $(y)$.

Alex Loukas
Alex Loukas
Numerade Educator
View

Problem 54

Refer to the Body Image:
An International Journal of Research (March 2010) study of the influence of reality TV shows on one's desire to undergo cosmetic surgery, Exercise 12.25 (p. 700 ). Recall that psychologists modeled desire to have cosmetic surgery $(y)$ as a function of gender $\left(x_{1}\right),$ self-esteem $\left(x_{2}\right),$ body satisfaction $\left(x_{3}\right),$ and impression of reality $\mathrm{TV}\left(x_{4}\right) .$ For this exercise, consider only the independent variables gender $\left(x_{1}\right)$ and impression of reality TV $\left(x_{4}\right)$.
a. The research psychologists theorize that the impact of one's impression of reality $\mathrm{TV}$ on level of desire for cosmetic surgery will be greater for females than for males. Does this theory imply that the independent variables $x_{1}$ and $x_{4}$ interact, or that there is no interaction? Explain.
b. Fit the interaction model, $E(y)=\beta_{0}+\beta_{1} x_{1}+\beta_{2} x_{4}+$ $\beta_{3} x_{1} x_{4},$ to the data.
c. Use the results, part $\mathbf{b},$ to carry out a test for interaction. Make your conclusion using $\alpha=.05$.

Rashmi Sinha
Rashmi Sinha
Numerade Educator
01:09

Problem 55

A study was conducted to determine the effects of linguistic delivery style and client credibility on auditors' judgments (Advances in Accounting and Behavioral Research, 2003). Each of 200 auditors performed an analytical review of a client's financial statement. The researchers gave the auditors different information on the client's credibility and the linguistic delivery style of the client's explanation. Each auditor then provided an assessment of the likelihood that the client's explanation accounts for the fluctuation in the financial statement. The three variables of interest-credibility $\left(x_{1}\right),$ linguistic delivery style $\left(x_{2}\right),$ and likelihood $(y)-$ were all measured on a numerical scale. Regression analysis was used to fit the interaction model $y=\beta_{0}+\beta_{1} x_{1}+\beta_{2} x_{2}+\beta_{3} x_{1} x_{2}+\varepsilon .$ The results
are summarized in the table at the bottom of the page.
a. Interpret the phrase "client credibility and linguistic delivery style interact" in the words of the problem.
b. Give the null and alternative hypotheses for testing the overall adequacy of the model.
c. Conduct the test suggested in part $\mathbf{b}$, using the information in the table.
d. Give the null and alternative hypotheses for testing whether client credibility and linguistic delivery style interact.
e. Conduct the test suggested in part $\mathbf{d}$, using the information in the table.
f. The researchers estimated the slope of the likelihoodlinguistic delivery style line at a low level of client credibility $\left(x_{1}=22\right) .$ Obtain this estimate and interpret it in the words of the problem.
g. The researchers also estimated the slope of the likelihood-linguistic delivery style line at a high level of client credibility $\left(x_{1}=46\right) .$ Obtain this estimate and interpret it in the words of the problem.

Dominador Tan
Dominador Tan
Numerade Educator
03:00

Problem 56

Psychologists define $i m$ plicit self-esteem as unconscious evaluations of one's worth or value. In contrast, explicit self-esteem refers to the extent to which a person consciously considers oneself as valuable and worthy. An article published in Journal of Articles in Support of the Null Hypothesis (Mar. 2006$)$ investigated whether implicit self-esteem is really unconscious. A sample of 257 college undergraduate students completed a questionnaire designed to measure implicit self-esteem and explicit self-esteem. Thus, an implicit selfesteem score $\left(x_{1}\right)$ and explicit self-esteem score $\left(x_{2}\right)$ were obtained for each. (Note: Higher scores indicate higher levels of self-esteem.) Also, a second questionnaire was administered in order to obtain each subject's estimate of his/her level of implicit self-esteem. The score obtained from this questionnaire was called an estimated implicit self-esteem score $\left(x_{3}\right) .$ Finally, the researchers computed two measures of accuracy in estimating implicit selfesteem: $y_{1}=\left(x_{3}-x_{1}\right)$ and $y_{2}=\left|x_{3}-x_{1}\right| .$
a. The researchers fit the interaction model $E\left(y_{1}\right)=\beta_{0}+\beta_{1} x_{1}+\beta_{2} x_{2}+\beta_{3} x_{1} x_{2} .$ The $t$ -test of the interaction term, $\beta_{3}$, was "nonsignificant," with a $p$ -value $>.10 .$ However, both $t$ -tests of $\beta_{1}$ and $\beta_{2}$ were statistically significant $(p$ -value $<.001)$. Interpret these results practically.
b. The researchers also fit the interaction model $E\left(y_{2}\right)=\beta_{0}+\beta_{1} x_{1}+\beta_{2} x_{2}+\beta_{3} x_{1} x_{2} .$ The $t$ -test on the interaction term, $\beta_{3}$, was "significant," with a $p$ -value $<.001$. Interpret this result practically.

Marc Lauzon
Marc Lauzon
Numerade Educator
03:29

Problem 57

Refer to the Journal of Engineering for Gas Turbines and Power (Jan. 2005) study of a high-pressure inlet fogging method for a gas turbine engine, presented in Exercise 12.27 (p. 701). Recall that you fit a first-order model for heat rate $(y)$ as a function of speed $\left(x_{1}\right),$ inlet temperature $\left(x_{2}\right),$ exhaust temperature $\left(x_{3}\right),$ cycle pressure ratio $\left(x_{4}\right),$ and air mass flow rate $\left(x_{5}\right)$ to the data.
a. Researchers hypothesize that the linear relationship between heat rate $(y)$ and temperature (both inlet and exhaust) depends on air mass flow rate. Write a model for heat rate that incorporates the researchers' theories.
b. Use statistical software to fit the interaction model you wrote in part a. Give the least squares prediction equation.
c. Conduct a test (at $\alpha=.05$ ) to determine whether inlet temperature and air mass flow rate interact to affect heat rate.
d. Conduct a test (at $\alpha=.05$ ) to determine whether exhaust temperature and air mass flow rate interact to affect heat rate.
e. Interpret practically the results of the tests you conducted in parts $\mathbf{c}$ and $\mathbf{d}$.

Lucas Finney
Lucas Finney
Numerade Educator
03:40

Problem 58

Refer to the Environmental Science \& Technology (Jan. 2005) study of the reliability of a commercial kit to test for arsenic in groundwater, presented in Exercise 12.23 (p. 700 ). Recall that you fit a first-order model for arsenic level $(y)$ as a function of latitude $\left(x_{1}\right),$ longitude $\left(x_{2}\right),$ and depth $\left(x_{3}\right)$ to the data.
a. Write a model for arsenic level $(y)$ that includes firstorder terms for latitude, longitude, and depth, as well as terms for interaction between latitude and depth and interaction between longitude and depth.
b. Use statistical software to fit the interaction model you wrote in part a. Give the least squares prediction equation.
c. Conduct a test (at $\alpha=.05$ ) to determine whether latitude and depth interact to affect arsenic level.
d. Conduct a test (at $\alpha=.05$ ) to determine whether longitude and depth interact to affect arsenic level.
e. Interpret practically the results of the tests you conducted in parts $\mathbf{c}$ and $\mathbf{d}$.

Robin Corrigan
Robin Corrigan
Numerade Educator
01:55

Problem 59

Refer to the Heat Transfer Engineering (Vol. 34, 2013) study of bubble behavior in subcooled flow boiling, Exercise 12.28 (p. 701 ). Recall that bubble density (liters per meters squared) was modeled as a function of mass flux (kilograms per meters squared per second) and heat flux (megawatts per meters squared) using data saved in the BUBBLE2 file.
a. Write an interaction model for bubble density $(y)$ as a function of $x_{1}=$ mass flux and $x_{2}=$ heat flux.
b. Fit the interaction model, part a, to the data using statistical software. Give the least squares prediction equation.
c. Evaluate overall model adequacy by conducting a global $F$ -test $($ at $\alpha=.05)$ and interpreting the model statistics, $R_{a}^{2}$ and $2 s .$
d. Conduct a test (at $\alpha=.05$ ) to determine whether mass flux and heat flux interact.
e. How much do you expect bubble density to decrease for every $1 \mathrm{~kg} / \mathrm{m}^{2}$ -sec increase in mass flux when heat flux is set at .5 megawatts $/ \mathrm{m}^{2} ?$

Dominador Tan
Dominador Tan
Numerade Educator
03:49

Problem 60

In the model $E(y)=\beta_{0}+\beta_{1} x+\beta_{2} x^{2},$
a. Which $\beta$ represents the $y$ -intercept?
b. Which $\beta$ represents the shift?
c. Which $\beta$ represents the rate of curvature?

Nick Johnson
Nick Johnson
Numerade Educator
View

Problem 61

Write a second-order model relating the mean of $y$, $E(y),$ to
a. one quantitative independent variable
b. two quantitative independent variables
c. three quantitative independent variables [Hint: Include all possible two-way cross-product terms and squared terms.]

Shu Naito
Shu Naito
Numerade Educator
02:35

Problem 62

Suppose you fit the quadratic model
$$
E(y)=\beta_{0}+\beta_{1} x+\beta_{2} x^{2}
$$
to a set of $n=20$ data points and found
$$
R^{2}=.85, \mathrm{SS}_{y y}=25.59, \text { and } \mathrm{SSE}=3.82 .
$$
a. Is there sufficient evidence to indicate that the model contributes information for predicting $y$ ? Test using $\alpha=.05$. Write the hypotheses for the test.
b. What null and alternative hypotheses would you test to determine whether upward curvature exists?
c. What null and alternative hypotheses would you test to determine whether downward curvature exists?

Nick Johnson
Nick Johnson
Numerade Educator
02:35

Problem 63

Suppose you fit the second-order model
$$
y=\beta_{0}+\beta_{1} x+\beta_{2} x^{2}+\varepsilon
$$
to $n=30$ data points. Your estimate of $\beta_{2}$ is $\hat{\beta}_{2}=.52,$ and the estimated standard error of the estimate is .18 .
a. Test $H_{0}: \beta_{2}=0$ against $H_{\mathrm{a}}: \beta_{2} \neq 0 .$ Use $\alpha=.05$.
b. Suppose you want to determine only whether the quadratic curve opens upward; that is, as $x$ increases, the slope of the curve would increase. Give the test statistic and the rejection region for the test for $\alpha=.05 .$ Do the data support the theory that the slope of the curve increases as $x$ increases? Explain.

Nick Johnson
Nick Johnson
Numerade Educator
02:39

Problem 64

MINITAB was used to fit the complete second-order model
$$
E(y)=\beta_{0}+\beta_{1} x_{1}+\beta_{2} x_{2}+\beta_{3} x_{1} x_{2}+\beta_{4} x_{1}^{2}+\beta_{5} x_{2}^{2}
$$
to $n=39$ data points. (See the printout on p. 722)
a. Is there sufficient evidence to indicate that at least one of the parameters $\beta_{1}, \beta_{2}, \beta_{3}, \beta_{4},$ and $\beta_{5}$ is nonzero? Test, using $\alpha=.05$.
b. Test $H_{0}: \beta_{4}=0$ against $H_{\mathrm{a}}: \beta_{4} \neq 0 .$ Use $\alpha=.01$.
c. Test $H_{0}: \beta_{5}=0$ against $H_{\mathrm{a}}: \beta_{5} \neq 0 .$ Use $\alpha=.01$.
d. Use graphs to explain the consequences of the tests in parts $\mathbf{b}$ and $\mathbf{c}$.

Dominador Tan
Dominador Tan
Numerade Educator
03:51

Problem 65

Consider the following quadratic models:
(1) $y=1+4 x+2 x^{2}$
$$
\begin{array}{l}
\text { (2) } y=1-4 x+2 x^{2} \\
\text { (3) } y=1+2 x^{2} \\
\text { (4) } y=1-2 x^{2} \\
\text { (5) } y=1+5 x^{2}
\end{array}
$$
a. Graph each of these quadratic models, side by side, on the same sheet of graph paper.
b. What effect does the first-order term $(4 x)$ have on the graph of the curve?
c. What effect does the second-order term $\left(2 x^{2}\right)$ have on the graph of the curve?

James Kiss
James Kiss
Numerade Educator
02:21

Problem 66

The role of maintenance in energy savings in commercial refrigeration was the topic of an article in the Journal of Quality in Maintenance Engineering (Vol. 18,2012 ). The authors provided the following illustration of data relating the efficiency (relative performance) of a refrigeration system to the fraction of total charges for cooling the system required for optimal performance. Based on the data shown in the graph, hypothesize an appropriate model for relative performance $(y)$ as a function of fraction of charge $(x)$. What is the hypothesized sign (positive or negative) of the $\beta_{2}$ parameter in the model?

Katelyn Chen
Katelyn Chen
Numerade Educator
05:50

Problem 67

When attempting to predict job performance using personality traits, researchers typically assume that the relationship is linear. A study published in the Journal of Applied Psychology (Jan. 2011) investigated a curvilinear relationship between job task performance and a specific personality trait: conscientiousness. Using data collected for 602 employees of a large public organization, task performance was measured on a 30 -point scale (where higher scores indicate better performance) and conscientiousness was measured on a scale of -3 to +3 (where higher scores indicate a higher level of conscientiousness).
a. The coefficient of correlation relating task performance score to conscientiousness score was reported as $r=.18 .$ Explain why the researchers should not use this statistic to investigate the curvilinear relationship between task performance and conscientiousness.
b. Give the equation of a curvilinear (quadratic) model relating task performance score $(y)$ to conscientiousness score $(x)$
c. The researchers theorize that task performance will increase as level of conscientiousness increases but at a decreasing rate. Draw a sketch of this relationship.
d. If the theory in part $\mathbf{c}$ is supported, what is the expected sign of $\beta_{2}$ in the model, part $\mathbf{b}$ ?
e. The researchers reported $\hat{\beta}_{2}=-.32$ with an associated $p$ -value of less than $.05 .$ Use this information to test the researchers' theory at $\alpha=.05$.

Heather Duong
Heather Duong
Numerade Educator
03:40

Problem 68

The eating patterns of families of overweight preschool children were the subject of an article published in the Journal of Education and Human Development (Vol. 3, 2009). A sample of 10 overweight children living in a rural area of the United States was selected. A portion of the research focused on the body mass index of each child and his/her parent. (Body mass indexor BMI-is determined by dividing weight by height squared.) These data are provided in the table on p. $723 .$ The researchers were interested in determining whether parent BMI could be used as a predictor of child BMI for overweight children.
a. For this study, identify the dependent and independent variables.
b. Construct a scatterplot for the data. What trend do you observe?
c. A quadratic model was fit to the data, with the results shown in the SPSS printout below. Give the least squares prediction equation.
d. Is the overall model statistically useful? Test using $\alpha=.05 .$
e. Is there evidence of upward curvature in the relationship between child BMI and parent BMI? Test using $\alpha=.05 .$
f. Do you agree with the statement, "for obese children, child BMI increases at an increasing rate as parent BMI increases"?

Lucas Finney
Lucas Finney
Numerade Educator
02:52

Problem 69

Refer to the Chance (Winter 2009 ) study of fourth-down decisions by coaches in the National Football League (NFL), Exercise 11.87 (p. 652). Recall that statisticians at California State University, Northridge, fit a straight-line model for predicting the number of points scored $(y)$ by a team that has a first down with a given number of yards $(x)$ from the opposing goal line. A second model fit to data collected on five NFL teams from a recent season was the quadratic regression model, $E(y)=\beta_{0}+\beta_{1} x+\beta_{2} x^{2}$. The regression yielded the following results: $\hat{y}=6.13+.141 x-.0009 x^{2}, R^{2}=.226$.
a. If possible, give a practical interpretation of each of the $\beta$ -estimates in the model.
b. Give a practical interpretation of the coefficient of determination, $R^{2}$.
c. In Exercise $11.87,$ the coefficient of correlation for the straight-line model was reported as $R^{2}=.18$. Does this
statistic alone indicate that the quadratic model is a better fit than the straight-line model? Explain.
d. What test of hypothesis would you conduct to determine if the quadratic model is a better fit than the straight-line model?

Lucas Finney
Lucas Finney
Numerade Educator
02:15

Problem 70

Management professors at Columbia University examined the relationship between assertiveness and leadership (Journal of Personality and Social Psychology, Feb. 2007). The sample comprised 388 people enrolled in a full-time master's in business administration (MBA) program. On the basis of answers to a questionnaire, the researchers measured two variables for each subject: assertiveness score $(x)$ and leadership ability score $(y)$. A quadratic regression model was fit to the data, with the following results:
a. Conduct a test of overall model utility. Use $\alpha=.05$.
b. The researchers hypothesized that leadership ability will increase at a decreasing rate with assertiveness. Set up the null and alternative hypotheses to test this theory.
c. Use the reported results to conduct the test you set up in part $\mathbf{b}$. Give your conclusion (at $\alpha=.05$ ) in the words of the problem.

Jameson Kuper
Jameson Kuper
Numerade Educator
01:23

Problem 71

Underinflated or overinflated tires can increase tire wear. A new tire was tested for wear at different pressures, with the results shown in the following table.
$$
\begin{array}{cc}
\hline \begin{array}{c}
\text { Pressure } x \text { (pounds } \\
\text { per square inch) }
\end{array} & \begin{array}{c}
\text { Mileage } y \\
\text { (thousands) }
\end{array} \\
\hline 30 & 29 \\
31 & 32 \\
32 & 36 \\
33 & 38 \\
34 & 37 \\
35 & 33 \\
36 & 26 \\
\hline
\end{array}
$$
a. Plot the data on a scatterplot.
b. If you were given only the information for $x=30,31$, 32, and $33,$ what kind of model would you suggest? For $x=33,34,35,$ and 36 ? For all the data?

Carson Merrill
Carson Merrill
Numerade Educator
00:01

Problem 72

Do chief executive officers (CEOs) and their top managers always agree on the goals of the company? Goal importance congruence between CEOs and vice presidents (VPs) was studied in the Academy of Management Journal (Feb. 2008). The researchers used regression to model a VP's attitude toward the goal of improving efficiency $(y)$ as a function of the two quantitative independent variables, level of CEO leadership $\left(x_{1}\right)$ and level of congruence between the $\mathrm{CEO}$ and the VP $\left(x_{2}\right)$. A complete second-order model in $x_{1}$ and $x_{2}$ was fit to data collected for $n=517$ top management team members at U.S. credit unions.
a. Write the complete second-order model for $E(y)$
b. The coefficient of determination for the model, part a, was reported as $R^{2}=.14$. Interpret this value.
c. The estimate of the $\beta$ -value for the $\left(x_{2}\right)^{2}$ term in the model was found to be negative. Interpret this result, practically.
d. A $t$ -test on the $\beta$ -value for the interaction term in the model, $x_{1} x_{2},$ resulted in a $p$ -value of .02. Practically interpret this result, using $\alpha=.05$.

Oluwadamilola Ameobi
Oluwadamilola Ameobi
Numerade Educator
04:59

Problem 73

Refer to the Geographical Analysis (Jan. 2007 ) study that demonstrated the use of satellite image maps for estimating urban population, presented in Exercise 12.18 (p. 698 ). A first-order model for census block population density $(y)$ was fit as a function of the proportion of a block with low-density residential areas $\left(x_{1}\right)$ and the proportion of a block with high-density residential areas $\left(x_{2}\right) .$ Now consider a complete second-order model for $y$.
a. Write the equation of the model.
b. Identify the terms in the model that allow for curvilinear relationships.

Sneha Ravi
Sneha Ravi
Numerade Educator
01:04

Problem 74

Refer to the IHS Journal of Hydraulic Engineering (Sept. 2012) study of the repair and replacement of water pipes, Exercise 11.28 (p. 628 ). Recall that a team of civil engineers used regression analysis to model $y=$ the ratio of repair to replacement cost of commercial pipe as a function of $x=$ the diameter (in millimeters) of the pipe. Data for a sample of 13 different pipe sizes are reproduced in the accompanying table. In Exercise 11.28 , you fit a straightline model to the data. Now consider the quadratic model $E(y)=\beta_{0}+\beta_{1} x+\beta_{2} x^{2} .$ A MINITAB printout of the analysis follows in the next column.
$$
\begin{array}{cc}
\hline \text { Diameter } & \text { Ratio } \\
\hline 80 & 6.58 \\
100 & 6.97 \\
125 & 7.39 \\
150 & 7.61 \\
200 & 7.78 \\
250 & 7.92 \\
300 & 8.20 \\
350 & 8.42 \\
400 & 8.60 \\
450 & 8.97 \\
500 & 9.31 \\
600 & 9.47 \\
700 & 9.72 \\
\hline
\end{array}
$$
a. Give the least squares prediction equation relating ratio of repair to replacement cost $(y)$ to pipe diameter $(x)$.
b. Conduct a global $F$ -test for the model using $\alpha=.01$. What do you conclude about overall model adequacy?
c. Evaluate the adjusted coefficient of determination, $R_{a}^{2}$, for the model.
d. Give the null and alternative hypotheses for testing if the rate of increase of ratio $(y)$ with diameter $(x)$ is slower for larger pipe sizes.
e. Carry out the test, part $\mathbf{d}$, using $\alpha=.01$.
f. Locate on the printout a $95 \%$ prediction interval for the ratio of repair to replacement cost for a pipe with a diameter of 250 millimeters. Interpret the result.

Nick Johnson
Nick Johnson
Numerade Educator
03:51

Problem 75

A standard method for studying toxic substances and their effects on humans is to observe the responses of rodents exposed to various doses of the substance over time. In the Journal of Agricultural, Biological, and Environmental Statistics (June 2005), researchers used least squares regression to estimate the "change-point" dosage, defined as the largest dose level that has no adverse effects. Data were obtained from a dose-response study of rats exposed to the toxic substance aconiazide. A sample of 50 rats was evenly divided into five dosage groups: $0,100,200,500,$ and 750 milligrams per kilogram of body weight. The dependent variable $y$ measured was the weight change (in grams) after a 2-week exposure. The researchers fit the quadratic model $E(y)=\beta_{0}+\beta_{1} x+\beta_{2} x^{2},$ where $x=$ dosage level, with the following results: $\hat{y}=10.25+.0053 x-.0000266 x^{2}$.
a. Construct a rough sketch of the least squares prediction equation. Describe the nature of the curvature in the estimated model.
b. Estimate the weight change $(y)$ for a rat given a dosage of $500 \mathrm{mg} / \mathrm{kg}$ of aconiazide.
c. Estimate the weight change $(y)$ for a rat given a dosage of $0 \mathrm{mg} / \mathrm{kg}$ of aconiazide. (This dosage is called the "control" dosage level.)
d. Of the five groups in the study, find the largest dosage level $x$ that yields an estimated weight change that is closest to, but below, the estimated weight change for the control group. This value is the change-point dosage.

Sheryl Ezze
Sheryl Ezze
Numerade Educator
02:35

Problem 76

Refer to the International Journal of Retail and Distribution Management $($ Vol. 39,2011$)$ study of shopping on Black Friday (the day after Thanksgiving), Exercise 7.22 (p. 353). Recall that researchers conducted interviews with a sample of 38 women shopping on Black Friday to gauge their shopping habits. Two of the variables measured for each shopper were age $(x)$ and number of years shopping on Black Friday $(y)$. Data on these two variables for the 38 shoppers are listed in the accompanying table.
a. Fit the quadratic model, $E(y)=\beta_{0}+\beta_{1} x+\beta_{2} x_{2},$ to the data using statistical software. Give the prediction equation.
b. Conduct a test of the overall adequacy of the model. Use $\alpha=.01$.
c. Conduct a test to determine if the relationship between age $(x)$ and number of years shopping on Black Friday $(y)$ is best represented by a linear or quadratic function. Use $\alpha=.01$.

Christopher Stanley
Christopher Stanley
Numerade Educator
11:41

Problem 77

How satisfied are people who have recently joined a new religious movement? To answer this question, German researchers collected data for a sample of 58 believers who had recently joined a new religious group (Applied Psychology:
An International Review, Apr. 2010). The dependent variable of interest was satisfaction level $(y),$ measured quantitatively on an 11 -point scale (where $0=$ totally dissatisfied and $10=$ totally satisfied). Two independent variables were used to predict satisfaction level: Needs $\left(x_{1}\right)-$ a measure of the level of needs one requires in a religion, and Supplies $\left(x_{2}\right)-$ a measure of the level of supplies provided by the religion. In theory, if the level of needs matches the level of supplies, one will be highly satisfied with the religion.
a. The researchers fitted a complete second-order model for $E(y)$ as a function of $x_{1}$ and $x_{2}$. Write the equation of this model.
b. The regression results are reported in the table below. Interpret the value of $R^{2}$
c. Use the $R^{2}$ statistic to conduct a test of overall model adequacy. Test using $\alpha=.10$.
d. Conduct a test to determine whether needs $\left(x_{1}\right)$ is curvilinearly related to satisfaction $(y)$. Test using $\alpha=.10$.
e. Conduct a test to determine whether supplies $\left(x_{2}\right)$ is curvilinearly related to satisfaction $(y) .$ Test using $\alpha=.10 .$

Heather Duong
Heather Duong
Numerade Educator
06:28

Problem 78

Researchers at National Semiconductor experimented with tin-lead solder bumps used to manufacture silicon wafer integrated circuit chips (International Wafer Level Packaging Conference, Nov. $3-4,2005$ ). The failure times of the microchips (in hours) were determined at different solder temperatures (degrees Celsius). The data for one experiment are given in the following table. The researchers want to predict failure time $(y)$ based on solder temperature $(x)$.
a. Construct a scatterplot for the data. What type of relationship, linear or curvilinear, appears to exist between failure time and solder temperature?
b. Fit the model, $E(y)=\beta_{0}+\beta_{1} x+\beta_{2} x^{2},$ to the data. Give the least squares prediction equation.
c. Conduct a test to determine if there is upward curvature in the relationship between failure time and solder temperature. (Use $\alpha=.05 .)$

PG
Patrick Garavaglia
Numerade Educator
03:36

Problem 79

In the Journal of Experimental Psychology: Learning, Memory, and Cognition (July 2005), University of Basel (Switzerland) psychologists tested the ability of people to judge the risk of an infectious disease. The researchers asked German college students to estimate the number of people who are infected with a certain disease in a typical year. The median estimates, as well as the actual incidence of the disease for each in a sample of 24 infections, are listed in the table. Consider the quadratic model $E(y)=\beta_{0}+\beta_{1} x+\beta_{2} x^{2},$ where $y=$ actual incidence rate and $x=$ estimated rate.

a. Fit the quadratic model to the data, and then conduct a test to determine whether the actual incidence is curvilinearly related to the estimated incidence. (Use $\alpha=.05 .)$
b. Construct a scatterplot of the data. Locate the data point for botulism on the graph. What do you observe?
c. Repeat part a, but omit the data point for botulism from the analysis. Has the fit of the model improved? Explain.

Jon Southam
Jon Southam
Numerade Educator
02:38

Problem 80

The optomotor responses of tree frogs were studied in the Journal of Experimental Zoology (Sept. 1993). Microspectrophotometry was used to measure the threshold quantal flux (the light intensity at which the optomotor response was first observed) of tree frogs tested at different spectral wavelengths. The data revealed the relationship between the logarithm of quantal flux $(y)$ and wavelength $(x),$ shown in the following graph.

Noah Musser
Noah Musser
Numerade Educator
06:02

Problem 81

Write a regression model relating the mean value of $y$ to a qualitative independent variable that can assume two levels. Interpret all the terms in the model.

Sneha Ravi
Sneha Ravi
Numerade Educator
01:44

Problem 82

Write a regression model relating $E(y)$ to a qualitative independent variable that can assume three levels. Interpret all the terms in the model.

Shu Naito
Shu Naito
Numerade Educator
View

Problem 83

The model $E(y)=\beta_{0}+\beta_{1} x_{1}+\beta_{2} x_{2}+\beta_{3} x_{3},$ where
$$
\begin{array}{l}
x_{1}=\left\{\begin{array}{ll}
1 & \text { if level } 2 \\
0 & \text { if } \mathrm{not}
\end{array}\right. \\
x_{2}=\left\{\begin{array}{ll}
1 & \text { if level } 3 \\
0 & \text { if } \mathrm{not}
\end{array}\right. \\
x_{3}=\left\{\begin{array}{ll}
1 & \text { if level } 4 \\
0 & \text { if } \mathrm{not}
\end{array}\right.
\end{array}
$$
was used to relate $E(y)$ to a single qualitative variable with four levels. This model was fit to $n=45$ data points, and the result was and the following result was obtained:
$$
\hat{y}=16.3-7 x_{1}+15 x_{2}+3 x_{3}
$$
a. Use the least squares prediction equation to find the estimate of $E(y)$ for each level of the qualitative independent variable.
b. Specify the null and alternative hypotheses you would use to test whether $E(y)$ is the same for all four levels of the independent variable.

Victor Salazar
Victor Salazar
Numerade Educator
02:39

Problem 84

MINITAB was used to fit the model
$$
y=\beta_{0}+\beta_{1} x_{1}+\beta_{2} x_{2}+\varepsilon
$$
where
$$
\begin{array}{l}
x_{1}=\left\{\begin{array}{ll}
1 & \text { if level } 2 \\
0 & \text { if not }
\end{array}\right. \\
x_{2}=\left\{\begin{array}{ll}
1 & \text { if level } 3 \\
0 & \text { if } \text { not }
\end{array}\right.
\end{array}
$$
to $n=15$ data points. The results are shown in the accompanying MINITAB printout.

a. Report the least squares prediction equation.
b. Interpret the values of $\beta_{1}$ and $\beta_{2}$.
c. Interpret the following hypotheses in terms of $\mu_{1}, \mu_{2}$, and $\mu_{3}$ :
$H_{0}: \beta_{1}=\beta_{2}=0$
$H_{\mathrm{a}}:$ At least one of the parameters $\beta_{1}$ and $\beta_{2}$ differs from 0
d. Conduct the hypothesis test of part $\mathbf{c}$.

Dominador Tan
Dominador Tan
Numerade Educator
02:25

Problem 85

As part of a study on how children affect parents' depression in old age, researchers fit a model for total number of children in a family (Social Science \& Medicine, Jan. 2014). One of the key independent variables used in the model is whether the first birth was a single birth or a multiple (twin, triplet) birth.
a. Create a dummy variable for type of first birth.
b. Write a model for total number of children $(y)$ as a function of type of first birth.
c. Give a practical interpretation of each of the $\beta$ -coefficients in the model.

AG
Ankit Gupta
Numerade Educator
View

Problem 86

The production of quality wine is strongly influenced by the natural endowments of the grape-growing region called the "terroir." The Economic Journal (May 2008) published an empirical study of the factors that yield a quality Bordeaux wine. A quantitative measure of wine quality
(y) was modeled as a function of several qualitative independent variables, including grape-picking method (manual or automated), soil type (clay, gravel, or sand), and slope orientation (east, south, west, southeast, or southwest).
a. Create the appropriate dummy variables for each of the qualitative independent variables.
b. Write a model for wine quality $(y)$ as a function of grape-picking method. Interpret the $\beta$ 's in the model.
c. Write a model for wine quality $(y)$ as a function of soil type. Interpret the $\beta$ 's in the model.
d. Write a model for wine quality (y) as a function of slope orientation. Interpret the $\beta$ 's in the model.

Victor Salazar
Victor Salazar
Numerade Educator
04:34

Problem 87

Refer to the Marine Mammal Science (Apr. 2010) study of whales entangled in fishing gear, Exercise 12.17 (p. 698 ). These entanglements involved one of three types of fishing gear: set nets, pots, and gill nets. Consequently, the researchers used gear type as a predictor of the body length $(y,$ in meters $)$ of the entangled whale. Consider the regression model $E(y)=\beta_{0}+\beta_{1} x_{1}+\beta_{2} x_{2},$ where $x_{1}=\{1$ if set net, 0 if not $\}$ and $x_{2}=\{1$ if pots, 0 if not $\}$ [Note: Gill nets is the "base" level of gear type.]
a. The researchers want to know the mean body length of whales entangled in gill nets. Give an expression for this value in terms of the $\beta$ 's in the model.
b. Practically interpret the value of $\beta_{1}$ in the model.
c. In terms of the $\beta$ 's in the model, how would you test to determine if the mean body lengths of entangled whales differ for the three types of fishing gear?

Jameson Kuper
Jameson Kuper
Numerade Educator
04:21

Problem 88

During fundraising, does the physical appearance of the solicitor affect the level of capital raised? An economist at the University of NevadaReno designed an experiment to answer this question and published the results in Economic Letters (Vol. 100,2008 ). Each in a sample of 955 households was contacted by a female solicitor and asked to contribute to the Center for Natural Hazards Mitigation Research. The level of contribution (in dollars) was recorded as well as the hair color of the solicitor (blonde Caucasian, brunette Caucasian, or minority female).
a. Consider a model for the mean level of contribution, $E(y),$ that allows for different means depending on the hair color of the solicitor. Create the appropriate number of dummy variables for hair color. (Use minority female as the base level.)
b. Write the equation of the model, part a, incorporating the dummy variables.
c. In terms of the $\beta$ 's in the model, what is the mean level of contribution for households contacted by a blonde Caucasian solicitor?
d. In terms of the $\beta$ 's in the model, what is the difference between the mean level of contribution for households contacted by a blonde solicitor and those contacted by a minority female?

Lucas Finney
Lucas Finney
Numerade Educator
00:01

Problem 89

University of Colorado sociologists investigated the impact of race on the value of professional football players' "rookie" cards (Electronic Journal of Sociology, 2007 ). The sample consisted of 148 rookie cards of National Football League (NFL) players who were inducted into the Football Hall of Fame. The price of a card (in dollars) was modeled as a function of several qualitative independent variables: race of player (black or white), availability of the card (high or low), and position of the player (quarterback, running back, wide receiver, tight end, defensive lineman, linebacker, defensive back, or offensive lineman).
a. Create the appropriate dummy variables for each of the qualitative independent variables.
b. Write a model for price $(y)$ as a function of race. Interpret the $\beta$ 's in the model.
c. Write a model for price $(y)$ as a function of the availability of the card. Interpret the $\beta$ 's in the model.
d. Write a model for price $(y)$ as a function of the player's position. Interpret the $\beta$ 's in the model.

Pritesh Ranjan
Pritesh Ranjan
Numerade Educator
02:47

Problem 90

Researchers at the University of Aberdeen (Scotland) developed a statistical model for estimating the chemical composition of water (Journal of Agricultural, Biological, and Environmental Statistics, Mar. 2005 ). For one application, the nitrate concentration $y$ (milligrams per liter) in a water sample collected after a heavy rainfall was modeled as a function of water source (groundwater, subsurface flow, or overground flow).
a. Write a model for $E(y)$ as a function of the qualitative independent variable.
b. Give an interpretation of each of the $\beta$ parameters in the model you wrote in part a.

James Kiss
James Kiss
Numerade Educator
01:46

Problem 91

In gene therapy, it is important to know the location of a gene for a disease on the genome (genetic map). Although many genes yield a specific trait (e.g., disease or not), others cannot be categorized, since they are quantitative in nature (e.g., extent of disease). Researchers at the University of North Carolina at Wilmington developed statistical models that link quantitative genetic traits to locations on the genome (Chance, Summer 2006). The extent of a certain disease is determined by the absence (A) or presence (B) of a gene marker at each of two locations, L1 and L2, on the genome. For example, AA represents absence of the marker at both locations, while AB represents absence at location $\mathrm{L} 1,$ but presence at location $\mathrm{L} 2$
a. How many different gene marker combinations are possible at the two locations?
b. Using dummy variables, write a model for extent of the disease, $y,$ as a function of gene marker combination.
c. Interpret the $\beta$ -values in the model you wrote in part $\mathbf{b}$.
d. Give the null hypothesis for testing whether the overall model from part $\mathbf{b}$ is statistically useful for predicting extent of the disease, $y$.

James Kiss
James Kiss
Numerade Educator
View

Problem 92

Refer to the Animal Science Journal (May 2014) study on the use of ascorbic acid (AA) to reduce the stress in goats during transportation, Exercise 10.7 (p. 537 ). Recall that 24 healthy goats were randomly divided into four groups $(\mathrm{A}, \mathrm{B}, \mathrm{C},$ and $\mathrm{D})$ of six animals each. Goats in group A were administered a dosage of AA 30 minutes prior to transportation; goats in group $\mathrm{B}$ were administered a dosage of $\mathrm{AA} 30$ minutes following transportation; group C goats were not given any AA prior to or following transportation; and goats in group D were not given any AA and were not transported. Weight was measured before and after transportation and the weight loss (in kilograms) determined for each goat.
a. Write a model for mean weight loss, $E(y),$ as a function of AA dosage group $(\mathrm{A}, \mathrm{B}, \mathrm{C},$ or $\mathrm{D})$. Use group $\mathrm{D}$ as the base level.
b. Interpret the $\beta$ 's in the model, part a.
c. Recall that the researchers discovered that mean weight loss is reduced in goats administered AA compared with goats not given any AA. Based on this result, determine the sign (positive or negative) of as many of the $\beta$ 's in the model, part a, as possible.

Shu Naito
Shu Naito
Numerade Educator
02:44

Problem 93

A team of physicians, psychiatrists, and psychologists investigated whether depressed patients exhibit more or fewer personality disorder symptoms than nondepressed patients in the American Journal of Psychiatry (May 2010). A study group of over 400 psychiatric patients was monitored over a six-year period. At the start of the study, each was diagnosed as having (1) major depression only, (2) personality disorder only, or (3) both major depression and personality disorder. Of interest to the researchers was the number of personality disorder criteria met at the end of the study. Consider a regression model for the number of personality disorders $(y)$.
a. Write a model for $E(y)$ as a function of the qualitative variable, patient diagnosis group.
b. If there are no differences among the mean number of personality disorders for the three patient groups, what are the values of the $\beta$ 's in the model, part a?
c. How could you test to determine if the mean number of personality disorders for the major depression-only patients is less than the corresponding mean for the patients with both major depression and personality disorder?

Joshua Argo
Joshua Argo
Numerade Educator
03:34

Problem 94

Refer to the Journal of Applied Psychology (June 2002) study of recall of television commercials, presented in Exercise 10.37 Participants were assigned to watch one of three types of TV programs, with nine commercials embedded in each show. Group $\mathrm{V}$ watched a TV program with a violent-content code rating, Group S viewed a show with a sex-content code rating, and Group $\mathrm{N}$ watched a neutral TV program with neither a V nor an S rating. The dependent variable measured for each participant was the score $(y)$ on his/her recall of the brand names mentioned in the commercial messages, with scores ranging from 0 (no brands recalled) to 9 (all brands recalled).
a. Write a model for $E(y)$ as a function of viewer group.
b. Fit the model you wrote in part a to the data. Give the least squares prediction equation.
c. Conduct a test of overall model utility at $\alpha=.01$. Interpret the results. Show that the results agree with the analysis performed in Exercise 10.37 .
d. The sample mean recall scores for the three groups were $\bar{y}_{\mathrm{V}}=2.08, \bar{y}_{\mathrm{S}}=1.71,$ and $\bar{y}_{\mathrm{N}}=3.17 .$ Show how to find
these sample means by using only the $\beta$ -estimates obtained in part $\mathbf{b}$.

Anne Glasgow
Anne Glasgow
Numerade Educator
02:42

Problem 95

An article published in the Duke Journal of Gender Law \& Policy (Summer 2003) examined the impact of expert testimony on the outcome of homicide trials that involve battered woman syndrome. On the basis of data collected on individual juror votes from past trials, the article reported that "when expert testimony was present, women jurors were more likely than men to change a verdict from not guilty to guilty after deliberations." Assume that when no expert testimony was present, male jurors were more likely than women to change a verdict from not guilty to guilty after deliberations. These results were obtained from a multiple-regression model for likelihood of changing a verdict from not guilty to guilty after deliberations, $y,$ as a function of juror gender (male or female) and expert testimony (yes or no). Give the model for $E(y)$ that hypothesizes the relationships reported in the article. Illustrate the model with a sketch.

Haggai Liu
Haggai Liu
Numerade Educator
03:25

Problem 96

Do college professors who provide their students with assistance on homework help improve student grades? This was the research question of interest in the Journal of Accounting Education (Vol. 25,2007 ). A sample of 175 accounting students took a pretest on a topic not covered in class, then each was given a homework problem to solve on the same topic. The students were assigned to one of three homework assistance groups. Some students received the completed solution, some were given check figures at various steps of the solution, and some received no help at all. After finishing the homework, the students were all given a posttest on the subject. The dependent variable of interest was the knowledge gain (or test score improvement). These data are saved in the ACCHW file.
a. Propose a model for the knowledge gain $(y)$ as a function of the qualitative variable, homework assistance group.
b. In terms of the $\beta$ 's in the model, give an expression for the difference between the mean knowledge gains of students in the "completed solution" and "no help" groups.
c. Fit the model to the data and give the least squares prediction equation.
d. Conduct the global $F$ -test for model utility using $\alpha=.05$. Interpret the results practically.

Nick Johnson
Nick Johnson
Numerade Educator
01:45

Problem 97

Refer to the Evolutionary Ecology Research (July 2003) study of the patterns of extinction in the New Zealand bird population, presented in Exercise 2.24 (p. 70). Recall that the NZBIRDS file contains qualitative data on flight capability (volant or flightless), habitat (aquatic, ground terrestrial, or aerial terrestrial), nesting site (ground, cavity within ground, tree, or cavity above ground), nest density (high or low), diet (fish, vertebrates, vegetables, or invertebrates), and extinct status (extinct, absent from island, present), and quantitative data on body mass (grams) and egg length (millimeters) for 132 bird species at the time of the Maori colonization of New Zealand.
a. Write a model for mean body mass as a function of flight capability.
b. Write a model for mean body mass as a function of diet.
c. Write a model for mean egg length as a function of nesting site.
d. Fit the model you wrote in part a to the data and interpret the estimates of the $\beta$ 's.
e. Conduct a test to determine whether the model from part a is statistically useful (at $\alpha=.01$ ) for estimating mean body mass.
f. Fit the model you wrote in part b to the data and interpret the estimates of the $\beta$ 's.
g. Conduct a test to determine whether the model from part $\mathbf{b}$ is statistically useful (at $\alpha=.01$ ) for estimating mean body mass.
h. Fit the model you wrote in part $\mathbf{c}$ to the data and interpret the estimates of the $\beta$ 's.
i. Conduct a test to determine whether the model from part $\mathbf{c}$ is statistically useful $($ at $\alpha=.01)$ for estimating mean egg length.

Lucas Finney
Lucas Finney
Numerade Educator
01:15

Problem 98

Refer to the Geographical Analysis (Vol. 42, 2010) study of the permeability of sandstone exposed to the weather, Exercise 2.69 (p. 91). Recall that blocks of sandstone were cut into 300 equal-sized slices and the slices randomly divided into three groups of 100 slices each. Slices in group A were not exposed to any type of weathering; slices in group $\mathrm{B}$ were repeatedly sprayed with a $10 \%$ salt solution (to simulate wetting by driven rain) under temperate conditions; and slices in group C were soaked in a $10 \%$ salt solution and then dried (to simulate blocks of sandstone exposed during a wet winter and dried during a hot summer). All sandstone slices were then tested for permeability, a measure of pressure decay in milliDarcies $(\mathrm{mD}) .$ The data for the study (simulated) are saved in the SAND file.
a. Write a model for mean permeability, $E(y),$ as a function of sandstone $\operatorname{group}(\mathrm{A}, \mathrm{B},$ or $\mathrm{C})$
b. Measures of central tendency for the permeability measurements of each sandstone group are displayed in the MINITAB printout below. Use this information to estimate the $\beta$ parameters of the model, part a.
c. Fit the model, part a, to the data in the SAND file. Use the output to verify your $\beta$ estimates in part $\mathbf{b}$.

Dominador Tan
Dominador Tan
Numerade Educator
04:01

Problem 99

How communities respond to a disaster or a violent crime was the subject of research published in the American Journal of Community Psychology (Vol. 44, 2009). Psychologists at the University of California tracked monthly violent crime incidents in two Texas cities, Jasper and Center, before and after the murder of a Jasper citizen that had racial overtones and heavy media coverage. (Center, Texas, was selected as comparison city since it had roughly the same population and racial makeup as Jasper.) Using monthly data on violent crimes, the researchers fit the regression model:
$$
E(y)=\beta_{0}+\beta_{1} x_{1}+\beta_{2} x_{2}+\beta_{3} x_{1} x_{2}
$$
where $y=$ violent crime rate (number of crimes per 1,000 population) $, x_{1}=\{1$ if Jasper, 0 if Center $\},$ and $x_{2}=\{1$ if after the murder, 0 if before the murder $\}$.
a. In terms of the $\beta$ 's in the model, what is the mean violent crime rate in Center, Texas, for months following the murder?
b. In terms of the $\beta$ 's in the model, what is the mean violent crime rate in Jasper, Texas, for months following the murder?
c. For months following the murder, find the difference between the mean violent crime rate for Jasper and Center. (Use your answers to parts a and b.)
d. Repeat part $\mathbf{c}$ for months before the murder.
e. Note that the differences, parts $\mathbf{c}$ and $\mathbf{d}$, are not the same. Explain why this illustrates the notion of interaction between $x_{1}$ and $x_{2}$.
f. A test for $H_{0}: \beta_{3}=0$ yielded a $p$ -value $<.001$. Using $\alpha=.01$, interpret this result.
g. The regression resulted in the following $\beta$ -estimates:
$\hat{\beta}_{1}=-429, \hat{\beta}_{2}=-169, \hat{\beta}_{3}=255 .$ Use these estimates
to illustrate that average monthly violent crime decreased in Center after the murder but increased in Jasper.

Sheryl Ezze
Sheryl Ezze
Numerade Educator
01:44

Problem 100

Consider a multiple-regression model for a response $y$ with one quantitative independent variable $x_{1}$ and one qualitative variable at three levels.
a. Write a first-order model that relates the mean response $E(y)$ to the quantitative independent variable.
b. Add the main-effect terms for the qualitative independent variable to the model of part a. Specify the coding scheme you use.
c. Add terms to the model of part $\mathbf{b}$ to allow for interaction between the quantitative and qualitative independent variables.
d. Under what circumstances will the response lines of the model in part $\mathbf{c}$ be parallel?
e. Under what circumstances will the model in part $\mathbf{c}$ have only one response line?

Shu Naito
Shu Naito
Numerade Educator
03:34

Problem 101

Refer to Exercise 12.100
a. Write a complete second-order model that relates $E(y)$ to the quantitative variable.
b. Add the main-effect terms for the qualitative variable (at three levels) to the model of part a.
c. Add terms to the model of part $\mathbf{b}$ to allow for interaction between the quantitative and qualitative independent variables.
d. Under what circumstances will the response curves of the model have the same shape, but different $y$ -intercepts?
e. Under what circumstances will the response curves of the model be parallellines?
f. Under what circumstances will the response curves of the model be identical?

Raymond Matshanda
Raymond Matshanda
Numerade Educator
View

Problem 102

Write a model that relates $E(y)$ to two independent variables, one quantitative and one qualitative (at four levels). Construct a model that allows the associated response curves to be second order but does not allow for interaction between the two independent variables.

Shu Naito
Shu Naito
Numerade Educator
01:33

Problem 103

Consider the model
$$
y=\beta_{0}+\beta_{1} x_{1}+\beta_{2} x_{2}+\beta_{3} x_{3}+\varepsilon
$$
where $x_{1}$ is a quantitative variable and $x_{2}$ and $x_{3}$ are dummy variables describing a qualitative variable at three levels, using the coding scheme
$$
x_{2}=\left\{\begin{array}{ll}
1 & \text { if level } 2 \\
0 & \text { otherwise }
\end{array} \quad x_{3}=\left\{\begin{array}{ll}
1 & \text { if level } 3 \\
0 & \text { otherwise }
\end{array}\right.\right.
$$
The resulting least squares prediction equation is
$$
\hat{y}=44.8+2.2 x_{1}+9.4 x_{2}+15.6 x_{3}
$$
a. What is the response line (equation) for $E(y)$ when $x_{2}=x_{3}=0 ?$ When $x_{2}=1$ and $x_{3}=0 ?$ When
$$
x_{2}=0 \text { and } x_{3}=1 ?
$$
b. What is the least squares prediction equation associated with level 1? Level 2? Level 3? Plot these on the same graph.

James Kiss
James Kiss
Numerade Educator
01:33

Problem 104

Consider the model
$$
\begin{aligned}
y=& \beta_{0}+\beta_{1} x_{1}+\beta_{2} x_{1}^{2}+\beta_{3} x_{2}+\beta_{4} x_{3}+\beta_{5} x_{1} x_{2} \\
&+\beta_{6} x_{1} x_{3}+\beta_{7} x_{1}^{2} x_{2}+\beta_{8} x_{1}^{2} x_{3}+\varepsilon
\end{aligned}
$$
where $x_{1}$ is a quantitative variable and
$$
x_{2}=\left\{\begin{array}{ll}
1 & \text { if level } 2 \\
0 & \text { otherwise }
\end{array} \quad x_{3}=\left\{\begin{array}{ll}
1 & \text { if level } 3 \\
0 & \text { otherwise }
\end{array}\right.\right.
$$
The resulting least squares prediction equation is
$$
\begin{aligned}
\hat{y}=& 45.6-4.7 x_{1}+.07 x_{1}^{2}-3.7 x_{2}-5.9 x_{3}+3.1 x_{1} x_{2} \\
&+2.5 x_{1} x_{3}-.09 x_{1}^{2} x_{2}-.03 x_{1}^{2} x_{3}
\end{aligned}
$$
a. What is the response curve for $E(y)$ when $x_{2}=0$ and $x_{3}=0 ?$ When $x_{2}=1$ and $x_{3}=0 ?$ When $x_{2}=0$ and $x_{3}=1 ?$
b. On the same graph, plot the least squares prediction equation associated with level 1 , with level 2 , and with level 3 . Choose the correct graph below.

James Kiss
James Kiss
Numerade Educator
View

Problem 105

Refer to the Body Image:
An International Journal of Research (Mar. 2010) study of the impact of reality TV shows on a college student's decision to undergo cosmetic surgery, Exercise 12.25 (p. 700). Consider the interaction model $E(y)=\beta_{0}+\beta_{1} x_{1}+$ $\beta_{2} x_{4}+\beta_{3} x_{1} x_{4},$ where $y=$ desire to have cosmetic surgery (25-point scale), $x_{1}=\{1$ if male, 0 if female $\},$ and $x_{4}=$ impression of reality TV (7-point scale). The model was fit to the data and the resulting SAS printout appears at the bottom of the page.
a. Give the least squares prediction equation.
b. Find the predicted level of desire $(y)$ for a male college student with an impression-of-reality-TV-scale score of 5
c. Conduct a test of overall model adequacy. Use $\alpha=.10$
d. Give a practical interpretation of $R_{a}^{2}$.
e. Give a practical interpretation of $s$.
f. Conduct a test (at $\alpha=.10$ ) to determine if gender $\left(x_{1}\right)$ and impression of reality TV show $\left(x_{4}\right)$ interact in the prediction of level of desire for cosmetic surgery $(y)$.
g. Give an estimate of the change in desire $(y)$ for every 1-point increase in impression of reality TV show $\left(x_{4}\right)$ for female students.
h. Repeat part $\mathbf{g}$ for male students.

Victor Salazar
Victor Salazar
Numerade Educator
View

Problem 106

Refer to the American Journal of Physics (Mar. 2014) study of the impact of dropping ping-pong balls, Exercise $11.122(\mathrm{p} .666)$ Recall that 19 standard ping-pong balls were dropped vertically onto a force plate. The data on $y=$ coefficient of restitution (COR, measured as a ratio of the speed at impact and rebound speed), $x_{1}=$ speed at impact (meters/second), and $x_{2}=\{1$ if ball buckled, 0 if not $\}$ are reproduced in the table on p. 741. Consider the interaction model $E(y)=\beta_{0}+\beta_{1} x_{1}+\beta_{2} x_{2}+\beta_{3} x_{1} x_{2}$.
A MINITAB printout of the regression analysis is also shown on p. 741 .
a. Give the equation of the hypothesized line relating $\mathrm{COR}(y)$ to impact speed $\left(x_{1}\right)$ for ping-pong balls that did not buckle. What is the estimated slope of this line?
b. Repeat part a for ping-pong balls that buckled.
c. The researcher believes that the rate of increase in COR with impact speed differs depending on whether the ping-pong ball buckles. Do the data support this hypothesis? Conduct the appropriate test using $\alpha=.05$.

Lainey Roebuck
Lainey Roebuck
Numerade Educator
06:01

Problem 107

Refer to the Economic Letters (Vol. 100,2008 ) study of whether the color of a female solicitor's hair affects the level of capital raised, Exercise 12.88 (p. 731). Recall that 955 households were contacted by a female solicitor to raise funds for hazard mitigation research. In addition to the household's level of contribution (in dollars) and the hair color of the solicitor (blonde Caucasian, brunette Caucasian, or minority female), the researcher also recorded the beauty rating of the solicitor (measured quantitatively on a 10 -point scale).
a. Write a first-order model (with no interaction) for mean contribution level, $E(y)$, as a function of a solicitor's hair color and her beauty rating.
b. Refer to the model, part a. For each hair color, express the change in contribution level for each 1 -point increase in solicitor's beauty rating in terms of the model parameters.
c. Write an interaction model for mean contribution level, $E(y)$, as a function of a solicitor's hair color and her beauty rating.
d. Refer to the model, part $\mathbf{c}$. For each hair color, express the change in contribution level for each 1 -point increase in solicitor's beauty rating in terms of the model parameters.
e. Refer to the model, part $\mathbf{c}$. Illustrate the interaction with a graph.

Beth Stone
Beth Stone
Numerade Educator
01:09

Problem 108

Refer to the Electronic Journal of Sociology (2007) study of the impact of race on the value of professional football players' "rookie" cards, presented in Exercise 12.89 (p. 732). Recall that the sample consisted of 148 rookie cards of National Football League (NFL) players who were inducted into the Football Hall of Fame. The researchers modeled the natural logarithm of card price $(y)$ as a function of the following independent variables:
a. The model $E(y)=\beta_{0}+\beta_{1} x_{1}+\beta_{2} x_{2}+\beta_{3} x_{3}+\beta_{4} x_{4}+$
$\beta_{5} x_{5}+\beta_{6} x_{6}+\beta_{7} x_{7}+\beta_{8} x_{8}+\beta_{9} x_{9}+\beta_{10} x_{10}+\beta_{11} x_{11}$
$+\beta_{11} x_{11}+\beta_{12} x_{12}$ was fit to the data, with the following results: $R^{2}=.705,$ adj. $R^{2}=.681, F=26.9$. Interpret the results practically. Make an inference about the overall adequacy of the model.
b. Refer to part
a. Statistics for the race variable were reported as follows: $\hat{\beta}_{1}=-147$, $s_{\hat{\beta}_{1}}=.145, t=-1.014, p$ -value $=.312 .$ Use this information to make an inference about the impact of race on the value of professional football players' rookie cards.
c. Refer to part a. Statistics for the card vintage variable were reported as follows: $\hat{\beta}_{3}=-.074, s_{\hat{\beta}_{2}}=.007$, $t=-10.92, p$ -value $=.000 .$ Use this information to make an inference about the impact of card vintage on the value of professional football players' rookie cards.
d. Write a first-order model for $E(y)$ as a function of card vintage $\left(x_{4}\right)$ and position $\left(x_{5}-x_{12}\right)$ that allows for the relationship between price and vintage to vary with position.

Dominador Tan
Dominador Tan
Numerade Educator
00:54

Problem 109

Refer to the Marine Mammal Science (Apr. 2010) study of whales entangled in fishing gear, Exercise $12.87,(\mathrm{p} .731)$. Now consider a model for the length $(y)$ of an entangled whale (in meters) that is a function of water depth of the entanglement (in meters) and gear type (set nets, pots, or gill nets).
a. Write a main-effects-only model for $E(y)$.
b. Sketch the relationships hypothesized by the model, part a. (Hint: Plot length on the vertical axis and water depth on the horizontal axis.)
c. Add terms to the model, part a, that include interaction between water depth and gear type. (Hint: Be sure to interact each dummy variable for gear type with water depth.)
d. Sketch the relationships hypothesized by the model, part $\mathbf{c}$
e. In terms of the $\beta$ 's in the model of part $\mathbf{c},$ give the rate of change of whale length with water depth for set nets.
f. Repeat part e for pots.
g. Repeat part e for gill nets.
h. In terms of the $\beta$ 's in the model of part $\mathbf{c}$, how would you test to determine if the rate of change of whale length with water depth is the same for all three types of fishing gear?

Nick Auwerda
Nick Auwerda
Numerade Educator
31:17

Problem 110

Refer to the Chance (Winter 2000 ) study of men's and women's winning times in the Boston Marathon, presented in Exercise 11.146 (p. 675$)$. Suppose the researchers want to build a model for predicting winning time $(y)$ of the marathon as a function of year $\left(x_{1}\right)$ in which race is run and gender of winning runner $\left(x_{2}\right)$
a. Set up the appropriate dummy variables (if necessary) for $x_{1}$ and $x_{2}$.
b. Write the equation of a model that proposes parallel straight-line relationships between winning time (y) and year $\left(x_{1}\right),$ one line for each gender.
c. Write the equation of a model that proposes nonparallel straight-line relationships between winning time ( $y$ ) and year $\left(x_{1}\right)$, one line for each gender.
d. Which of the models do you think will provide the best predictions of winning time $(y)$ ? Base your answer on the graph displayed in Exercise 11.146 .

Nathalie Luna
Nathalie Luna
The University of Texas Rio Grande Valley
05:50

Problem 111

Refer to the Journal of Applied Psychology (Jan. 2011) study of the relationship between task performance and conscientiousness, Exercise 12.67 (p. 722). Recall that the researchers used a quadratic model to relate $y=$ task performance score (measured on a 30 -point scale) to $x_{1}=$ conscientiousness score (measured on a scale of -3 to +3 ). In addition, the researchers included job complexity in the model, where $x_{2}=\{1$ if highly complex job, 0 if not $\}$. The complete model took the form
$E(y)=\beta_{0}+\beta_{1} x_{1}+\beta_{2}\left(x_{1}\right)^{2}+\beta_{3} x_{2}+\beta_{4} x_{1} x_{2}+\beta_{5}\left(x_{1}\right)^{2} x_{2}$
a. For jobs that are not highly complex, write the equation of the model for $E(y)$ as a function of $x_{1}$. (Substitute $x_{2}=0$ into the equation.)
b. Refer to part a. What do each of the $\beta$ 's represent in the model?
c. For highly complex jobs, write the equation of the model for $E(y)$ as a function of $x_{1} .$ (Substitute $x_{2}=1$ into the equation.)
d. Refer to part $\mathbf{c}$. What do each of the $\beta$ 's represent in the model?
e. Does the model support the researchers' theory that the curvilinear relationship between task performance score $(y)$ and conscientiousness score $\left(x_{1}\right)$ depends on job complexity $\left(x_{2}\right) ?$ Explain.

Heather Duong
Heather Duong
Numerade Educator
13:46

Problem 112

Do agreeable individuals get paid less, on average, than those who are less agreeable on the job? And is this gap greater for males than for females? These questions were addressed in the Journal of Personality and Social Psychology (Feb. 2012). Several variables were measured for each in a sample of individuals enrolled in the National Survey of Midlife Development in the United States. Three of these variables are: (1) level of agreeableness score (where higher scores indicate a greater level of agreeableness),
(2) gender (male or female), and (3) annual income (dollars). The researchers modeled mean income, $E(y),$ as a function of both agreeableness score $\left(x_{1}\right)$ and a dummy variable for gender $\left(x_{2}=1\right.$ if male, 0 if female $) .$ Data for a sample of 100 individuals (simulated, based on information provided in the study) are saved in the WAGAP file. The first 10 observations are listed in the accompanying table.
$$
\begin{aligned}
&\text { Data for First } 10 \text { Individuals in Study }\\
&\begin{array}{lcc}
\hline \text { Income } & \text { Agree Score } & \text { Gender } \\
\hline 44,770 & 3.0 & 1 \\
51,480 & 2.9 & 1 \\
39,600 & 3.3 & 1 \\
24,370 & 3.3 & 0 \\
15,460 & 3.6 & 0 \\
43,730 & 3.8 & 1 \\
48,330 & 3.2 & 1 \\
25,970 & 2.5 & 0 \\
17,120 & 3.5 & 0 \\
20,140 & 3.2 & 0 \\
\hline
\end{array}
\end{aligned}
$$
a. Consider the model $E(y)=\beta_{0}+\beta_{1} x_{1}+\beta_{2} x_{2} .$ The researchers theorize that for either gender, income will decrease as agreeableness score increases. If this theory is true, what is the expected sign of $\beta_{1}$ in the model?
b. The researchers also theorize that the rate of decrease of income with agreeableness score will be steeper for males than for females (i.e., the income gap between males and females will be greater the less agreeable the individuals are). Can this theory be tested using the model, part a? Explain.
c. Consider the interaction model $E(y)=\beta_{0}+\beta_{1} x_{1}+$ $\beta_{2} x_{2}+\beta_{3} x_{1} x_{2} .$ If the theory, part $\mathbf{b},$ is true, give the expected sign of $\beta_{1}$ and the expected sign of $\beta_{3}$
d. Fit the model, part $\mathbf{c},$ to the sample data. Check the signs of the estimated $\beta$ coefficients. How do they compare to the expected values, part $\mathbf{c}$ ?
e. Refer to the interaction model, part c. Give the null and alternative hypotheses for testing whether the rate of decrease of income with agreeableness score is steeper for males than for females.
f. Conduct the test, part e. Use $\alpha=.05$. Is the researchers' theory supported?

Heather Duong
Heather Duong
Numerade Educator
01:56

Problem 113

Workplace bullying is defined as harassment, persistent criticism, withholding of key information, spreading of rumors, or intimidation on the job. In Human Resource Management Journal (Oct. 2008), researchers employed multiple regression to examine whether perceived organizational support would moderate the relationship between workplace bullying and victims' intention to leave the firm. The dependent variable in the analysis, intention to leave (y), was measured on a quantitative scale. The two key independent variables in the study were bullying $\left(x_{1},\right.$ measured on a quantitative scale) and perceived organizational support (measured qualitatively as "low," "neutral," or "high").
a. Set up the dummy variables required to represent perceived organizational support (POS) in the regression model.
b. Write a model for $E(y)$ as a function of bullying and POS that hypothesizes three parallel straight lines, one for each level of POS.
c. Write a model for $E(y)$ as a function of bullying and POS that hypothesizes three nonparallel straight lines, one for each level of $\mathrm{POS}$.
d. The researchers discovered that the effect of bullying on intention to leave was greater at the low level of POS than at the high level of POS. Which of the two models, parts $\mathbf{b}$ and $\mathbf{c}$, supports these findings?

Jake Zanazzi
Jake Zanazzi
Numerade Educator
03:26

Problem 114

A study of the atmospheric pollution on the slopes of the Blue Ridge Mountains (in Tennessee) was conducted. The file MOSS contains the levels of lead found in 70 fern moss specimens (in micrograms of lead per gram of moss tissue) collected from the mountain slopes, as well as the elevation of the moss specimen (in feet) and the direction (1 if east, 0 if west) of the slope face. The first five and last five observations of the data set are listed in the following table.
$$
\begin{array}{cccc}
\hline \text { Specimen } & \text { Lead Level } & \text { Elevation } & \text { Slope Face } \\
\hline 1 & 3.475 & 2,000 & 0 \\
2 & 3.359 & 2,000 & 0 \\
3 & 3.877 & 2,000 & 0 \\
4 & 4.000 & 2,500 & 0 \\
5 & 3.618 & 2,500 & 0 \\
\vdots & \vdots & \vdots & \vdots \\
66 & 5.413 & 2,500 & 1 \\
67 & 7.181 & 2,500 & 1 \\
68 & 6.589 & 2,500 & 1 \\
69 & 6.182 & 2,000 & 1 \\
70 & 3.706 & 2,000 & 1 \\
\hline
\end{array}
$$
a. Write the equation of a first-order model relating mean lead level $E(y)$ to elevation $\left(x_{1}\right)$ and slope face $\left(x_{2}\right)$. Include interaction between elevation and slope face in the model.
b. Graph the relationship between mean lead level and elevation for the different slope faces that is hypothesized by the model you wrote in part a.
c. In terms of the $\beta$ 's of the model from part a give the change in lead level for every 1 -foot increase in elevation for moss specimens on the east slope.
d. Fit the model from part a to the data, using an available statistical software package. Is the overall model statistically useful in predicting lead level? Test, using $\alpha=.10$

Nick Johnson
Nick Johnson
Numerade Educator
02:30

Problem 115

Refer to the Journal of Agricultural, Biological, and Environmental Statistics (Mar. 2005$)$ study of the chemical composition of rainwater, presented in Exercise 12.90 (p. 732). Recall that the nitrate concentration $y$ (milligrams per liter) in a sample of rainwater was modeled as a function of water source (groundwater, subsurface flow, or overground flow). Now consider adding a second independent variable, silica concentration (milligrams per liter), to the model.
a. Write a first-order model for $E(y)$ as a function of the independent variables. Assume that the rate of increase of nitrate concentration with silica concentration is the same for all three water sources. Sketch the relationships hypothesized by the model on a graph.
b. Write a first-order model for $E(y)$ as a function of the independent variables, but now assume that the rate of increase of nitrate concentration with silica concentration differs for the three water sources. Sketch the relationships hypothesized by the model on a graph.

Chengyu Li
Chengyu Li
Numerade Educator
05:17

Problem 116

Many women suffer from anemia. A female physician who is also an avid jogger wanted to know if women who exercise regularly have a different mean red blood cell count than women who do not. She also wanted to know if the amount of a particular iron supplement a woman takes has any effect and whether the effect is the same for both groups. Write a model that will reflect the relationship between red blood cell count and the two independent variables just described, assuming that
a. the effect of the iron supplement on mean blood cell count is the same regardless of whether a woman exercises regularly.
b. the effect of the iron supplement on mean blood cell count depends on whether a woman exercises regularly.

Jacob Fry
Jacob Fry
Numerade Educator
View

Problem 117

Determine which pairs of models that follow are nested models. For each pair of nested models, identify the complete and reduced model.
a. $E(y)=\beta_{0}+\beta_{1} x_{1}+\beta_{2} x_{2}$
b. $E(y)=\beta_{0}+\beta_{1} x_{1}$
c. $E(y)=\beta_{0}+\beta_{1} x_{1}+\beta_{2} x_{1}^{2}$
d. $E(y)=\beta_{0}+\beta_{1} x_{1}+\beta_{2} x_{2}+\beta_{3} x_{1} x_{2}$
e. $E(y)=\beta_{0}+\beta_{1} x_{1}+\beta_{2} x_{2}+\beta_{3} x_{1} x_{2}+\beta_{4} x_{1}^{2}+\beta_{5} x_{2}^{2}$

Shu Naito
Shu Naito
Numerade Educator
02:40

Problem 118

Explain why the $F$ -test used to compare complete and reduced models is a one-tailed, upper-tailed test.

Erin Moser
Erin Moser
Numerade Educator
01:20

Problem 119

What is a parsimonious model?

Sneha Ravi
Sneha Ravi
Numerade Educator
02:35

Problem 120

Suppose you fit the regression model
$$
y=\beta_{0}+\beta_{1} x_{1}+\beta_{2} x_{2}+\beta_{3} x_{1} x_{2}+\beta_{4} x_{1}^{2}+\beta_{5} x_{2}^{2}+\varepsilon
$$
to $n=30$ data points and you wish to test
$$
H_{0}: \beta_{3}=\beta_{4}=\beta_{5}=0
$$
a. State the alternative hypothesis $H_{\mathrm{a}}$
b. Give the reduced model appropriate for conducting the test.
c. What are the numerator and denominator degrees of freedom associated with the $F$ -statistic?
d. Suppose the SSE's for the reduced and complete models are $\mathrm{SSE}_{R}=1,250.2$ and $\mathrm{SSE}_{C}=1,125.2,$ respectively. Conduct the hypothesis test and interpret the results. Use $\alpha=.05$.

Nick Johnson
Nick Johnson
Numerade Educator
02:35

Problem 121

The complete model
$$
y=\beta_{0}+\beta_{1} x_{1}+\beta_{2} x_{2}+\beta_{3} x_{3}+\beta_{4} x_{4}+\varepsilon
$$
was fit to $n=22$ data points, with $\mathrm{SSE}=146.99 .$ The reduced model $y=\beta_{0}+\beta_{1} x_{1}+\beta_{2} x_{2}+\varepsilon$ was also fit,
$$
\text { with } \mathrm{SSE}=165.37
$$
a. How many $\beta$ parameters are in the complete model? The reduced model?
b. Specify the null and alternative hypotheses you would use to investigate whether the complete model contributes more information for the prediction of $y$ than the reduced model.
c. Conduct the hypothesis test of part b. Use $\alpha=.10$. Determine the test statistic.

Nick Johnson
Nick Johnson
Numerade Educator
05:29

Problem 122

How willing are you to donate your organs upon death? This was the question of interest in research published in Death Studies (Vol. 38,2014 ). A sample of $n=307$ Australians participated in the study. Willingness to donate organs upon death was measured for each participant on a 7 -point scale $(1=$ very unlikely $, 7=$ very likely $) .$ Multiple regression was used to model willingness $(y)$ as a function of general attitude $\left(x_{1}\right)$ toward donating organs (28-point scale), subjective norm $\left(x_{2}\right)-$ assessment of whether others would approve of organ donation (14-point scale), prototype favorability $\left(x_{3}\right)-$ opinion of those who do donate organs (14-point scale), prototype similarity $\left(x_{4}\right)-$ similarity to organ donors (14-point scale), age $\left(x_{5}\right)$ in years, and gender $\left(x_{6}\right)$.
a. Write a first-order, main-effects model for $E(y)$ as a function of the six independent variables.
b. The model, part a, resulted in $R^{2}=.39$ and global $F$ $p$ -value $<.001$. Interpret these results.
c. What term would you add to the model, part a, to allow for the effect of prototype favorability on willingness to depend on prototype similarity?
d. The model, part $\mathbf{c},$ resulted in $R^{2}=.41$ and global $F$ $p$ -value $<.001 .$ Interpret these results.
e. Are the models, parts a and $\mathbf{c},$ nested models?
f. A nested model $F$ -test comparing the models resulted in a test statistic of $F=10.94$ with a $p$ -value of .001 . Interpret these results.

Leah Lampen
Leah Lampen
Numerade Educator
02:57

Problem 123

Refer to the Human Factors (Mar. 2014) study of shared leadership by the cockpit and cabin crews of a commercial airplane, Exercise 9.13 (p. 476). Recall that simulated flights were taken by 84 six-person crews, where each crew consisted of a two-person cockpit (captain and first officer) and a four-person cabin team (three flight attendants and a purser). During the simulation, smoke appeared in the cabin and the reactions of the crew were monitored for teamwork. One key variable in the study was team goal attainment score, measured on a 0 -to-60-point scale. Multiple regression analysis was used to model team goal attainment $(y)$ as a function of the independent variables job experience of purser $\left(x_{1}\right),$ job experience of head flight attendant $\left(x_{2}\right),$ gender of purser $\left(x_{3}\right)$, gender of head flight attendant $\left(x_{4}\right)$, leadership score of purser $\left(x_{5}\right),$ and leadership score of head flight attendant $\left(x_{6}\right)$.
a. Write a complete, first-order model for $E(y)$ as a function of the six independent variables.
b. Consider a test of whether the leadership score of either the purser or head flight attendant (or both) is statistically useful for predicting team goal attainment. Give the null and alternative hypotheses as well as the reduced model for this test.
c. The two models were fit to the data for the $n=60$ successful cabin crews with the following results:
$R^{2}=.02$ for reduced model, $R^{2}=.25$ for complete model. Based on this information only, give your opinion regarding the null hypothesis for successful cabin crews.
d. The $p$ -value of the subset $F$ -test for comparing the two models for successful cabin crews was reported in the article as $p<.05 .$ Formally test the null hypothesis using $\alpha=.05 .$ What do you conclude?
e. The two models were also fit to the data for the $n=24$ unsuccessful cabin crews with the following results: $R^{2}=.14$ for reduced model, $R^{2}=.15$ for complete model. Based on this information only, give your opinion regarding the null hypothesis for unsuccessful cabin crews.
f. The $p$ -value of the subset $F$ -test for comparing the two models for unsuccessful cabin crews was reported in the article as $p>.10$. Formally test the null hypothesis using $\alpha=.05 .$ What do you conclude?

Nick Johnson
Nick Johnson
Numerade Educator
03:59

Problem 124

Mental health of a community. An article in the Community Mental Health Journal (Aug. 2000) used multiple-regression analysis to model the level of community adjustment of clients of the Department of Mental Health and Addiction Services in Connecticut. The dependent variable, community adjustment $(y),$ was measured quantitatively on the basis of staff ratings of the clients. (Lower scores indicate better adjustment.) The complete model was a first-order model with 21 independent variables. The independent variables were categorized as demographic (four variables), diagnostic (seven variables), treatment (four variables), and community (six variables).
a. Write the equation of $E(y)$ for the complete model.
b. Give the null hypothesis for testing whether the seven diagnostic variables contribute information relevant to the prediction of $y$.
c. Give the equation of the reduced model appropriate for the test suggested in part $\mathbf{b}$.
d. The test in part b was carried out and resulted in a test statistic of $F=59.3$ and $p$ -value $<.0001$. Interpret this result in the words of the problem.

Lucas Finney
Lucas Finney
Numerade Educator
01:56

Problem 125

Refer to the Human Resource Management Journal (Oct. 2008) study of workplace bullying, Exercise 12.113 (p. 743 ). Recall that multiple regression was used to model an employee's intention to leave $(y)$ as a function of bullying $\left(x_{1},\right.$ measured on a quantitative scale) and perceived organizational support (measured qualitatively as "low POS," "neutral POS," or "high POS"). In Exercise 12.113b, you wrote a model for $E(y)$ as a function of bullying and POS that hypothesizes three parallel straight lines, one for each level of POS. In Exercise $12.113 \mathbf{c},$ you wrote a model for $E(y)$ as a function of bullying and POS that hypothesizes three nonparallel straight lines, one for each level of $\mathrm{POS}$.
a. Explain why the two models are nested. Which is the complete model? Which is the reduced model?
b. Give the null hypothesis for comparing the two models.
c. If you reject $H_{0}$ in part $\mathbf{b}$, which model do you prefer? Why?
d. If you fail to reject $H_{0}$ in part $\mathbf{b}$, which model do you prefer? Why?

Jake Zanazzi
Jake Zanazzi
Numerade Educator
03:29

Problem 126

Refer to the Journal of Engineering for Gas Turbines and Power (Jan. 2005) study of a high-pressure inlet fogging method for a gas turbine engine, presented in Exercise 12.27 (p. 701). Consider a model for the heat rate (kilojoules per kilowatt per hour) produced by a gas turbine as a function of cycle speed (revolutions per minute) and cycle pressure ratio.
a. Write a complete second-order model for heat rate $(y)$.
b. Give the null and alternative hypotheses for determining whether the curvature terms in the complete second-order model are statistically useful in predicting the heat rate $(y)$.
c. For the test in part $\mathbf{b}$, identify the complete and reduced models.
d. Portions of the MINITAB printouts for the two models are shown above. Find the values of $\mathrm{SSE}_{R}, \mathrm{SSE}_{C}$ and $\mathrm{MSE}_{C}$ on the printouts.
e. Compute the value of the test statistics for the test of part b.
f. Find the rejection region for the test of part
b. Use $\alpha=.10$
g. State the conclusion of the test in the words of the problem.

Lucas Finney
Lucas Finney
Numerade Educator
19:12

Problem 127

Study of supervisor-targeted aggression. "Moonlighters" are workers who hold two jobs at the same time. What are the factors that affect the likelihood of a moonlighting worker becoming aggressive toward his or her supervisor? This was the research question of interest in the Journal of Applied Psychology (July 2005). Completed questionnaires were obtained from $n=105$ moonlighters, and the data were used to fit several multiple-regression models for supervisor-targeted aggression score $(y)$. Two of the models (with $R^{2}$ values in parentheses) are shown in the table below.
a. Interpret the $R^{2}$ values for the models.
b. Give the null and alternative hypotheses for comparing the fits of Models 1 and 2 .
c. Are the two models nested? Explain.
d. The nested $F$ -test for comparing the two models resulted in $F=42.13$ and $p$ -value $<.001 .$ What can you conclude from these results?
e. A third model was fit, one that hypothesizes all possible pairs of interactions between self-esteem, history of aggression, interactional injustice at primary job, and abusive supervisor at primary job. Give the equation of this model (Model 3).
f. A nested $F$ -test to compare Models 2 and 3 resulted in $p$ -value $>.10$. What can you conclude from this result?

Heather Duong
Heather Duong
Numerade Educator
View

Problem 128

Reality TV and cosmetic surgery. Refer to the Body Image:
An International Journal of Research (Mar. 2010) study of the influence of reality TV shows on one's desire to undergo cosmetic surgery, Exercise 12.25 (p. 700 ). Recall that psychologists modeled desire to have cosmetic surgery $(y)$ as a function of gender $\left(x_{1}\right),$ self-esteem $\left(x_{2}\right),$ body satisfaction $\left(x_{3}\right),$ and impression of reality $\mathrm{TV}\left(x_{4}\right) .$ The psychologists theorize that one's impression of reality TV will "moderate" the impact that each of the first three independent variables has on one's desire to have cosmetic surgery. If so, then $x_{4}$ will interact with each of the other independent variables.
a. Give the equation of the model for $E(y)$ that matches the theory.
b. Fit the model, part a, to the data saved in the IMAGE file. Evaluate the overall utility of the model.
c. Give the null hypothesis for testing the psychologists? theory.
d. Conduct a nested model $F$ -test to test the theory. What do you conclude?

Rashmi Sinha
Rashmi Sinha
Numerade Educator
05:50

Problem 129

Personality traits and job performance. Refer to the Journal of Applied Psychology (Jan. 2011) study of the relationship between task performance and conscientiousness, Exercise 12.111 (p. 742). Recall that $y=$ task performance score (measured on a 30 -point scale) was modeled as a function of $x_{1}=$ conscientiousness score (measured on a scale of -3 to +3 ) and $x_{2}=\{1$ if highly complex job, 0 if not $\}$ using the complete model
$E(y)=\beta_{0}+\beta_{1} x_{1}+\beta_{2}\left(x_{1}\right)^{2}+\beta_{3} x_{2}+\beta_{4} x_{1} x_{2}+\beta_{5}\left(x_{1}\right)^{2} x_{2}$
a. Specify the null hypothesis for testing the overall adequacy of the model.
b. Specify the null hypothesis for testing whether task performance score $(y)$ and conscientiousness score $\left(x_{1}\right)$ are curvilinearly related.
c. Specify the null hypothesis for testing whether the curvilinear relationship between task performance score $(y)$ and conscientiousness score $\left(x_{1}\right)$ depends on job complexity $\left(x_{2}\right)$
d. Explain how each of the tests, parts a-c, should be conducted (i.e., give form of the test statistic and the reduced model).

Heather Duong
Heather Duong
Numerade Educator
05:24

Problem 130

Glass as a waste encapsulant. The encapsulation of waste in glass is considered to be a promising solution to the problem of low-level nuclear waste. A study was undertaken jointly by the Department of Materials Science and Engineering at the University of Florida and the U.S. Department of Energy to assess the utility of glass as a waste encapsulant.* Corrosive chemical solutions (called corrosion baths) were prepared and applied directly to glass samples containing one of three types of waste $(\mathrm{TDS}-3 \mathrm{~A}, \mathrm{FE},$ and $\mathrm{AL}) ;$ the chemical reactions were observed over time. A few of the key variables measured were
$y=$ Amount of silicon (in parts per million) found in solution at end of experiment (This is both a measure of the degree of breakdown in the glass and a proxy for the amount of radioactive species released into the environment.)
$x_{1}=$ Temperature $\left({ }^{\circ} \mathrm{C}\right)$ of the corrosion bath $x_{2}=1$ if waste type $\operatorname{TDS}-3 \mathrm{~A}, 0$ if not $x_{3}=1$ if waste type $\mathrm{FE}, 0$ if not
(Waste type $\mathrm{AL}$ is the base level.) Suppose we want to model amount $y$ of silicon as a function of temperature $\left(x_{1}\right)$ and type of waste $\left(x_{2}, x_{3}\right)$
a. Write a model that proposes parallel straight-line relationships between amount of silicon and temperature, one line for each of the three types of waste.
b. Add terms for the interaction between temperature and waste type to the model from part a.
c. Refer to the model from part $\mathbf{b}$. For each type of waste, give the slope of the line relating amount of silicon to temperature.
d. Explain how you could test for the presence of temperature-type-of-waste interaction.

Khoobchandra Agrawal
Khoobchandra Agrawal
Numerade Educator
04:34

Problem 131

Refer to the Marine Mammal Science (Apr. 2010) study of whales entangled in fishing gear, Exercise 12.109 (p. 741). A first-order model for the length $(y)$ of an entangled whale that is a function of water depth of the entanglement $\left(x_{1}\right)$ and gear type (set nets, pots, or gill nets) is written as follows:
$E(y)=\beta_{0}+\beta_{1} x_{1}+\beta_{2} x_{2}+\beta_{3} x_{3}+\beta_{4} x_{1} x_{2}+\beta_{5} x_{1} x_{3}$
where $x_{2}=\{1$ if set net $, 0$ if $n o t\}$ and $x_{2}=\{1$ if pot, 0 if $n o t\}$. Consider this model the complete model in a nested model $F$ -test.
a. Suppose you want to determine if there are any differences in the mean lengths of entangled whales for the three gear types. Give the appropriate null hypothesis to test.
b. Refer to part
a. Give the reduced model for the test.
c. Refer to parts a and $\mathbf{b}$. If you reject the null hypothesis, what would you conclude?
d. Suppose you want to determine if the rate of change of whale length $(y)$ with water depth $\left(x_{1}\right)$ is the same for all three types of fishing gear. Give the appropriate null hypothesis to test.
e. Refer to part
d. Give the reduced model for the test.
f. Refer to parts $\mathbf{d}$ and $\mathbf{e}$. If you fail to reject the null hypothesis, what would you conclude?

Jameson Kuper
Jameson Kuper
Numerade Educator
View

Problem 132

Agreeableness, gender, and wages. Refer to the Journal of Personality and Social Psychology (Feb. 2012) study of on-the-job agreeableness and wages, Exercise 12.112 (p. 742). The researchers modeled mean income, $E(y)$, as a function of both agreeableness score $\left(x_{1}\right)$ and a dummy variable for gender $\left(x_{2}=1\right.$ if male, 0 if female). Suppose the researchers theorize that for either gender, income will decrease at a decreasing rate as agreeableness score increases. Consequently, they want to fit a second-order model.
a. Consider the model $E(y)=\beta_{0}+\beta_{1} x_{1}+$ $\beta_{2}\left(x_{1}\right)^{2}+\beta_{3} x_{2} .$ If the researchers' belief is true, what is the expected sign of $\beta_{2}$ in the model?
b. Draw a sketch of the model, part a, showing how gender affects the income-agreeableness score relationship.
c. Write a complete second-order model for $E(y)$ as a function of $x_{1}$ and $x_{2}$.
d. Draw a sketch of the model, part $\mathbf{c}$, showing how gender affects the income-agreeableness score relationship.
e. What null hypothesis would you test in order to compare the two models, parts a and $\mathbf{c}$ ?
f. Fit the models to the sample data saved in the WAGAP file and carry out the test, part e. What do you conclude? (Test, using $\alpha=.10 .)$

Victor Salazar
Victor Salazar
Numerade Educator
03:08

Problem 134

Emotional distress in firefighters. The Journal of Human Stress (Summer 1987) reported on a study of the "psychological response of firefighters to chemical fire." It is thought that the complete second-order model
$$
E(y)=\beta_{0}+\beta_{1} x_{1}+\beta_{2} x_{1}^{2}+\beta_{3} x_{2}+\beta_{4} x_{1} x_{2}+\beta_{5} x_{1}^{2} x_{2}
$$
where
$$
\begin{aligned}
y &=\text { Emotional distress } \\
x_{1} &=\text { Experience (years) }
\end{aligned}
$$
$x_{2}=1$ if exposed to chemical fire, 0 if not
will be adequate to describe the relationship between emotional distress and years of experience for two groups of firefighters: those exposed to a chemical fire and those not exposed.
a. How would you determine whether the rate of increase of emotional distress with experience is different for the two groups of firefighters?
b. How would you determine whether there are differences in mean emotional distress levels that are attributable to exposure group?

Nick Johnson
Nick Johnson
Numerade Educator
01:26

Problem 135

Explain the difference between a stepwise model and a standard regression model.

Arwa Ali
Arwa Ali
Numerade Educator
01:18

Problem 136

Give two caveats associated with using the stepwise regression results as the final model for predicting $y$.

Arwa Ali
Arwa Ali
Numerade Educator
View

Problem 137

Suppose there are six independent variables $x_{1}, x_{2}, x_{3},$ $x_{4}, x_{5},$ and $x_{6}$ that might be useful in predicting a response $y .$ A total of $n=50$ observations is available, and it is decided to employ stepwise regression to help in selecting the independent variables that appear to be useful. The computer fits all possible one-variable models of the form
$$
E(y)=\beta_{0}+\beta_{1} x_{i}
$$
where $x_{i}$ is the $i$ th independent variable $, i=1,2, \ldots, 6$. The information in the following table is provided from the computer printout:

a. Which independent variable is declared the best one-variable predictor of $y$ ? Explain.
b. Would this variable be included in the model at this stage? Explain.
c. Describe the next phase that a stepwise procedure would execute.

Victor Salazar
Victor Salazar
Numerade Educator
08:37

Problem 138

Distribution of yellowhammer birds. An article in the Journal of Animal Ecology (July 2006) discussed some of the caveats of using stepwise regression. An example presented used stepwise regression to find the best predictors of the number of yellowhammers inhabiting lowland farms in the United Kingdom. Nine potential independent variables were considered: (1) hedge presence,
(2) tree-line presence,
(3) ditch presence,
(4) adjacent road, (5) width of margin, (6) adjacent pasture, (7) adjacent silage ley, (8) winter rape, and (9) adjacent beans.
a. How many one-variable models are fit in step 1 of the stepwise regression?
b. After selecting ditch presence in step 1 , how many two-variable models are fit in step 2 of the stepwise regression?
c. After selecting adjacent pasture in step $2,$ how many three-variable models are fit in step 3 of the stepwise regression?
d. Through the first three steps of the stepwise regression, determine the total number of $t$ -tests performed. Assuming each test uses an $\alpha=.05$ level of significance, give an estimate of the probability of at least one Type I error in the stepwise regression.

Donald Albin
Donald Albin
Numerade Educator
07:27

Problem 139

Study of sex offenders. A study of sex offenders in the Canadian Federal Prison System was published in the British Journal of Criminology (May 2014). The following data were collected for each of 59 male sex offenders: race, type of crime (violent or nonviolent), age at first conviction, and total number of criminal convictions. Suppose you want to use the data to model the total number of convictions $(y)$ as a function of race, type of crime, and age at first conviction. Do you recommend using stepwise regression to find the "best" model for predicting $y$ ? Explain. If not, outline a strategy for finding the best model.

Jeremiah Mbaria
Jeremiah Mbaria
Numerade Educator
02:29

Problem 140

Accuracy of software effort estimates. Periodically, software engineers must provide estimates of their effort in developing new software. In the Journal of Empirical Software Engineering (Vol. 9,2004$)$, multiple regression was used to predict the accuracy of these effort estimates. The dependent variable, defined as the relative error in estimating effort,
$y=($ Actual effort $-$ Estimated effort $) /($ Actual effort $)$
was determined for each in a sample of $n=49$ software development tasks. Eight independent variables were evaluated as potential predictors of relative error using stepwise regression. Each of these was formulated as a dummy variable, as shown in the table.
c. In step 2 of the stepwise regression, how many different two-variable models (where $x_{1}$ is one of the variables) are fitted to the data?
d. The only two variables selected for entry into the stepwise regression model were $x_{1}$ and $x_{8} .$ The stepwise regression yielded the following prediction equation:
$$
\hat{y}=.12-.28 x_{1}+.27 x_{8}
$$
Give a practical interpretation of the $\beta$ estimates multiplied by $x_{1}$ and $x_{8}$.
e. Why should a researcher be wary of using the model, part $\mathbf{d}$, as the final model for predicting effort $(y)$ ?

Heather Duong
Heather Duong
Numerade Educator
06:15

Problem 141

An analysis of footprints in sand. Fossilized human footprints provide a direct source of information on the gait dynamics of extinct species. To gain insight into this phenomenon, a group of scientists used human subjects (16 young adults) to generate footprints in sand ( American Journal of Physical Anthropology, Apr. 2010). One dependent variable of interest was heel depth $(y)$ of the footprint (in millimeters). The scientists wanted to find the best predictors of depth from among six possible independent variables. Three variables were related to the human subject (foot mass, leg length, and foot type) and three variables were related to walking in sand (velocity, pressure, and impulse). A stepwise regression run on these six variables yielded the following results:
Selected independent variables: pressure and leg length $R^{2}=.771,$ Global $F$ -test $p$ -value $<.001$
a. Write the hypothesized equation of the final stepwise regression model.
b. Interpret the value of $R^{2}$ for the model.
c. Conduct a test of the overall utility of the final stepwise model.
d. At minimum, how many $t$ -tests on individual $\beta$ 's were conducted to arrive at the final stepwise model?
e. Based on your answer to part $\mathbf{d}$, comment on the probability of making at least one Type I error during the stepwise analysis.

Trent Speier
Trent Speier
Numerade Educator
07:01

Problem 142

Yield strength of steel alloy. Industrial engineers at the University of Florida used regression modeling as a tool to reduce the time and cost associated with developing new metallic alloys (Modelling and Simulation in Materials Science and Engineering, Vol. 13,2005$)$. To illustrate, the engineers build a regression model for the tensile yield strength $(y)$ of a new steel alloy. The potential important predictors of yield strength are listed in the following table.
$x_{1}=$ Carbon amount (\% weight)
$x_{2}=$ Manganese amount (\% weight)
$x_{3}=$ Chromium amount (\% weight)
$x_{4}=$ Nickel amount (\% weight)
$x_{5}=$ Molybdenum amount (\% weight)
$x_{6}=$ Copper amount (\% weight)
$x_{7}=$ Nitrogen amount (\% weight)
$x_{8}=$ Vanadium amount (\% weight)
$x_{9}=$ Plate thickness (millimeters)
$x_{10}=$ Solution treating (millimeters)
$x_{11}=$ Ageing temperature ( degrees Celsius)
a. The engineers used stepwise regression to search for a parsimonious set of predictor variables. Do you agree with this decision? Explain.
b. The stepwise regression selected the following independent variables: $x_{1}=$ Carbon, $x_{2}=$ Manganese, $x_{3}=$ Chromium, $x_{5}=$ Molybdenum, $x_{6}=$ Copper, $x_{8}=$ Vanadium, $x_{9}=$ Plate thickness, $x_{10}=$ Solution treating, and $x_{11}=$ Ageing temperature. On the basis of this information, determine the total number of first-order models that were fit in the stepwise routine.
c. Refer to part b. All the variables listed there were statistically significant in the stepwise model, with $R^{2}=.94$. Consequently, the engineers used the estimated stepwise model to predict yield strength. Do you agree with this decision? Explain.

Raymond Matshanda
Raymond Matshanda
Numerade Educator
04:05

Problem 143

Bus rapid-transit study. The Center for Urban Transportation Research (CUTR) at the University of South Florida conducted a survey of Bus rapid transit (BRT) customers in Miami (Transportation Research Board Annual Meeting, Jan. 2003). Data on the following variables (all measured on a five-point scale, where $1=$ "very unsatisfied" and $5=$ "very satisfied") were collected for a sample of over 500 bus riders: overall satisfaction with BRT $(y),$ safety on bus $\left(x_{1}\right),$ seat availability $\left(x_{2}\right),$ dependability $\left(x_{3}\right),$ travel time $\left(x_{4}\right), \operatorname{cost}\left(x_{5}\right),$ information/maps $\left(x_{6}\right),$ convenience of routes $\left(x_{7}\right),$ traffic signals $\left(x_{8}\right),$ safety at bus stops $\left(x_{9}\right),$ hours of service $\left(x_{10}\right),$ and frequency of service $\left(x_{11}\right) .$ CUTR analysts used stepwise regression to model overall satisfaction $(y)$.
a. How many models are fitted at step 1 of the stepwise regression?
b. How many models are fitted at step 2 of the stepwise regression?
c. How many models are fitted at step 11 of the stepwise regression?
d. The stepwise regression selected the following eight variables to include in the model (in order of selection $): x_{11}, x_{4}, x_{2}, x_{7}, x_{10}, x_{1}, x_{9},$ and $x_{3} .$ Write the equation for $E(y)$ that results.
e. The model in part d was tested and resulted in $R^{2}=.677 .$ Interpret this value.
f. Explain why the CUTR analysts should be cautious in concluding that the "best" model for $E(y)$ has been found.

Maxime Rossetti
Maxime Rossetti
Numerade Educator
01:45

Problem 144

Modeling species abundance. A marine biologist was hired by the EPA to determine whether the hot-water runoff from a particular power plant located near a large gulf is having an adverse effect on the marine life in the area. The biologist's goal is to acquire a prediction equation for the number of marine animals located at certain designated areas, or stations, in the gulf. On the basis of past experience, the EPA considered the following environmental factors as predictors for the number of animals at a particular station:
$x_{1}=$ Temperature of water (TEMP) $x_{2}=$ Salinity of water (SAL) $x_{3}=$ Dissolved oxygen content of water $(\mathrm{DO})$ $x_{4}=$ Turbidity index, a measure of the turbidity of the water (TI)
$$
x_{5}=\text { Depth of the water at the station }\left(\mathrm{ST}_{-} \mathrm{DEPTH}\right)
$$
$x_{6}=$ Total weight of sea grasses in sampled area (TGRSWT)
As a preliminary step in the construction of this model, the biologist used a stepwise regression procedure to identify the most important of these six variables. A total of 716 samples was taken at different stations in the gulf, producing the SPSS printout shown above. (The response measured was $y,$ the logarithm of the number of marine animals found in the sampled area.)
a. According to the printout, which of the independent variables should be used in the model?
b. Are we able to assume that the marine biologist has identified all the important independent variables for the prediction of $y$ ? Why?
c. Using the variables identified in part $\mathbf{a}$, write the firstorder model with interaction that may be used to predict $y$.
d. How would the marine biologist determine whether the model specified in part $\mathbf{c}$ is better than the first-order model?
e. Note the small value of $R^{2}$. What action might the biologist take to improve the model?

Aadit Sharma
Aadit Sharma
Numerade Educator
View

Problem 145

Using corn in a duck diet. Corn is high in starch content; consequently, it is considered excellent feed for domestic chickens. Does corn possess the same potential in feeding ducks bred for broiling? This was the subject of research published in Animal Feed Science and Technology (Apr. 2010). The objective of the study was to establish a prediction model for the true metabolizable energy (TME) of corn regurgitated from ducks. The researchers considered 11 potential predictors of TME: dry matter (DM), crude protein (CP), ether extract (EE), ash (ASH), crude fiber (CF), neutral detergent fiber (NDF), acid detergent fiber (ADF), gross energy (GE), amylose (AM), amylopectin (AP), and amylopectin/amylose (AMAP). Stepwise regression was used to find the best subset of predictors. The final stepwise model yielded the following results:
$$
\begin{array}{c}
\widehat{\mathrm{TME}}=7.70+2.14(\mathrm{AMAP})+.16(\mathrm{NDF}) \\
R^{2}=.988, s=.07, \text { global } F p \text { -value }=.001
\end{array}
$$
a. Determine the number of $t$ -tests performed in step 1 of the stepwise regression.
b. Determine the number of $t$ -tests performed in step 2 of the stepwise regression.
c. Give a full interpretation of the final stepwise model regression results.
d. Explain why it is dangerous to use the final stepwise model as the "best" model for predicting TME.
e. Using the independent variables selected by the stepwise routine, write a complete second-order model for TME.
f. Refer to part $\mathbf{e}$. How would you determine if the terms in the model that allow for curvature are statistically useful for predicting TME?

Lainey Roebuck
Lainey Roebuck
Numerade Educator
02:26

Problem 146

Define a regression residual.

Ernest Castorena
Ernest Castorena
Numerade Educator
00:39

Problem 147

Define an outlier.

Emily Himsel
Emily Himsel
Numerade Educator
01:18

Problem 148

Give two properties of the regression residuals from a model.

Arwa Ali
Arwa Ali
Numerade Educator
01:46

Problem 149

True or False. Regression models fit to time-series data typically result in uncorrelated errors.

Oluwadamilola Ameobi
Oluwadamilola Ameobi
Numerade Educator
02:26

Problem 150

Define multicollinearity in regression.

Ernest Castorena
Ernest Castorena
Numerade Educator
00:51

Problem 151

Give three indicators of a multicollinearity problem.

Emily Himsel
Emily Himsel
Numerade Educator
00:51

Problem 152

Give three indicators of a multicollinearity problem.

Emily Himsel
Emily Himsel
Numerade Educator
00:51

Problem 153

Consider fitting the multiple regression model
$E(y)=\beta_{0}+\beta_{1} x_{1}+\beta_{2} x_{2}+\beta_{3} x_{3}+\beta_{4} x_{4}+\beta_{5} x_{5}$
A matrix of correlations for all pairs of independent variables is given on the right. Do you detect a multicollinearity problem? Explain.

Victor Salazar
Victor Salazar
Numerade Educator
02:25

Problem 154

Identify the problem(s) in each of the residual plots shown on page 752 .

Eric Mockensturm
Eric Mockensturm
Numerade Educator
02:22

Problem 155

Dating and disclosure. Refer to the Journal of Adolescence (Apr. 2010) study of adolescents' disclosure of their dating and romantic relationships, Exercise 12.16 (p. 698). Recall that multiple regression was used to model $y=$ level of an adolescent's disclosure of a date's identity to his/her mother (measured on a 5 -point scale). The independent variables in the study were gender $\left(x_{1}=1\right.$ if female, 0 if male), age $\left(x_{2},\right.$ years $),$ dating experience $\left(x_{3},\right.$ years $),$ and level of trust in parents $\left(x_{4}, 5\right.$ -point scale $)$ The highest correlation (in absolute value) for any pair of independent variables was $r=-.16$ for gender and level of trust. Do you believe that the regression analysis will exhibit multicollinearity problems? Explain.

Lucas Finney
Lucas Finney
Numerade Educator
02:03

Problem 156

State casket sales restrictions. Some states permit only licensed firms to sell funeral goods (e.g., caskets, urns) to the consumer, while other states have no restrictions. States with casket sales restrictions are being challenged in court to lift these monopolistic restrictions. A paper in the Journal of Law and Economics (Feb. 2008 ) used multiple regression to investigate the impact of lifting casket sales restrictions on the cost of a funeral. Data collected for a sample of 1,437 funerals were used to fit the model. A simpler version of the model estimated by the researchers is $E(y)=\beta_{0}+\beta_{1} x_{1}+\beta_{2} x_{2}+\beta_{3} x_{1} x_{2},$ where $y$ is the price

(in dollars) of a direct burial, $x_{1}=\{1$ if funeral home is in a restricted state, 0 if $n o t\}$, and $x_{2}=\{1$ if price includes a basic wooden casket, 0 if no casket $\}$. The estimated equation (with standard errors in parentheses) is
$$
\hat{y}=1,432+793 x_{1}-252 x_{2}+261 x_{1} x_{2}, R^{2}=.78
$$
(70)$\quad(134) \quad(109)$
a. Calculate the predicted price of a direct burial with a basic wooden casket at a funeral home in a restricted
state.
b. The data include a direct burial funeral with a basic wooden casket at a funeral home in a restricted state that costs $\$ 2,200$. Assuming the standard deviation of the model is $\$ 50$, is this data value an outlier?
c. The data also include a direct burial funeral with a basic wooden casket at a funeral home in a restricted state that costs $\$ 2,500$. Again, assume that the standard deviation of the model is $\$ 50$. Is this data value an outlier?

Dominador Tan
Dominador Tan
Numerade Educator
01:52

Problem 157

Personality traits and job performance. Refer to the Journal of Applied Psychology (Jan. 2011) study of the determinants of task performance, Exercise 12.111 (p. 742). In addition to $x_{1}=$ conscientiousness score and $x_{2}=\{1$ if highly complex job, 0 if $\operatorname{not}\},$ the researchers also used $x_{3}=$ emotional stability score, $x_{4}=$ organizational citizenship behavior score, and $x_{5}=$ counterproductive work behavior score to model $y=$ task performance score. One of their concerns is the level of multicollinearity in the data. A matrix of correlations for all possible pairs of independent variables follows. Based on this information, do you detect a moderate or high level of multicollinearity? If so, what are your recommendations?

Hunza Gilgit
Hunza Gilgit
Numerade Educator
02:39

Problem 158

Agreeableness, gender, and wages. Refer to the Journal of Personality and Social Psychology (Feb. 2012) study of on-the-job agreeableness and wages, Exercise 12.132 (p. 753). Recall that the researchers modeled mean income, $E(y),$ as a function of both agreeableness score $\left(x_{1}\right)$ and a dummy variable for gender $\left(x_{2}=1\right.$ if male, 0 if female). The data was used to fit the model $E(y)=\beta_{0}+\beta_{1} x_{1}+\beta_{2}\left(x_{1}\right)^{2}+\beta_{3} x_{2} .$ A MINITAB nor-
mal probability plot of the residuals is displayed here. Interpret the graph. Does the random error term in the model appear to be normally distributed?

Dominador Tan
Dominador Tan
Numerade Educator
03:25

Problem 159

Personality and aggressive behavior. Psychological Bulletin (Vol. 132,2006 ) reported on a study linking personality and aggressive behavior. Four of the variables measured in the study were aggressive behavior, irritability, trait anger, and narcissism. Pairwise correlations for these four variables are given in the following table.

a. Suppose aggressive behavior is the dependent variable in a regression model and the other variables are independent variables. Is there evidence of extreme multicollinearity? Explain.
b. Suppose narcissism is the dependent variable in a regression model and the other variables are independent variables. Is there evidence of extreme multicollinearity? Explain.

Emily Himsel
Emily Himsel
Numerade Educator
00:55

Problem 160

Yield strength of steel alloy. Refer to Exercise 12.142 (p. 759) and the Modelling and Simulation in Materials Science and Engineering (Vol. 13,2005 ) study in which engineers built a regression model for the tensile yield strength $(y)$ of a new steel alloy. The engineers discovered that the independent variable nickel $\left(x_{4}\right)$ was highly correlated with the other 10 potential independent variables. Consequently, nickel was dropped from the model. Do you agree with this decision? Explain.

Victor Salazar
Victor Salazar
Numerade Educator
11:58

Problem 161

Accuracy of software effort estimates. Refer to the Journal of Empirical Software Engineering (Vol. 9,2004$)$ study of the accuracy of new software effort estimates, Exercise 12.140 (p. 758). Recall that stepwise regression was used to develop a model for the relative error in estimating effort $(y)$ as a function of company role of estimator $\left(x_{1}=1\right.$ if developer, 0 if project leader) and previous accuracy $\left(x_{8}=1\right.$ if more than $20 \%$ accurate, 0 if less than $20 \%$ accurate $)$. The stepwise regression yielded the prediction equation $\hat{y}=.12-.28 x_{1}+.27 x_{8} .$ The researcher is concerned that the sign of the estimated $\beta$ multiplied by $x_{1}$ is the opposite from what is expected. (The researcher expects a project leader to have a smaller relative error of estimation than a developer.) Give at least one reason why this phenomenon occurred.

Heather Duong
Heather Duong
Numerade Educator
06:28

Problem 162

Failure times of silicon wafer microchips. Refer to the National Semiconductor study of manufactured silicon wafer integrated circuit chips, Exercise 12.78 (p. 725). Recall that the failure times of the microchips (in hours) were determined at different solder temperatures (degrees Celsius). The data are repeated in the table in the next column.
a. Fit the straight-line model $E(y)=\beta_{0}+\beta_{1} x$ to the data, where $y=$ failure time and $x=$ solder temperature.
b. Compute the residual for a microchip manufactured at a temperature of $152^{\circ} \mathrm{C}$.
c. Plot the residuals against solder temperature $(x) .$ Do you detect a trend?
d. In Exercise $12.78 \mathrm{c},$ you determined that failure time $(y)$ and solder temperature $(x)$ were curvilinearly related. Does the residual plot, part $\mathbf{c},$ support this conclusion?

PG
Patrick Garavaglia
Numerade Educator
03:40

Problem 163

Arsenic in groundwater. Refer to the Environmental Science \& Technology (Jan. 2005) study of the reliability of a commercial kit to test for arsenic in groundwater, Exercise 12.23 (p. 700 ). Recall that you fit a first-order model for arsenic level $(y)$ as a function of latitude $\left(x_{1}\right),$ longitude $\left(x_{2}\right),$ and depth $\left(x_{3}\right)$ to the data. Conduct a residual analysis of the data. Based on the results, comment on each of the following:
a. assumption of mean error $=0$
b. assumption of constant error variance
c. outliers
d. assumption of normally distributed errors
e. multicollinearity

Robin Corrigan
Robin Corrigan
Numerade Educator
07:01

Problem 164

A rubber additive made from cashew nut shells. Refer to the Industrial \& Engineering Chemistry Research (May 2013) study of the use of cardanol as an additive for natural rubber, Exercise 12.24 (p. 700 ). You analyzed a first-order model for $y=$ grafting efficiency as a function of $x_{1}=$ initiator concentration (parts per hundred resin), $x_{2}=$ cardanol concentration (parts per hundred resin), $x_{3}=$ reaction temperature (degrees Celsius) and $x_{4}=$ reaction time (hours).
a. Suppose an engineer wants to predict the grafting efficiency of chemical run with initiator concentration set at 5 parts per hundred resin, cardanol concentration at 20 parts per hundred resin, reaction temperature at 30 degrees, and reaction time at 5 hours. Would you recommend using the prediction equation from Exercise 12.24 to obtain this prediction? Explain.
b. Examine the data in the GRAFT file and determine if there is any evidence of multicollinearity. (Note: This result is due to the design of the experiment.)
c. Conduct a complete residual analysis for the model. Do you recommend any model modifications?

Raymond Matshanda
Raymond Matshanda
Numerade Educator
01:58

Problem 165

Reality TV and cosmetic surgery. Refer to the Body Image: An International Journal of Research (Mar. 2010) study of the influence of reality TV shows on one's desire to undergo cosmetic surgery, Exercise 12.25 (p. 700 ). In Exercise $12.25,$ you fit the first-order model, $E(y)=\beta_{0}+\beta_{1} x_{1}+\beta_{2} x_{2}+\beta_{3} x_{3}+\beta_{4} x_{4},$ where $y=$
desire to have cosmetic surgery, $x_{1}$ is a dummy variable for gender, $x_{2}=$ level of self-esteem, $x_{3}=$ level of body satisfaction, and $x_{4}=$ impression of reality TV. Conduct a complete residual analysis for the model. Do you detect any violations of the assumptions?

Clarissa Noh
Clarissa Noh
Numerade Educator
03:29

Problem 166

Cooling method for gas turbines. Refer to the Journal of Engineering for Gas Turbines and Power (Jan. 2005) study of a high-pressure inlet fogging method for a gas turbine engine, presented in Exercise 12.27 (p. 701). Now consider the interaction model $E(y)=\beta_{0}+\beta_{1} x_{1}+\beta_{2} x_{2}+\beta_{3} x_{1} x_{2}$
for heat rate $(y)$ of a gas turbine as a function of cycle speed $\left(x_{1}\right)$ and cycle pressure ratio $\left(x_{2}\right)$. Use the data to conduct a complete residual analysis of the model. Do you recommend making modifications to the model?

Lucas Finney
Lucas Finney
Numerade Educator
07:16

Problem 167

Emotional intelligence and team performance. Refer to the Engineering Project Organizational Journal (Vol. 3, 2013) study of the relationship between emotional intelligence of individual team members and their performance during an engineering project, Exercise 12.39 (p. 708 ). Using data on $n=23$ teams (saved in the EINTELL file), you fit a first-order model for mean project score $(y)$ as a function of range of interpersonal scores $\left(x_{1}\right),$ range of stress management scores $\left(x_{2}\right),$ and range of mood scores $\left(x_{3}\right) .$ Do you detect any signs of multicollinearity in the data?

Jerelyn Nevil
Jerelyn Nevil
Numerade Educator
04:22

Problem 168

Bubble behavior in subcooled flow boiling. Refer to the Heat Transfer Engineering (Vol. 34,2013 ) study of bubble behavior in subcooled flow boiling, Exercise 12.59 (p. 716). You fit an interaction model for bubble density
(y) as a function of $x_{1}=$ mass flux and $x_{2}=$ heat flux. Conduct a complete residual analysis for the model. Do you recommend any model modifications?

Khoobchandra Agrawal
Khoobchandra Agrawal
Numerade Educator
01:43

Problem 169

Write a model relating $E(y)$ to one qualitative independent variable that is at four levels. Define all the terms in your model.

Harsh Gadhiya
Harsh Gadhiya
Numerade Educator
01:59

Problem 170

Explain why stepwise regression is used. What is its value in the model-building process?

Shiksha Dutta
Shiksha Dutta
Numerade Educator
05:55

Problem 171

It is desired to relate $E(y)$ to a quantitative variable $x_{1}$ and a qualitative variable at three levels.
a. Write a first-order model.
b. Write a model that will graph as three different second-order curves, one for each level of the qualitative variable.

Mihir Nayar
Mihir Nayar
Numerade Educator
View

Problem 172

a. Write a first-order model relating $E(y)$ to two quantitative independent variables $x_{1}$ and $x_{2}$.
b. Write a complete second-order model.

Shu Naito
Shu Naito
Numerade Educator
02:35

Problem 173

Suppose you fit the model
$$
y=\beta_{0}+\beta_{1} x_{1}+\beta_{2} x_{1}^{2}+\beta_{3} x_{2}+\beta_{4} x_{1} x_{2}+\varepsilon
$$
to $n=25$ data points and find that
$$
\begin{array}{r}
\hat{\beta}_{0}=1.24 \quad \hat{\beta}_{1}=-2.43 \quad \hat{\beta}_{2}=.03 \quad \hat{\beta}_{3}=.65 \quad \hat{\beta}_{4}=1.75 \\
s_{\hat{\beta}_{1}}=1.15, \quad s_{\hat{\beta}_{2}}=.18, \quad s_{\hat{\beta}_{3}}=.30, \quad s_{\hat{\beta}_{4}}=1.49
\end{array}
$$
$\mathrm{SSE}=.47, \quad R^{2}=.85$
a. Is there sufficient evidence to conclude that at least one of the parameters $\beta_{1}, \beta_{2}, \beta_{3},$ and $\beta_{4}$ is nonzero? Test, using $\alpha=.05$.
b. Test $H_{0}: \beta_{1}=0$ against $H_{\mathrm{a}}: \beta_{1}<0 .$ Use $\alpha=.05$.
c. Test $H_{0}: \beta_{2}=0$ against $H_{\mathrm{a}}: \beta_{2}>0 .$ Use $\alpha=.05$.
d. Test $H_{0}: \beta_{3}=0$ against $H_{\mathrm{a}}: \beta_{3} \neq 0 .$ Use $\alpha=.05$.

a. What is the least squares prediction equation?
b. Find $R^{2}$ and interpret its value.
c. Is there sufficient evidence to indicate that the model is useful in predicting $y ?$ Conduct an $F$ -test, using $\alpha=.05 .$
d. Test the null hypothesis $H_{0}: \beta_{1}=0$ against the alternative hypothesis $H_{\mathrm{a}}: \beta_{1} \neq 0 .$ Use $\alpha=.05 .$ Draw the appropriate conclusions.
e. Find the standard deviation of the regression model and interpret it.

Nick Johnson
Nick Johnson
Numerade Educator
02:39

Problem 174

Suppose you used MINITAB to fit the model
$$
y=\beta_{0}+\beta_{1} x_{1}+\beta_{2} x_{2}+\varepsilon
$$
to $n=15$ data points and you obtained the printout below.

a. What is the least squares prediction equation?
b. Find $R^{2}$ and interpret its value.
c. Is there sufficient evidence to indicate that the model is useful in predicting $y$ ? Conduct an $F$ -test, using $\alpha=.05 .$
d. Test the null hypothesis $H_{0}: \beta_{1}=0$ against the alternative hypothesis $H_{\mathrm{a}}: \beta_{1} \neq 0 .$ Use $\alpha=.05 .$ Draw the appropriate conclusions.
e. Find the standard deviation of the regression model and interpret it.

Dominador Tan
Dominador Tan
Numerade Educator
01:43

Problem 175

Suppose you have developed a regression model to explain the relationship between $y$ and $x_{1}, x_{2},$ and $x_{3}$. The ranges of the variables you observed were as follows: $\quad 50 \leq y \leq 100,10 \leq x_{1} \leq 75, .25 \leq x_{2} \leq .75,$
and $2,000 \leq x_{3} \leq 2,500$. Will the error of prediction be smaller when you use the least squares equation to predict $y$ when $x_{1}=40, x_{2}=.5,$ and $x_{3}=2,300,$ or when $x_{1}=80, x_{2}=.9,$ and $x_{3}=1,850 ?$ Why?

Dominador Tan
Dominador Tan
Numerade Educator
View

Problem 176

The first-order model $E(y)=\beta_{0}+\beta_{1} x_{1}$ was fit to $n=19$ data points. A plot of the residuals of the model is shown. Is the need for a quadratic term in the model evident from the plot? Explain.

Shu Naito
Shu Naito
Numerade Educator
02:35

Problem 177

To model the relationship between $y$ (a dependent variable) and $x$ (an independent variable), a researcher has taken one measurement of $y$ at each of three different $x$ values. Drawing on his mathematical expertise, the researcher realizes that he can fit the second-order model
$$
E(y)=\beta_{0}+\beta_{1} x+\beta_{2} x^{2}
$$
and it will pass exactly through all three points, yielding $\mathrm{SSE}=0 .$ The researcher, delighted with the "excellent" fit of the model, eagerly sets out to use it to make inferences. What problems will he encounter in attempting to make inferences?

Nick Johnson
Nick Johnson
Numerade Educator
02:35

Problem 178

Suppose you fit the regression model
$$
E(y)=\beta_{0}+\beta_{1} x_{1}+\beta_{2} x_{2}+\beta_{3} x_{2}^{2}+\beta_{4} x_{1} x_{2}+\beta_{5} x_{1} x_{2}^{2}
$$
to $n=30$ data points and wish to test the null hypothesis $H_{0}: \beta_{4}=\beta_{5}=0$
a. State the alternative hypothesis.
b. Explain in detail how to compute the $F$ -statistic needed to test the null hypothesis.
c. What are the numerator and denominator degrees of freedom associated with the $F$ -statistic in part $\mathbf{b}$ ?
d. Give the rejection region for the test if $\alpha=.05$.

Nick Johnson
Nick Johnson
Numerade Educator
04:26

Problem 179

Global warming and foreign investments. Scientists believe that a major cause of global warming is higher levels of carbon dioxide $\left(\mathrm{CO}_{2}\right)$ in the atmosphere. In the Journal of World-Systems Research (Summer 2003), sociologists examined the impact of a dependence on foreign investment on $\mathrm{CO}_{2}$ emissions in $n=66$ developing countries. In particular, the researchers modeled the level of $\mathrm{CO}_{2}$ emissions in a particular year on the basis of foreign investments made 16 years earlier and several other independent variables. The variables and the model results are listed in the following table.
a. Interpret the value of $R^{2}$.
b. Use the value of $R^{2}$ to test the null hypothesis, $H_{0}: \beta_{1}=\beta_{2}=\cdots=\beta_{7}=0$ at $\alpha=.01 .$ Give the
appropriate conclusion.
c. Do you advise conducting $t$ -tests on each of the independent variables to test the overall adequacy of the model? Explain.
Correlations for Exercise 12.179
d. What null hypothesis would you test to determine whether the number of foreign investments made 16 years earlier is a statistically useful predictor of $\mathrm{CO}_{2}$ emissions the current year?
e. Conduct the test mentioned in part $\mathbf{d}$ at $\alpha=.05$. Give the appropriate conclusion.
f. A matrix giving the correlation $(r)$ for each pair of independent variables is shown in the table above. Identify the independent variables that are highly correlated. What problems may result from including these highly correlated variables in the regression model?

James Kiss
James Kiss
Numerade Educator
01:09

Problem 180

Students' ability in science. An article published in the American Educational Research Journal (Fall 1998) used multiple regression to model the students' perceptions of their ability in science classes. The sample consisted of 165 Grade $5-8$ students. The dependent variable of interest, the student's perception of his or her ability (y), was measured on a four-point scale (where $1=$ little or no ability and $4=$ high ability). Two types of independent variables were included in the model, control variables and performance behavior variables. The control variables are: prior science attitude (measured on a 4 -point scale), score on standardized science test, gender, and classroom $(1,3,4,5,$ or 6$) .$ The performance behavior variables (all measured on a numerical scale between 0 and 1 ) are: active-leading behavior, passive-assisting behavior, and active-manipulating behavior.
a. Identify the independent variables as quantitative or qualitative.
b. Individual $\beta$ -tests on the independent variables all had $p$ -values greater than .10 except for prior science attitute, gender, and active-leading behavior. Which variables appear to contribute to the prediction of a student's perception of his or her ability in science?
c. The estimated $\beta$ -value for the active-leading behavior variable is .88 with a standard error of .34. Use this information to construct a $95 \%$ confidence interval for this $\beta$. Interpret the interval.
d. The following statistics for evaluating the overall predictive power of the model were reported:
$R^{2}=.48, F=12.84, p<.001 .$ Interpret the results.
e. Hypothesize the equation of the first-order main effects model for $E(y)$
f. The researchers also considered a model that included all possible interactions between the control variables and the performance behavior variables. Write the equation for this model for $E(y)$.

Dominador Tan
Dominador Tan
Numerade Educator
03:08

Problem 181

Distress in EMS workers. The Journal of Consulting and Clinical Psychology (June 1995) reported on a study of emergency service (EMS) rescue workers who responded to the I-880 freeway collapse during a San Francisco earthquake. The goal of the study was to identify the predictors of symptomatic distress in the workers. One of the distress variables studied was the Global Symptom Index (GSI). Several models for GSI (y) based on the following independent variables were considered:
$\begin{aligned} x_{1}=& \text { Critical Incident Exposure scale (CIE) } \\ x_{2}=& \text { Hogan Personality Inventory-Adjustment } \\ & \text { scale (HPI-A) } \\ x_{3}=& \text { Years of experience (EXP) } \\ x_{4}=& \text { Locus of Control scale (LOC) } \\ x_{5}=& \text { Social Support scale (SS) } \\ x_{6}=& \text { Dissociative Experiences scale (DES) } \\ x_{7}=& \text { Peritraumatic Dissociation Experiences } \\ & \text { Questionnaire, self-report (PDEQ-SR) } \end{aligned}$
a. Write a first-order model for $E(y)$ as a function of the first five independent variables, $x_{1}-x_{5}$.
b. The model from part a, fitted to data collected on $n=147$ EMS workers, yielded the following results:
$R^{2}=.469, F=34.47, p$ -value $<.001 .$ Interpret these
results.
c. Write a first-order model for $E(y)$ as a function of all seven independent variables, $x_{1}-x_{7}$
d. The model from part $\mathbf{c}$ yielded $R^{2}=.603 .$ Interpret this result.
e. The $t$ -tests for testing the DES and PDEQ-SR variables both yielded a $p$ -value of .001 . Interpret this result.

Nick Johnson
Nick Johnson
Numerade Educator
02:57

Problem 182

Listen-and-look study. Where do you look when you are listening to someone speak? Researchers have discovered that listeners tend to gaze at the eyes or mouth of the speaker. In a study published in Perception \& Psychophysics (Aug. 1998), subjects watched a videotape of a speaker giving a series of short monologues at a social gathering (e.g., a party). The level of background noise (multilingual voices and music) was varied during the listening sessions. Each subject wore a pair of clear plastic goggles on which an infrared corneal detection system was mounted, enabling the researchers to monitor the subject's eye movements. One response variable of interest was the proportion $y$ of times the subject's eyes fixated on the speaker's mouth.
a. The researcherswanted toestimate $E(y)$ forfourdifferent noise levels: none, low, medium, and high. Hypothesize a model that will allow the researchers to obtain these estimates.
b. Interpret the $\beta$ 's in the model you hypothesized in part a
c. Explain how to test the hypothesis of no differences in the mean proportions of mouth fixations for the four background noise levels.

Lucas Finney
Lucas Finney
Numerade Educator
03:00

Problem 183

Defects in nuclear missile housing parts. The technique of multivariable testing was discussed in The Journal of the Reliability Analysis Center (First Quarter, 2004). Multivariable testing was shown to improve the quality of carbon-foam rings used in nuclear missile housings. The rings are produced via a casting process that involves mixing ingredients, oven curing, and carving the finished part. One type of defect analyzed was the number $y$ of black streaks in the manufactured ring. Two variables found to affect the number of defects were turntable speed (revolutions per minute) $x_{1}$ and cutting blade position (inches from center) $x_{2}$.
a. The researchers discovered "an interaction between blade position and turntable speed." Hypothesize a regression model for $E(y)$ that incorporates this interaction.
b. Interpret what it means practically to say that "blade position and turntable speed interact."
c. The researchers reported a positive linear relationship between number of defects $(y)$ and turntable speed $\left(x_{1}\right)$ but found that the slope of the relationship was much steeper for lower values of cutting blade position $\left(x_{2}\right)$. What does this imply about the interaction term in the model you hypothesized in part a? Explain.

Khoobchandra Agrawal
Khoobchandra Agrawal
Numerade Educator
03:26

Problem 184

Violent behavior in children. Refer to the Development Psychology (Mar. 2003) study of the behavior of elementary school children, presented in Exercise 8.149 (p. 454 ). The researchers used a quadratic equation to model the level $(y)$ of aggressive fantasies experienced by a child as a function of the child's age $(x)$. [Note: Level $y$ was measured as an average of responses to six questions (e.g. "Do you sometimes have daydreams about hitting or hurting someone you don't like?"). Responses were measured on a scale ranging from 1 (never) to 5 (always).]
a. Write the equation of the hypothesized model for $E(y)$
b. Research psychologists theorized that a child's aggressive fantasies increase with age but at a slower rate of acceleration in older children. Sketch the curve hypothesized by the researchers.
c. Set up $H_{0}$ and $H_{\mathrm{a}}$ for testing the researchers? theory.
d. The model was fitted to data collected for over 11,000 elementary school children, with the following results: $\hat{y}=1.926+.097 x-.003 x^{2},$ standard error of $\hat{\beta}_{2}=.001 .$ Compute the test statistic for the test you set up in part $\mathbf{c}$.
e. Use the result you found in part $\mathbf{d}$ to make the appropriate conclusion. Take $\alpha=.05$.

Sheryl Ezze
Sheryl Ezze
Numerade Educator
14:37

Problem 185

Improving SAT scores. Refer to the Chance (Winter
2001) study of students who paid a private tutor (or coach) to help them improve their SAT scores, presented in Exercise 2.197 (p. 136). Multiple regression was used to estimate the effect of coaching on SAT-Mathematics scores. Data on 3,492 students $(573$ of whom were coached $)$ were used to fit the model $E(y)=\beta_{0}+\beta_{1} x_{1}+\beta_{2} x_{2},$
where $y=$ SAT-Math score, $x_{1}=$ score on PSAT, and $x_{2}=\{1$ if student was coached, 0 if not $\}$
a. The fitted model had an adjusted $R^{2}$ value of .76 . Interpret this result.
b. The estimate of $\beta_{2}$ in the model was $19,$ with a standard error of $3 .$ Use this information to form a $95 \%$ confidence interval for $\beta_{2}$. Interpret the interval.
c. On the basis of the interval you found in part $\mathbf{b},$ what can you say about the effect of coaching on SAT-Math scores?
d. As an alternative model, the researcher added several "control" variables, including dummy variables for student ethnicity $\left(x_{3}, x_{4},\right.$ and $\left.x_{5}\right),$ a socioeconomic status index variable $\left(x_{6}\right),$ two variables that measured high school performance $\left(x_{7}\right.$ and $\left.x_{8}\right),$ the number of math courses taken in high school $\left(x_{9}\right),$ and the overall GPA for the math courses $\left(x_{10}\right) .$ Write the hypothesized equation for $E(y)$ for the alternative model.
e. Give the null hypothesis for a nested-model $F$ -test comparing the initial and alternative models.
f. The nested-model $F$ -test from part e was statistically significant at $\alpha=.05 .$ Interpret this result practically.
g. The alternative model from part d resulted in $R_{a}^{2}=.79, \hat{\beta}_{2}=14,$ and $s_{\hat{\beta}_{2}}=3 .$ Interpret the value
of $R_{a}^{2}$
h. Refer to part $\mathbf{g}$. Find and interpret a $95 \%$ confidence interval for $\beta_{2}$.
i. The researcher concluded that "the estimated effect of SAT coaching decreases from the baseline model when control variables are added to the model." Do you agree? Justify your answer.
j. As a modification to the model of part $\mathbf{d}$, the researcher added all possible interactions between the coaching variable $\left(x_{2}\right)$ and the other independent variables in the model. Write the equation for $E(y)$ for this modified model.
k. Give the null hypothesis for comparing the models from parts $\mathbf{d}$ and $\mathbf{j}$. How would you perform this test?

Heather Duong
Heather Duong
Numerade Educator
06:04

Problem 186

Factors identifying urban counties. The Professional Geographer (Feb. 2000) published a study of urban and rural counties in the western United States. The researchers used six independent variables-total county population $\left(x_{1}\right),$ population density $\left(x_{2}\right),$ population concentration $\left(x_{3}\right),$ population growth $\left(x_{4}\right),$ proportion of county land in farms $\left(x_{5}\right),$ and five-year change in agricultural land base $\left(x_{6}\right)-$ to model the urban/ rural rating $(y)$ of a county on a scale of 1 (most rural) to 10 (most urban). Prior to running the multiple-regression analysis, the researchers were concerned about possible multicollinearity in the data. Following is a MINITAB printout of correlations between all pairs of the independent variables:
a. On the basis of the correlation printout, is there any evidence of extreme multicollinearity?
b. The first-order model with all six independent variables was fit to the data. The multiple-regression results are shown in the accompanying MINITAB printout. On the basis of the reported tests, is there any evidence of extreme multicollinearity?

Jon Southam
Jon Southam
Numerade Educator
03:39

Problem 187

Growth of Japanese beetles. In the Journal of Insect Behavior (Nov. 2001), biologists at Eastern Illinois University published the results of their study on Japanese beetles. The biologists collected beetles over a period of $n=13$ summer days in a soybean field. For one portion of the study, the biologists modeled $y,$ the average size (in millimeters) of female beetles as a function of the average daily temperature $x_{1}$ (degrees) and Julian date $x_{2}$
a. Write a first-order model for $E(y)$ as a function of $x_{1}$ and $x_{2}$.
b. The model was fit to the data, with the results shown in the accompanying table. Interpret the estimate of $\beta_{1}$.
c. Conduct a test to determine whether the average size of female Japanese beetles decreases linearly as the temperature increases. Use $\alpha=.05$.

Carson Merrill
Carson Merrill
Numerade Educator
01:10

Problem 188

Deferred tax allowance study. A study was conducted to identify accounting choice variables that influence a manager's decision to change the level of the deferred tax asset allowance at a firm (The Engineering Economist, Jan./Feb. 2004). Data were collected on a sample of 329 firms that reported deferred tax assets. The dependent variable of interest (DTVA) is measured as the change in the deferred tax asset valuation allowance divided by the deferred tax asset. The independent variables used as predictors of DTVA are as follows:

A first-order model was fitted to the data, with the following results $(p$ -values are in parentheses):
$\hat{y}=.044+.006 x_{1}-.035 x_{2}-.001 x_{3}+.296 x_{4},+.010 x_{5}, R_{a}^{2}=.280$
$\begin{array}{lllll}(.070)(.228) & (.157) & (.678) & (.001) & (.869)\end{array}$
a. Interpret the estimate of the $\beta$ -coefficient for $x_{4}$.
b. The "Big Bath" theory proposed by the researchers states that the mean DTVA for firms with negative earnings and earnings lower than last year will exceed the mean DTVA of other firms. Is there evidence to support this theory? Test, using $\alpha=.05$.
c. Interpret the value of $R_{a}^{2}$.

Dominador Tan
Dominador Tan
Numerade Educator
03:51

Problem 189

Snow geese feeding trial. Refer to the Journal of Applied Ecology (Vol. 32, 1995) study of the feeding habits of baby snow geese, presented in Exercise 11.150 (p. 676 ). Data on gosling weight change, digestion efficiency, acid-detergent fiber (all measured as percentages), and diet (plants or duck chow) for 42 feeding trials are saved in the SNOW file. Selected observations are shown in the table (p. 789 ). The botanists were interested in predicting weight change $(y)$ as a function of the other variables. Consider the first-order model $E(y)=\beta_{0}+\beta_{1} x_{1}+B_{2} x_{2},$ where $x_{1}$ is digestion efficiency and $x_{2}$ is acid-detergent fiber.
a. Find the least squares prediction equation for weight change $y$.
b. Interpret the $\beta$ -estimates in the equation you found in part a.
c. Conduct a test to determine whether digestion efficiency, $x_{1},$ is a useful linear predictor of weight change. Use $\alpha=.01$
d. Form a $99 \%$ confidence interval for $\beta_{2} .$ Interpret the result.
e. Find and interpret $R^{2}$ and $R_{a}^{2}$. Which statistic is the preferred measure of model fit? Explain.
f. Is the overall model statistically useful in predicting weight change? Test, using $\alpha=.05$.
g. Write a first-order model relating gosling weight change $(y)$ to digestion efficiency $\left(x_{1}\right)$ and diet (plants or duck chow) that allows for different slopes for each diet.
h. Fit the model you wrote in part $g$ to the data saved in the SNOW file. Give the least squares prediction equation.
i. Refer to part $\mathbf{g}$. Find the estimated slope of the line for goslings fed a diet of plants. Interpret its value.
j. Refer to part $\mathbf{g}$. Find the estimated slope of the line for goslings fed a diet of duck chow. Interpret its value.
k. Refer to part $\mathbf{g}$. Conduct a test to determine whether the slopes associated with the two diets are significantly different. Use $\alpha=.05$.

Sheryl Ezze
Sheryl Ezze
Numerade Educator
09:40

Problem 190

Optimizing semiconductor material processing. Fluorocarbon plasmas are used in the production of semiconductor materials. In the Journal of Applied Physics (Dec. 1,2000 ), electrical engineers at Nagoya University (Japan) studied the kinetics of fluorocarbon plasmas in order to optimize material processing. In one portion of the study, the surface production rate of fluorocarbon radicals emitted from the production process was measured at various points in time (in milliseconds) after the radio frequency power was turned off. The data are given in the table below. Consider a model relating surface production rate $(y)$ to time $(x)$
a. Graph the data in a scatterplot. What trend do you observe?
b. Fit a quadratic model to the data. Give the least squares prediction equation.
c. Is there sufficient evidence of upward curvature in the relationship between surface production rate and time after turnoff? Use $\alpha=.05$.

Ameer Said
Ameer Said
Numerade Educator
08:01

Problem 191

Women in top management. The Journal of Organizational Culture, Communications and Conflict (July 2007) published a study on women in upper management positions at U.S. firms. Monthly data $(n=252$ months) were collected for several variables in an attempt to model the number of females in managerial positions $(y)$. The independent variables included the number of females with a college degree $\left(x_{1}\right),$ the number of female high school graduates with no college degree $\left(x_{2}\right),$ the number of males in managerial positions $\left(x_{3}\right),$ the number of males with a college degree $\left(x_{4}\right),$ and the number of male high school graduates with no college degree $\left(x_{5}\right)$. Determine which of the correlations reported in parts a-d results in a potential multicollinearity problem for the regression analysis.
a. The correlation relating number of females in managerial positions and number of females with a college degree: $r=.983$.
b. The correlation relating number of females in managerial positions and number of female high school graduates with no college degree: $r=.074$.
c. The correlation relating number of males in managerial positions and number of males with a college degree: $r=.722$.
d. The correlation relating number of males in managerial positions and number of male high school graduates with no college degree: $r=.528$.

SS
Sarvesh Somasundaram
Numerade Educator
05:48

Problem 192

Comparing mosquito repellents. Which insect repellents protect best against mosquitoes? Periodically, consumer groups conduct tests to compare insect repellents (e.g., Consumer Reports, June 2000). Consider a similar study of 14 popular mosquito repellents. Each product was classified as either an aerosol spray or as a lotion. The cost of the product (in dollars) was divided by the amount of the repellent needed to cover exposed areas of the skin (about $1 / 3$ ounce) to obtain a cost-per-use value. Effectiveness was measured as the maximum number of hours of protection (in half-hour increments) provided when human testers exposed their arms to 200 mosquitoes. The simulated data are listed in the table on p. 790 .
a. Suppose you want to use repellent type to model the cost per use $(y)$. Create the appropriate number of dummy variables for repellent type, and write the model.
b. Fit the model you wrote in part a to the data.
c. Give the null hypothesis for testing whether repellent type is a useful predictor of cost per use $(y)$
d. Conduct the test suggested in part $\mathbf{c},$ and give the appropriate conclusion. Use $\alpha=.10$.
e. Repeat parts a-d if the dependent variable is maximum number of hours of protection $(y)$.

Sheryl Ezze
Sheryl Ezze
Numerade Educator
View

Problem 193

Rating funny cartoons. Newspaper cartoons, although designed to be funny, often invoke hostility, pain, or aggression in readers, especially when those cartoons depict violence. A study was undertaken to determine how violence in cartoons is related to aggression or pain (Motivation and Emotion, Vol. 10,1986 ). A group of volunteers (psychology students) rated each of 32 violent newspaper cartoons (16 "Herman" and 16 "Far Side" cartoons) on three dimensions:
$\begin{aligned} y=& \text { Funniness }(0=\text { not funny }, \ldots, 9=\text { very funny }) \\ x_{1}=& \text { Pain }(0=\text { none }, \ldots, 9=\text { a very great deal }) \\ x_{2}=& \text { Aggression/hostility }(0=\text { none }, \ldots, 9=\text { a very }\\ & \text { great deal) } \end{aligned}$
The ratings of the students on each dimension were averaged, and the resulting $n=32$ observations were subjected to a multiple-regression analysis. On the basis of the underlying theory (called the inverted-U theory) that the funniness of a joke will increase at low levels of aggression or pain, level off, and then decrease at high levels of aggression or pain, the following quadratic models were proposed:
Model $1: E(y)=\beta_{0}+\beta_{1} x_{1}+\beta_{2} x_{1}^{2}, R^{2}=.099, F=1.60$
Model $2: E(y)=\beta_{0}+\beta_{1} x_{2}+\beta_{2} x_{2}^{2}, R^{2}=.100, F=1.61$
a. According to the theory, what is the expected sign of $\beta_{2}$ in each model?
b. Is there sufficient evidence to indicate that the quadratic model relating pain to funniness rating is useful? Test at $\alpha=.05$.
c. Is there sufficient evidence to indicate that the quadratic model relating aggression/hostility to funniness rating is useful? Test at $\alpha=.05$.

Victor Salazar
Victor Salazar
Numerade Educator
05:16

Problem 194

"Sun safety" study. Excessive exposure to solar radiation is known to increase the risk of developing skin cancer, yet many people do not practice "sun safety." A group of University of Arizona researchers examined the feasibility of educating preschool (four- to five-year-old) children about sun safety (American Journal of Public Health, July 1995). A sample of 122 preschool children was divided into two groups: the control group and the intervention group. Children in the intervention group received a Be Sun Safe curriculum in preschool, while the control group did not. All children were tested for their knowledge, comprehension, and application of sun safety at two points in time: prior to the sun safety curriculum (pretest, $x_{1}$ ) and seven weeks following the curriculum (posttest, $y$ )
a. Write a first-order model for mean posttest score $E(y)$ as a function of pretest score $x_{1}$ and group. Assume that no interaction exists between pretest score and group.
b. For the model you wrote in part a, show that the slope of the line relating posttest score to pretest score is the same for both groups of children.
c. Repeat part $\mathbf{a}$, but assume that pretest score and group interact.
d. For the model of part $\mathbf{c}$, show that the slope of the line relating posttest score to pretest score differs for the two groups of children.
e. Assuming that interaction exists, give the reduced model for testing whether the mean posttest scores differ for the intervention and control groups.
f. With sun safety knowledge as the dependent variable, the test presented in part e was carried out and resulted in a $p$ -value of .03. Interpret this result.
g. With sun safety comprehension as the dependent variable, the test presented in part e was carried out and resulted in a $p$ -value of .033 . Interpret this result.
h. With sun safety application as the dependent variable, the test presented in part e was carried out and resulted in a $p$ -value of .322. Interpret this result.

Lucas Finney
Lucas Finney
Numerade Educator
01:52

Problem 195

Revenues of popular movies. The Internet Movie Database (www.imdb.com) monitors the gross revenues of all major motion pictures. The table on p. 791 gives both the domestic (United States and Canada) and international gross revenues for a sample of 22 popular movies.
a. Write a first-order model for international gross revenues $y$ as a function of domestic gross revenues $x$.
b. Write a second-order model for international gross revenues $y$ as a function of domestic gross revenues $x$.
c. Construct a scatterplot of these data. Which of the models appears to be a better choice for explaining variation in international gross revenues?
d. Fit the model of part $\mathbf{b}$ to the data and investigate its usefulness. Is there evidence of a curvilinear relationship between international and domestic gross revenues? Test, using $\alpha=.05$.
e. On the basis of your analysis in part $\mathbf{d}$, which of the two models better explains the variation in international gross revenues?

Sheryl Ezze
Sheryl Ezze
Numerade Educator
01:57

Problem 196

Habitats of grizzly bears. Do grizzly bears segregate on the basis of sex? One hypothesis is that female grizzlies avoid male-occupied habitats because of competition for food and cannibalism. A competing theory is that females do not avoid males, but simply have different habitats available to them. These hypotheses were investigated in the Journal of Wildlife Management (July 1995). Grizzly bears were trapped, fitted with a radio collar, and released in the Highwood trapping zone (HTZ) in Alberta, Canada. The percentage of time, $y,$ each bear used the HTZ as a habitat over a four-year period was recorded. The researchers modeled $E(y)$ as a function of reproductive class at five levels: estrous adult females, adult females with offspring, independent subadult females,

adult males, and independent subadult males. One goal was to compare the mean percentage use of HTZ for the five classes of grizzly bears.
a. Write a model for $E(y)$ that will enable the researchers to carry out the comparison.
b. The sample sizes and sample means for the five classes are shown in the table below. Use this information to find estimates of the $\beta$ 's in the model you wrote in part a
c. Give the null hypothesis for a test to determine whether the mean percentage of use of HTZ differs among the grizzly bear classes.
d. The $p$ -value for the test mentioned in part $\mathbf{c}$ was reported as .15. Interpret this result.

Sheryl Ezze
Sheryl Ezze
Numerade Educator
02:03

Problem 197

Sale prices of apartments. A Minneapolis, Minnesota, real-estate appraiser used regression analysis to explore the relationship between the sale prices of apartment buildings sold in Minneapolis and various characteristics of the properties. Twenty-five apartment buildings were randomly sampled from all apartment buildings that were sold during a recent year. The table on p. 792 lists the data collected by the appraiser. [Note:
The physical condition of each apartment building is coded $\mathrm{E}($ excellent $), \mathrm{G}$ (good), or $\mathrm{F}$ (fair).]
a. Write a model that describes the relationship between sale price and number of apartment units as three parallel lines, one for each level of physical condition. Be sure to specify the dummy-variable coding scheme you use.
b. Plot $y$ against $x_{1}$ (number of apartment units) for all buildings in excellent condition. On the same graph, plot $y$ against $x_{1}$ for all buildings in good condition. Do this again for all buildings in fair condition. Does it appear that the model you specified in part a is appropriate? Explain.
c. Fit the model from part a to the data. Report the least squares prediction equation for each of the three building condition levels.
d. Plot the three prediction equations of part $\mathbf{c}$ on a scatterplot of the data.
e. Do the data provide sufficient evidence to conclude that the relationship between sale price and number of units varies with the physical condition of the apartments? Test, using $\alpha=.05$.
f. Check the data set for multicollinearity. How does your result affect your choice of independent variables to use in a model for sale price?
g. Consider the first-order model $E(y)=\beta_{0}+\beta_{1} x_{1}$ $+\cdots+\beta_{5} x_{5} .$ Conduct a complete residual analysis for the model to check the assumptions on $\varepsilon$.

Dominador Tan
Dominador Tan
Numerade Educator
View

Problem 198

RNA analysis of wheat genes. Engineers from the Department of Crop and Soil Sciences at Washington State University used regression to estimate the number of copies of a gene transcript in an aliquot of ribonucleic acid (RNA) extracted from a wheat plant (Electronic Journal of Biotechnology, Apr. 15,2004 ). The proportion $\left(x_{1}\right)$ of RNA extracted from a wheat plant exposed to the cold was varied, and the transcript copy number $(y,$ in thousands) was measured for each of two cloned genes:
Mn superoxide dismutase (MnSOD) and phospholipase D (PLD). The data are listed in the table on p. 792 .
a. Write a first-order model for number of copies $(y)$ as a function of proportion $\left(x_{1}\right)$ of RNA extracted and gene type (MnSOD or PLD). Assume that proportion of RNA and gene type interact to affect $y$.
b. Fit the model you wrote in part a to the data. Give the least squares prediction equation for $y$.
c. Conduct a test to determine whether, in fact, proportion of RNA and gene type interact. Test using $\alpha=.01$.
d. Use the results from part b to estimate the rate of increase of number of copies $(y)$ with proportion $\left(x_{1}\right)$ of RNA extracted for the MnSOD gene type.
e. Repeat part $\mathbf{d}$ for the PLD gene type.

Sana Riaz
Sana Riaz
Numerade Educator
09:18

Problem 199

Entry-level job preferences. Benefits Quarterly periodical publishes a study of entry-level job preferences. A number of independent variables are used to model the job preferences (measured on a 10 -point scale) of business school graduates. Suppose stepwise regression is used to build a model for job preference score $(y)$ as a function of the following independent variables:
$x_{1}=\left\{\begin{array}{ll}1 & \text { if flextime position } \\ 0 & \text { if not }\end{array}\right.$
$x_{2}=\left\{\begin{array}{ll}1 & \text { if day care support required } \\ 0 & \text { if not }\end{array}\right.$
$x_{3}=\left\{\begin{array}{ll}1 & \text { if spousal transfer support required } \\ 0 & \text { if not }\end{array}\right.$
$x_{4}=$ Number of sick days allowed
$x_{5}=\left\{\begin{array}{ll}1 & \text { if applicant married } \\ 0 & \text { if not }\end{array}\right.$
$x_{6}=$ Number of children of applicant
$x_{7}=\left\{\begin{array}{ll}1 & \text { if male applicant } \\ 0 & \text { if female applicant }\end{array}\right.$
a. How many models are fitted to the data in step $1 ?$ Give the general form of these models.
b. How many models are fitted to the data in step $2 ?$ Give the general form of these models.
c. How many models are fitted to the data in step $3 ?$ Give the general form of these models.
d. Explain how the procedure determines when to stop adding independent variables to the model.
e. Describe two major drawbacks to using the final stepwise model as the "best" model for job preference score $(y)$

Heather Duong
Heather Duong
Numerade Educator
View

Problem 200

Characteristics of sea-ice melt ponds. Surface albedo is defined as the ratio of solar energy directed upward from a surface over energy incident upon the surface. Surface albedo is a critical climatological parameter of sea ice. The National Snow and Ice Data Center (NSIDC) collects data on the albedo, depth, and physical characteristics of ice-melt ponds in the Canadian Arctic. Data on 504 ice-melt ponds located in the Barrow Strait in the Canadian Arctic are saved in the PONDICE file. Environmental engineers want to examine the relationship between the broadband surface albedo level $y$ of the ice and the pond depth $x$ (in meters).
a. Construct a scatterplot of the data. On the basis of the scatterplot, hypothesize a model for $E(y)$ as a function of $x$
b. Fit the model you hypothesized in part a to the data. Give the least squares prediction equation.
c. Conduct a test of the overall adequacy of the model. Use $\alpha=.01$.
d. Conduct tests (at $\alpha=.01$ ) on any important $\beta$ parameters in the model.
e. Find and interpret the values of adjusted $R^{2}$ and $s$.
f. Do you detect any outliers in the data? Explain.

Rashmi Sinha
Rashmi Sinha
Numerade Educator
04:01

Problem 201

Child abuse report. Licensed therapists are mandated by law to report child abuse by their clients. This law requires the therapist to breach confidentiality and possibly lose the client's trust. A national survey of licensed psychotherapists was conducted to investigate clients' reactions to legally mandated child-abuse reports ( American Journal of Orthopsychiatry, Jan. 1997). The sample consisted of 303 therapists who had filed a child-abuse report against at least one of their clients. The researchers were interested in finding the best predictors of a client's reaction $(y)$ to the report, where $y$ is measured on a 30 -point scale. (The higher the value, the more favorable was the client's response to the report.) The independent variables found to have the most predictive power are as follows:
$x_{1}:$ Therapist's age (years)
$x_{2}:$ Therapist's gender $(1$ if male, 0 if female $)$ $x_{3}:$ Degree of therapist's role strain ( 25 -point scale $)$ $x_{4}:$ Strength of client-therapist relationship ( 40 -point scale $\mathrm{x}_{5}:$ Type of case $(1$ if family, 0 if not $)$ $x_{1} x_{2}:$ Age $\times$ Gender interaction
a. Hypothesize a first-order model relating $y$ to each of the five independent variables and the interaction term.
b. Give the null hypothesis for testing the contribution of $x_{4},$ strength of client-therapist relationship, to the model.
c. The test statistic for the test suggested in part $\mathbf{b}$ was $t=4.408,$ with an associated $p$ -value of $.001 .$ Interpret this result.
d. The estimated $\beta$ coefficient for the $x_{1} x_{2}$ interaction term was positive and highly significant $(p<.001)$. According to the researchers, "This interaction suggests that $\ldots$ as the age of the therapist increased,... male therapists were less likely to get negative client reactions than were female therapists." Do you agree?
e. For the model presented here, $R^{2}=.2946 .$ Interpret this value.

Daniel Caproni
Daniel Caproni
Numerade Educator
05:02

Problem 202

Processed straw as thermal insulation. An article published in Engineering Structures and Technologies (Sept.
2012) presented the results of a study on the use of processed straw as thermal insulation for homes. Chopped straw specimens were prepared and tested for thermal conductivity at a temperature of $10^{\circ}$ Celsius. Two variables were measured for each of 25 straw specimens:
$y=$ thermal conductivity (watts per meter-Kelvin) and $x=$ density (kilograms per cubic meter). The data are provided in the accompanying table. Consider the quadratic model $E(y)=\beta_{0}+\beta_{1} x+\beta_{0} x^{2}$. Conduct a complete analysis (including a residual analysis) of this model. Summarize your results in a report.

Banhishikha Sinha
Banhishikha Sinha
Numerade Educator
08:37

Problem 203

Abundance of bird species. Multiple-regression analysis was used to model the abundance $y$ of an individual bird species in transects in the United Kingdom (Journal of Applied Ecology, Vol. 32,1995 ). Three of the independent variables used in the model, all field boundary attributes, are
1. Transect location (small pasture field, small arable field, or large arable field)
2. Land use (pasture or arable) adjacent to the transect
3. Total number of trees in the transect
a. Identify each of the independent variables as a quantitative or qualitative variable.
b. Write a first-order model for $E(y)$ as a function of the total number of trees.
c. Add main-effect terms for transect location to the model you wrote in part $\mathbf{b}$. Graph the hypothesized relationships of the new model.
d. Add main-effect terms for land use to the model you came up with in part $\mathbf{c}$. In terms of the $\beta$ 's of the new model, what is the slope of the relationship between $E(y)$ and number of trees for any combination of transect location and land use?
e. Add terms for interaction between transect location and land use to the model you arrived at in part $\mathbf{d}$. Do these interaction terms affect the slope of the relationship between $E(y)$ and number of trees? Explain.
f. Add terms for interaction between number of trees and all coded dummy variables to the model you formulated in part $\mathbf{e}$. In terms of the $\beta$ 's of the new model, give the slope of the relationship between $E(y)$ and number of trees for each combination of transect location and land use.

Donald Albin
Donald Albin
Numerade Educator
15:18

Problem 204

IQ and The Bell Curve. In Exercise 5.157 (p. 306), we introduced The Bell Curve (New York: Free Press,
1994) by Richard Herrnstein and Charles Murray (H\&M), a controversial book about race, genes, IQ, and economic mobility. The book heavily employs statistics and statistical methodology in an attempt to support the authors' positions on the relationships among these variables and their social consequences. The main theme of The Bell Curve can be summarized as follows:
1. Measured intelligence (IQ) is largely genetically inherited.
2. IQ is correlated positively with a variety of socioeconomic status success measures, such as a prestigious job, a high annual income, and high educational attainment.
3. From 1 and 2 , it follows that socioeconomic successes are largely genetically caused and therefore resistant to educational and environmental interventions (such as affirmative action).
The statistical methodology (regression) employed by the authors and the inferences derived from the statistics were critiqued in Chance (Summer 1995) and The Journal of the American Statistical Association (Dec.
1995). The following are just a few of the problems with H\&M's use of regression that have been identified:
Problem 1 H\&M consistently use a trio of independent variables $-\mathrm{IQ},$ socioeconomic status, and age $-\mathrm{in}$ a series of first-order models designed to predict dependent social outcome variables such as income and unemployment. (Only on a single occasion are interaction terms incorporated.) Consider, for example, the model
$$
E(y)=\beta_{0}+\beta_{1} x_{1}+\beta_{2} x_{2}+\beta_{3} x_{3}
$$
where $y=$ income, $x_{1}=\mathrm{IQ}, x_{2}=$ socioeconomic status, and $x_{3}=$ age. $\mathrm{H} \& \mathrm{M}$ utilize $t$ -tests on the individual $\beta$ parameters to assess the importance of the independent variables. As with most of the models considered in The Bell Curve, the estimate of $\beta_{1}$ in the income model is positive and statistically significant at $\alpha=.05,$ and the associated $t$ -value is larger (in absolute value) than the $t$ -values associated with the other independent variables. Consequently, $H \& M$ claim that $I Q$ is a better predictor of income than the other two independent variables. No attempt was made to determine whether the model was properly specified or whether the model provides an adequate fit to the data.
Problem 2 In an appendix, the authors describe multiple regression as a "mathematical procedure that yields coefficients for each of [the independent variables], indicating how much of a change in [the dependent variable] can be anticipated for a given change in any particular [independent] variable, with all the others held constant." Armed with this information and the fact that the estimate of $\beta_{1}$ in the model just described is positive, $H \& M$ infer that a high $I Q$ necessarily implies (or causes) a high income, and a low IQ inevitably leads to a low income. (Cause-and-effect inferences like this are made repeatedly throughout the book.)

Problem 3 The title of the book refers to the normal distribution and its well-known "bell-shaped" curve. There is a misconception among the general public that scores on intelligence tests (IQS) are normally distributed. In fact, most IQ scores have distributions that are decidedly skewed. Traditionally, psychologists and psychometricians have transformed these scores so that the resulting numbers have a precise normal distribution. H\&M make a special point to do this. Consequently, the measure of $I Q$ used in all the regression models is normalized (i.e., transformed so that the resulting distribution is normal), despite the fact that regression methodology does not require predictor (independent) variables to be normally distributed.
Problem 4 A variable that is not used as a predictor of social outcome in any of the models in The Bell Curve is level of education. H\&M purposely omit education from the models, arguing that IQ causes education, not the other way around. Other researchers who have examined H\&M's data report that when education is included as an independent variable in the model, the effect of $I Q$ on the dependent variable (say, income) is diminished.
a. Comment on each of the problems identified. Why do these problems cast a shadow on the inferences made by the authors?
b. Using the variables specified in the model presented, describe how you would conduct the multiple-regression analysis. (Propose a more complicated model and describe the appropriate model tests, including a residual analysis.)

Heather Duong
Heather Duong
Numerade Educator
10:51

Problem 205

FLAG study of bid collusion. Road construction contracts in the state of Florida are awarded on the basis of competitive, sealed bids; the contractor who bids the lowest price wins the contract. During the $1980 \mathrm{~s}$, the Office of the Florida Attorney General (FLAG) suspected numerous contractors of practicing bid collusion (i.e., setting the winning bid price above the fair, or competitive, price in order to increase their own profit margin). By comparing the prices bid (and other important bid variables) of the fixed (i.e., rigged) contracts with the competitively bid contracts, FLAG was able to establish invaluable benchmarks for detecting future bid rigging. FLAG collected data on 279 road construction contracts. For each contract, the following variables were measured (the data are saved in the FLAG file):

Jon Southam
Jon Southam
Numerade Educator