Chapter Questions
Complete the following for the $(x, y)$ pairs of data points $(1,5),(3,7),(4,6),(5,8)$, and $(7,9)$.a. Prepare a scatter plot of these data points.b. Compute $b_1$.c. Compute $b_{\mathrm{a}}$d. What is the equation of the regression line?
The following data give $X$, the price charged per piece of plywood, and $Y$, the quantity sold (in thousands).$$\begin{array}{cc}\hline \text { Price per Piece, } X & \text { Thousands of Pieces Sold, } Y \\\hline \$ 6 & 80 \\7 & 60 \\8 & 70 \\9 & 40 \\10 & 0 \\\hline\end{array}$$a. Prepare a scatter plot of these data points.b. Compute the covariance.c. Compute and interpret $b_1$.d. Compute $b_0$e. What quantity of plywood would you expect to sell if the price were $\$ 7$ per piece?
A random sample of data for 7 days of operation produced the following (price, quantity) data values:$$\begin{array}{cc}\hline \text { Price per Gallon of Paint, } X & \text { Quantity Sold, } Y \\\hline 10 & 100 \\8 & 120 \\5 & 200 \\4 & 200 \\10 & 90 \\7 & 110 \\6 & 150 \\\hline\end{array}$$a. Prepare a scatter plot of the data.b. Compute and interpret $b_1$.c. Compute and interpret $b_c$d. How many gallons of paint would you expect to sell if the price is $\$ 7$ per gallon?
A large consumer goods company has been studying the effect of advertising on total profits. As part of this study, data on advertising expenditures and total sales were collected for a five-month period and are as follows:$$(10,100)(15,200)(7,80)(12,120)(14,150)$$The first number is advertising expenditures and the second is total sales.a. Plot the data.b. Does the plot provide evidence that advertising has a positive effect on sales?c. Compute the regression coefficients, $b_0$ and $b_1$.
Abdul Hassan, president of Floor Coverings Unlimited, has asked you to study the relationship between market price and the tons of rugs supplied by his competitor, Best Floor, Inc. He supplies you with the following observations of price per ton and number of tons, obtained from his secret files:$$(2,5)(4,10)(3,8)(6,18)(3,6)(5,15)(6,20)(2,4)$$The first number for each observation is price and the second is quantity.a. Prepare a scatter plot.b. Determine the regression coefficients, $b_0$ and $b_1$.c. Write a short explanation of the regression equation that tells Abdul how the equation can be used to describe his competition. Include an indication of the range over which the equation can be applied.
The following ordered pairs provide data about some Nestlé snacks, where the first number is grams of sugar and the second is the number of calories for each snack.$(3,110),(14,180),(13,150),(11,120),(8,100)$, $(5,70),(7,140),(15,200),(12,130)$a. Construct a scatter plot of the data. Does a clear linear relationship exist between the two variables?b. Estimate the regression equation and identify the value of the slope.c. Which conclusion can you draw from your results?
Given the regression equation$$Y=100+10 X$$a. What is the change in $Y$ when $X$ changes by +3 ?b. What is the change in $Y$ when $X$ changes by -4 ?c. What is the predicted value of $Y$ when $X=12$ ?d. What is the predicted value of $Y$ when $X=23$ ?e. Does this equation prove that a change in $X$ causes a change in $\gamma$ ?
Given the regression equation$$Y=-50+12 X$$a. What is the change in $Y$ when $X$ changes by +3 ?b. What is the change in $Y$ when $X$ changes by -4 ?c. What is the predicted value of $Y$ when $X=12$ ?d. What is the predicted value of $Y$ when $X=23$ ?e. Does this equation prove that a change in $X$ causes a change in $r$ ?
Given the regression equation$$Y=43+10 X$$a. What is the change in $Y$ when $X$ changes by +8 ?b. What is the change in $Y$ when $X$ changes by -6 ?c. What is the predicted value of $Y$ when $X=11$ ?d. What is the predicted value of $Y$ when $X=29$ ?e. Does this equation prove that a change in $X$ causes a change in $Y$ ?
Given the regression equation$$Y=100+21 X$$a. What is the change in $Y$ when $X$ changes by +5 ?b. What is the change in $Y$ when $X$ changes by -7 ?c. What is the predicted value of $Y$ when $X=14$ ?d. What is the predicted value of $Y$ when $X=27$ ?e. Does this equation prove that a change in $X$ causes a change in $Y$ ?
In Example 11.1 a linear regression model was developed. Use that model to answer the following.a. Interpret the coefficient $b_1=2.545$ for the plant manager.b. How many tables would be produced on average with 19 workers?c. Suppose you were asked to estimate the number of tables produced if only five workers were available. Discuss your response to this request.
As the new market manger for Blue Crunchies breakfast cereal, you are asked to estimate the demand for next month using regression analysis. Two months ago the target market had 20,000 families and sales were 3,780 boxes and, 1 month ago the target market was 40,000 families and sales were 5,349 boxes. Next month you plan to target 75,000 families. How would you respond to the request to use regression analysis and the currently available data to estimate sales next month?
Consider the sales prediction model developed for Northern Household Goods in Example 11.2.a. Estimate per capita sales if the mean disposable income is $$\$ 56,000$$.b. Interpret the coefficients $b_0$ and $b_1$ for Northern's management.c. You have been asked to estimate per capita sales if mean disposable income grows to $$\$ 64,000$$. Discuss how you would proceed and indicate your cautions.
What is the difference between a population linear model and an estimated linear regression model?
Explain the difference between the residual $e_j$ and the model error $\varepsilon_i$.
Suppose that we obtained an estimated equation for the regression of weekly sales of palm pilots and the price charged during the week. Interpret the constant $b_0$ for the product brand manager.
A regression model of total grocery sales on disposable income was estimated using data from small, isolated towns in the western United States. Prepare a list of factors that might contribute to the random error term.
Compute the coefficients for a least squares regression equation and write the equation, given the following sample statistics.a. $\bar{x}=50, \bar{y}=100, s_x=25, s_y=75, r_{x y}=0.6, n=60$b. $\bar{x}=60, \bar{y}=210, s_x=35, s_y=65, r_{x y}=0.7, n=60$c. $\bar{x}=20, \bar{y}=100, s_x=60, s_y=78, r_{x y}=0.75, n=60$d. $\bar{x}=10, \bar{y}=50, s_x=100, s_y=75, r_{x y}=0.4, n=60$e. $\bar{x}=90, \bar{y}=200, s_x=80, s_y=70, r_{x y}=0.6, n=60$
A company sets different prices for a particular DVD system in eight different regions of the country. The accompanying table shows the numbers of units sold and the corresponding prices (in dollars).$$\begin{array}{lrrrrrrrr}\hline \text { Sales } & 420 & 380 & 350 & 400 & 440 & 380 & 450 & 420 \\\hline \text { Price } & 104 & 195 & 148 & 204 & 96 & 256 & 141 & 109 \\\hline\end{array}$$a. Graph these data, and estimate the linear regression of sales on price.b. What effect would you expect a $$\$ 50$$ increase in price to have on sales?
For a sample of $\mathbf{2 0}$ monthly observations, a financial analyst wants to regress the percentage rate of return $(Y)$ of the common stock of a corporation on the percentage rate of return $(X)$ of the Standard \& Poor's 500 index. The following information is available:$$\sum_{i=1}^{20} y_i=22.6 \quad \sum_{i=1}^{20} x_i=25.4 \quad \sum_{i=1}^{20} x_i^2=145.7 \quad \sum_{i=1}^{20} x_i y_i=150.5$$a. Estimate the linear regression of $Y$ on $X$.b. Interpret the slope of the sample regression line.c. Interpret the intercept of the sample regression line.
A corporation administers an aptitude test to all new sales representatives. Management is interested in the extent to which this test is able to predict sales representatives' eventual success. The accompanying table records average weekly sales (in thousands of dollars) and aptitude test scores for a random sample of eight representatives.$$\begin{array}{lllllllll}\hline \text { Weokly sales } & 10 & 12 & 28 & 24 & 18 & 16 & 15 & 12 \\\hline \text { Test score } & 55 & 60 & 85 & 75 & 80 & 85 & 65 & 60 \\\hline\end{array}$$a. Estimate the linear regression of weekly sales on aptitude test scores.b. Interpret the estimated slope of the regression line.
In Wanchai Computer Centers in Hong Kong, there are dozens of computer shops selling multiple laptop brands. After a survey in one of them, 10 were selected. The ordered pairs show the speed of each computer's CPU in gigahertz and its price in Hong Kong dollars ( $1 \mathrm{USD}=7.78 \mathrm{HKD})$.$$\begin{aligned}& (1.8,14,500),(1.6,12,290),(2.0,17,500),(1.6,16,500), \\& (1.8,19,650),(2.4,21,000),(1.2,7,500),(1.4,12,500), \\& (1.6,14,650),(2.0,18,350)\end{aligned}$$a. Determinate the regression equation of the sample.b. Find the intercept and the slope of the equation.c. Compute the coefficient of determination and interpret its meaning in this specific context.
Refer to the data file Dow Jones, which contains percentage change $(X)$ in the Dow Jones index over the first five trading days of the year and percentage change $(Y)$ in the index over the whole year.a. Estimate the linear regression of $Y$ on $X$.b. Provide interpretations of the intercept and slope of the sample regression line.
On Friday, November 13, 1989, prices on the New York Stock Exchange fell steeply; the Standard \& Poor's 500-share index was down $6.1 \%$ on that day. The data file New York Stock Exchange Gains and Losses shows the percentage losses (y) of the 25 largest mutual funds on November 13, 1989. Also shown are the percentage gains $(x)$, assuming reinvested dividends and capital gains, for these same funds for 1989 through November 11 .a. Estimate the linear regression of November 13 losses on pre-November 13, 1989, gains.b. Interpret the slope of the sample regression line.
Compute SSR, SSE, ser and the coefficient of determination, given the following statistics computed from a random sample of pairs of $X$ and $Y$ observations.a. $\sum_{i=1}^n\left(y_i-\bar{y}\right)^2=100,000, R^2=0.50, n=52$b. $\sum_{i=1}^n\left(y_i-\bar{y}\right)^2=90,000, R^2=0.70, n=52$c. $\sum_{i=1}^n\left(y_i-\bar{y}\right)^2=240, R^2=0.80, n=52$d. $\sum_{i=1}^n\left(y_i-\bar{y}\right)^2=200,000, R^2=0.30, n=74$e. $\sum_{i=1}^n\left(y_i-\bar{y}\right)^2=60,000, R^2=0.90, n=40$
Let the sample regression line be$$y_i=b_0+b_1 x_i+e_i=y_i+e_i(i=1,2, \ldots, n)$$and let $\bar{x}$ and $\bar{y}$ denote the sample means for the independent and dependent variables, respectively.a. Show that$$c_i=y_i-\bar{y}-b\left(x_i-\bar{x}\right)$$b. Using the result in part $a$, show that$$\sum_{i=1}^n e_i=0$$c. Using the result in part $\mathrm{a}$, show that$$\sum_{i=1}^n e_i^2=\sum_{i=1}^n\left(y_i-\bar{y}\right)^2-b^2 \sum_{i=1}^n\left(x_i-\bar{x}\right)^2$$d. Show that$$\hat{y}_i-\bar{y}=b_i\left(x_i-\bar{x}\right)$$e. Using the results in parts $\mathrm{c}$ and $d$, show that$$S S T=S S R+S S E$$f. Using the result in part $a$, show that$$\sum_{i=1}^n e_i\left(x_i-\bar{x}\right)=0$$
Let$$R^2=\frac{S S R}{S S T}$$denote the coefficient of determination for the sample regression line.a. Using part $d$ of the previous exercise, show that$$R^2=b_1^2 \frac{\sum_{i=1}^n\left(x_i-\bar{x}\right)^2}{\sum_{i=1}^n\left(y_i-\bar{y}\right)^2}$$b. Using the result in part a, show that the coefficient of determination is equal to the square of the sample correlation between $X$ and $Y$.c. Let $b_1$ be the slope of the least squares regression of $Y$ on $X, b_1^*$ be the slope of the least squares regression of $X$ on $Y$, and $r$ be the sample correlation between $X$ and $Y$. Show that $b_1 \cdot b_1^*=r^2$
Find and interpret the coefficient of determination for the regression of DVD system sales on price, using the following data.$$\begin{array}{rrrrrrrrr}\hline \text { Sales } & 420 & 380 & 350 & 400 & 440 & 380 & 450 & 420 \\\hline \text { Price } & 98 & 194 & 244 & 207 & 89 & 261 & 149 & 198 \\\hline\end{array}$$
Find and interpret the coefficient of determination for the regression of the percentage change in the Dow Jones index in a year based on the percentage change in the index over the first five trading days of the year. Compare your answer with the sample correlation found for these data. Use the data file Dow Jones.
Find the proportion of the sample variability in mutual fund percentage losses on November 13, 1989, explained by their linear dependence on 1989 percentage gains through November 12 , based on the data in the data file New York Stock Exchange Gains and Losses.
In a study it was shown that for a sample of 353 college faculty, the correlation was 0.11 between annual raises and teaching evaluations. What would be the coefficient of determination of a regression of annual raises on teaching evaluations for this sample? Interpret your result.
Given the simple regression model$$Y=\beta_0+\beta_1 X$$and the regression results that follow, test the null hypothesis that the slope coefficient is 0 versus the alternative hypothesis of greater than zero using probability of Type I error equal to 0.05 , and determine the two-sided $95 \%$ and $99 \%$ confidence intervals.a. A random sample of size $n=38$ with$$b_1=5 \quad s_{b_1}=2.1$$b. A random sample of size $n=46$ with$$b_1=5.2 \quad s_{b_1}=2.1$$c. A random sample of size $n=38$ with$$b_1=2.7 \quad s_{b_1}=1.87$$d. A random sample of size $n=29$ with$$b_1=6.7 \quad s_{b_1}=1.8$$
Use a simple regression model to test the hypothesis$$H_0: \beta_1=0$$versus$$H_1: \beta_1 \neq 0$$with $\alpha=0.05$, given the following regression statistics.a. The sample size is $35, S S T=100,000$, and the correlation between $X$ and $Y$ is 0.46 .b. The sample size is $61, S S T=123,000$, and the correlation between $X$ and $Y$ is 0.65 .c. The sample size is $25, S S T=128,000$, and the correlation between $X$ and $Y$ is 0.69 .
Mumbai Electronics is planning to extend its marketing region from the western United States to include the midwestern states. In order to predict its sales in this new region, the company has asked you to develop a linear regression of DVD system sales on price, using the following data supplied by the marketing department:$$\begin{array}{lrrrrrrrr}\hline \text { Sales } & 418 & 384 & 343 & 407 & 432 & 386 & 444 & 427 \\\hline \text { Price } & 98 & 194 & 231 & 207 & 89 & 255 & 149 & 195 \\\hline\end{array}$$a. Use an unbiased estimation procedure to find an estimate of the variance of the error terms in the population regression.b. Use an unbiased estimation procedure to find an estimate of the variance of the least squares estimator of the slope of the population regression line.c. Find a $90 \%$ confidence interval for the slope of the population regression line.
A fast-food chain decided to carry out an experiment to assess the influence of advertising expenditure on sales. Different relative changes in advertising expenditure, compared to the previous year, were made in eight regions of the country, and resulting changes in sales levels were observed. The accompanying table shows the results.$$\begin{array}{lllllllll}\hline \begin{array}{l}\text { Increase in } \\\text { advertising } \\\text { expenditure (\%) }\end{array} & 0 & 4 & 14 & 10 & 9 & 8 & 6 & 1 \\\hline\begin{array}{l}\text { Increase in } \\\text { sales }(\%)\end{array} & 2.4 & 7.2 & 10.3 & 9.1 & 10.2 & 4.1 & 7.6 & 3.5 \\\hline\end{array}$$a. Estimate by least squares the linear regression of increase in sales on increase in advertising expenditure.b. Find a $90 \%$ confidence interval for the slope of the population regression line.
You have been asked to determine the effect of per capita disposable income on retail sales using cross-section data by state. The data are contained in the data file Economic Activity. Estimate the appropriate regression equation and determine the $95 \%$ confidence interval for the expected change in retail sales that would result from a $$\$ 1,000$$ increase in per capita disposable income.
Estimate the regression equation for the percentage change in the Dow Jones index in a year on the percentage change in the index over the first five trading days of the year. Use the data file Dow Jones.a. Use an unbiased estimation procedure to find a point estimate of the variance of the error terms in the population regression.b. Use an unbiased estimation procedure to find a point estimate of the variance of the least squares estimator of the slope of the population regression line.c. Find and interpret a $95 \%$ confidence interval for the slope of the population regression line.d. Test at the $10 \%$ significance level, against a two-sided alternative, the null hypothesis that the slope of the population regression line is 0 .
Estimate a linear regression model for mutual fund losses on November 13, 1989, using the data file New York Stock Exchange Gains and Losses.a. Use an unbiased estimation procedure to obtain a point estimate of the variance of the error terms in the population regression.b. Use an unbiased estimation procedure to obtain a point estimate of the variance of the least squares estimator of the slope of the population regression line.c. Find $90 \%, 95 \%$, and $99 \%$ confidence intervals for the slope of the population regression line.
Given a simple regression analysis, suppose that we have obtained a fitted regression model$$\hat{y}_i=12+5 x_i$$and also$$s_e=9.67 \quad \bar{x}=8 \quad n=32 \sum_{i=1}^n\left(x_i-\bar{x}\right)^2=500$$Find the $95 \%$ confidence interval and $95 \%$ prediction interval for the point where $x=13$.
Given a simple regression analysis, suppose that we have obtained a fitted regression model$$\hat{y}_i=14+7 x_i$$and also$$s_e=7.45 \quad \bar{x}=8 \quad n=25 \sum_{i=1}^n\left(x_i-\bar{x}\right)^2=300$$Find the $95 \%$ confidence interval and $95 \%$ prediction interval for the point where $x=11$.
Given a simple regression analysis, suppose that we have obtained a fitted regression model$$\hat{y}_i=22+8 x_i$$and also$$s_e=3.45 \quad \bar{x}=11 \quad n=22 \sum_{i=1}^n\left(x_i-\bar{x}\right)^2=400$$Find the $95 \%$ confidence interval and $95 \%$ prediction interval for the point where $x=17$.
Given a simple regression analysis, suppose that we have obtained a fitted regression model$$\hat{y}_i=8+10 x_i$$and also$$s_e=11.23 \quad \bar{x}=8 \quad n=44 \sum_{i=1}^n\left(x_i-\bar{x}\right)^2=800$$Find the $95 \%$ confidence interval and $95 \%$ prediction interval for the point where $x=17$.
A sample of 25 blue-collar employees at a production plant was taken. Each employee was asked to assess his or her own job satisfaction $(x)$ on a scale of 1 to 10. In addition, the numbers of days absent (y) from work during the last year were found for these employees. The sample regression line$$\hat{y}_i=11.6-1.2 x$$was estimated by least squares for these data. Also found were$$\bar{x}=6.0 \sum_{i=1}^{25}\left(x_i-\bar{x}\right)^2=130.0 \text { SSE }=80.6$$a. Test, at the $1 \%$ significance level against the appropriate one-sided alternative, the null hypothesis that job satisfaction has no linear effect on absenteeism.b. A particular employee has job satisfaction level 4. Find a $90 \%$ interval for the number of days this employee would be absent from work in a year.
Doctors are interested in the relationship between the dosage of a medicine and the time required for a patient's recovery. The following table shows, for a sample of 10 patients, dosage levels (in grams) and recovery times (in hours). These patients have similar characteristics except for medicine dosages.$$\begin{array}{lrrrrrrrrrr}\hline \text { Dosage level } & 1.2 & 1.3 & 1.0 & 1.4 & 1.5 & 1.8 & 1.2 & 1.3 & 1.4 & 1.3 \\\hline \text { Recovery time } & 25 & 28 & 40 & 38 & 10 & 9 & 27 & 30 & 16 & 18 \\\hline\end{array}$$a. Estimate the linear regression of recovery time on dosage level.b. Find and interpret a $90 \%$ confidence interval for the slope of the population regression line.c. Would the sample regression derived in part a be useful in predicting recovery time for a patient given 2.5 grams of this drug? Explain your answer.
For a sample of 20 monthly observations, a financial analyst wants to regress the percentage rate of return $(Y)$ of the common stock of a corporation on the percentage rate of return $(X)$ of the Standard \& Poor's 500 index. The following information is available:$$\begin{gathered}\sum_{i=1}^{20} y_i=22.6 \sum_{i=1}^{20} x_i=25.4 \quad \sum_{i=1}^{20} x_i^2=145.7 \\\sum_{i=1}^{20} x_i y_i=150.5 \quad \sum_{i=1}^{20} y_i^2=196.2\end{gathered}$$a. Test the null hypothesis that the slope of the population regression line is 0 against the alternative that it is positive.b. Test against the two-sided alternative the null hypothesis that the slope of the population regression line is 1.
Estimate a linear regression model for mutual fund losses on November 13, 1989, on previous gains in 1989, using the data file New York Stock Exchange Gains and Losses. Test, against a two-sided alternative, the null hypothesis that mutual fund losses on Friday, November 13, 1989, did not depend linearly on previous gains in 1989.
Denote by $r$ the sample correlation between a pair of random variables.a. Show that$$\frac{1-r^2}{n-2}=\frac{s_e^2}{S S T}$$b. Using the result in part a, show that$$\frac{r}{\sqrt{\left(1-r^2\right) /(n-2)}}=\frac{b}{s_e / \sqrt{\sum\left(x_i-\bar{x}\right)^2}}$$
In a UK business school, lecturers have tried to determine if the number of hours students attend lectures has any measurable effect on the grades obtained by the students. The following data from a sample of 14 students in an international business class show hours of attendance and resulting grades.$$\begin{aligned}& (22,72),(20,64),(24,70),(8,34),(12,40),(16,40), \\& (18,52),(16,45),(20,68),(24,65),(28,72), \\& (20,64),(10,38),(16,44)\end{aligned}$$a. Estimate the regression line.b. Find a $95 \%$ confidence interval for the slope of the regression line.
For a sample of 74 monthly observations the regression of the percentage return on gold $(y)$ against the percentage change in the consumer price index $(x)$ was estimated. The sample regression line, obtained through least squares, was as follows:$$y=-0.003+1.11 x$$The estimated standard deviation of the slope of the population regression line was 2.31. Test the null hypothesis that the slope of the population regression line is 0 against the alternative that the slope is positive.
A liquor wholesaler is interested in assessing the effect of the price of a premium scotch whiskey on the quantity sold. The results in the accompanying table on price, in dollars, and sales, in cases, were obtained from a sample of 8 weeks of sales records.$$\begin{array}{lllllllll}\hline \text { Price } & 19.2 & 20.5 & 19.7 & 21.3 & 20.8 & 19.9 & 17.8 & 17.2 \\\hline \text { Sales } & 25.4 & 14.7 & 18.6 & 11.4 & 11.1 & 15.7 & 29.2 & 35.2 \\\hline\end{array}$$Test, at the $5 \%$ level against the appropriate one-sided alternative, the null hypothesis that sales do not depend linearly on price for this premium scotch whiskey.
The data file Dow Jones shows percentage changes $\left(x_i\right)$ in the Dow Jones index over the first five trading days of each of 13 years and also the corresponding percentage changes $\left(y_i\right)$ in the index over the whole year. If the Dow Jones index increases by $1.0 \%$ in the first five trading days of a year, find $90 \%$ confidence intervals for the actual and also the expected percentage changes in the index over the whole year. Discuss the distinction between these intervals.
You have been asked to study the relationship between mean health care costs and mean disposable income using the state level data contained in the data file Economic Activity. Estimate the regression of health and personal expenditures on disposable income. Compute the $95 \%$ prediction interval and the $95 \%$ confidence interval for health and personal expenditures when disposable income is $$\$ 32,000$$.
An economic policy research organization has asked you to study the relationship between disposable income and unemployment level. The data forthis study are contained in the data file Economic Activity. As a first step you estimate the regression model for the relationship between unemployment regressed on disposable income. Determine if there is a significant relationship between unemployment and disposable income and whether the relationship is increasing or decreasing. Compute the $95 \%$ prediction interval for unemployment when disposable income is $$\$ 30,000$$.
Given the following pairs of $(x, y)$ observations, compute the sample correlation.a. $(2,5),(5,8),(3,7),(1,2),(8,15)$b. $(7,5),(10,8),(8,7),(6,2),(13,15)$c. $(12,4),(15,6),(16,5),(21,8),(14,6)$d. $(2,8),(5,12),(3,14),(1,9),(8,22)$
Test the null hypothesis$$H_0: \rho=0$$versus$$H_1: \rho \neq 0$$given the following.a. A sample correlation of 0.35 for a random sample of size $n=40$b. A sample correlation of 0.50 for a random sample of size $n=60$c. A sample correlation of 0.62 for a random sample of size $n=45$d. A sample correlation of 0.60 for a random sample of size $n=25$
An instructor in a statistics course set a final examination and also required the students to do a data analysis project. For a random sample of 10 students, the scores obtained are shown in the table. Find the sample correlation between the examination and project scores.$$\begin{array}{lcccccccccc}\hline \text { Examination } & 81 & 62 & 74 & 78 & 93 & 69 & 72 & 83 & 90 & 84 \\\hline \text { Project } & 76 & 71 & 69 & 76 & 87 & 62 & 80 & 75 & 92 & 79 \\\hline\end{array}$$
In the study of 49 countries discussed in Example 11.4, the sample correlation between the experts' political riskiness score and the infant mortality rate in these countries was 0.75 . Test the null hypothesis of no correlation between these quantities against the alternative of positive correlation.
For a random sample of 353 high school teachers, the correlation between annual raises and teaching evaluations was found to be 0.11 . Test the null hypothesis that these quantities are uncorrelated in the population against the alternative that the population correlation is positive.
The sample correlation for 68 pairs of annual returns on common stocks in country $\mathrm{A}$ and country $\mathrm{B}$ was found to be 0.51. Test the null hypothesis that the population correlation is 0 against the alternative that it is positive.
The accompanying table and the data file Dow Jones show percentage changes $\left(x_i\right)$ in the Dow Jones index over the first five trading days of each of 13 years and also the corresponding percentage changes $\left(y_i\right)$ in the index over the whole year.a. Calculate the sample correlation.b. Test, at the $10 \%$ significance level against a two-sided alternative, the null hypothesis that the population correlation is 0 .$$\begin{array}{rrrr}\hline x & {y} & {x} & {y} \\\hline 1.5 & 14.9 & 5.6 & 2.3 \\0.2 & -9.2 & -1.4 & 11.9 \\-0.1 & 19.6 & 1.4 & 27.0 \\2.8 & 20.3 & 1.5 & -4.3 \\2.2 & -3.7 & 4.7 & 20.3 \\-1.6 & 27.7 & 1.1 & 4.2 \\-1.3 & 22.6 & & \\\hline\end{array}$$
A college administers a student evaluation questionnaire for all its courses. For a random sample of 12 courses, the accompanying table and the data file Student Evaluation show both the average student ratings of the instructor (on a scale of 1 to 5), and the average expected grades of the students (on a scale of $\mathrm{A}=4$ to $\mathrm{F}=0$ ).$$\begin{array}{ccccccccccccc}\hline \begin{array}{l}\text { Instructor } \\\text { rating }\end{array} & 2.8 & 3.7 & 4.4 & 3.6 & 4.7 & 3.5 & 4.1 & 3.2 & 4.9 & 4.2 & 3.8 & 3.3 \\\hline \begin{array}{l}\text { Expected } \\\text { grade }\end{array} & 26 & 2.9 & 3.3 & 3.2 & 3.1 & 2.8 & 2.7 & 2.4 & 3.5 & 3.0 & 3.4 & 2.5 \\\hline\end{array}$$a. Find the sample correlation between instructor ratings and expected grades.b. Test, at the $10 \%$ significance level, the hypothesis that the population correlation coefficient is zero against the alternative that it is positive.
In an advertising study the researchers wanted to determine if there was a relationship between the per capita cost and the per capita revenue. The following variables were measured for a random sample of advertising programs:$x_i=$ Cost of Advertisement $\div$ Number of Inquiries Received$y_i=$ Revenue from Inquiries + Number of Inquiries ReceivedThe sample data results are shown in the data file Advertising Revenue. Find the sample correlation and test, against a two-sided alternative, the null hypothesis that the population correlation is 0 .
As part of a process to build a new automotive portfolio, you have been asked to determine the beta coefficients for AB Volvo and General Motors. Data for this task are contained in the data file Return on Stock Price 60 Months. Compare the required return on the two stocks to compensate for the risk.
In this exercise you are asked to determine the beta coefficient for Senior Housing Properties Trust. Data for this task are contained in the data file Return on Stock Price 60 Months. Interpret this coefficient.
An investor is considering the possibility of including TCF Financial in her portfolio. Data for this task are contained in the data file Return on Stock Price 60 Months. Compare the mean and variance of the monthly return with the S & P 500 mean and variance. Then, estimate the beta coefficient. Based on this analysis, what would you recommend to the investor?
Allied Financial is considering the possibility of adding one or more computer industry stocks to its portfolio. You are asked to consider the possibility of Seagate, Microsoft, and Tata Information systems. Data for this task are contained in the data file Return on Stock Price 60 Months. Compare the return on these three stocks by computing the beta coefficients and the mean and variance of the returns. What is your recommendation regarding these three stocks?
Charlie Ching has asked you to analyze the possibility of including Seneca Foods and Safeco in his portfolio. Data for this task are contained in the data file Return on Stock Price 60 Months. Compute the beta coefficients for the stock price growth for each stock. Then construct a portfolio that includes equal dollar value for both stocks. Compute the beta coefficient for that portfolio. Compare the mean and variance for the portfolio with the S & P 500. What is your recommendation regarding the inclusion of these two stocks in Charlie's portfolio?
Frank Anscombe, senior research executive, has asked you to analyze the following four linear models using data contained in the data file Anscombe:$$\begin{aligned}& Y_1=\beta_0+\beta_1 X_1 \\& Y_2=\beta_0+\beta_1 X_2 \\& Y_3=\beta_0+\beta_1 X_3 \\& Y_4=\beta_0+\beta_1 X_4\end{aligned}$$Use your computer package to obtain a linear regression estimate for each model. Prepare a scatter plot for the data used in each model. Write a report, including regression and graphical outputs, that compares and contrasts the four models.
Josie Foster, president of Public Research, Inc., has asked for your assistance in a study of the occurrence of crimes in different states before and after a large federal government expenditure to reduce crime. As part of this study she wants to know if the crime rate for selected crimes after the expenditure can be predicted using the crime rate before the expenditure. She has asked you to test the hypothesis that crime before predicts crime after for total crime rate and for the murder, rape, and robbery rates. The data for your analysis are contained in the data file Crime Study. Perform appropriate analysis and write a report that summarizes your results.
For a random sample of 53 building supply stores in a chain, the correlation between annual sales per square meter of floor space and annual rent per square meter of floor space was found to be 0.37 . Test the null hypothesis that these two quantities are uncorrelated in the population against the alternative that the population correlation is positive.
For a random sample of 526 firms, the sample correlation between the proportion of a firm's officers who are directors and a risk-adjusted measure of return on the firm's stock was found to be 0.1398 . Test, against a two-sided alternative, the null hypothesis that the population correlation is 0 .
For a sample of 66 months, the correlation between the returns on Canadian and Singapore 10-year bonds was found to be 0.293 . Test the null hypothesis that the population correlation is 0 against the alternative that it is positive.
Based on a sample on $n$ observations, $\left(x_1, y_1\right)$, $\left(x_2, y_2\right), \ldots,\left(x_n, y_n\right)$, the sample regression of $y$ on $x$ is calculated. Show that the sample regression line passes through the point $(x=\bar{x}, y=\bar{y})$, where $\bar{x}$ and $\bar{y}$ are the sample means.
An attempt was made to evaluate the inflation rate as a predictor of the spot rate in the German treasury bill market. For a sample of 79 quarterly observations, the estimated linear regression$$\hat{y}=0.0027+0.7916 x$$was obtained, where$y=$ actual change in the spot rate$x=$ change in the spot rate predicted by the inflation rateThe coefficient of determination was 0.097 , and the estimated standard deviation of the estimator of the slope of the population regression line was 0.2759 .a. Interpret the slope of the estimated regression line.b. Interpret the coefficient of determination.c. Test the null hypothesis that the slope of the population regression line is 0 against the alternative that the true slope is positive, and interpret your result.d. Test, against a two-sided alternative, the null hypothesis that the slope of the population regression line is 1 , and interpret your result.
The following table shows, for eight vintages of select wine, purchases per buyer $(y)$ and the wine buyer's rating in a year $(x)$ :$$\begin{array}{lllllllll}\hline x & 3.6 & 3.3 & 2.8 & 2.6 & 2.7 & 2.9 & 2.0 & 2.6 \\\hline y & 24 & 21 & 22 & 22 & 18 & 13 & 9 & 6 \\\hline\end{array}$$a. Estimate the regression of purchases per buyer on the buyer's rating.b. Interpret the slope of the estimated regression line.c. Find and interpret the coefficient of determination.d. Find and interpret a $90 \%$ confidence interval for the slope of the population regression line.e. Find a $90 \%$ confidence interval for expected purchases per buyer for a vintage for which the buyer's rating is 20.
For a sample of 306 students in a basic business statistics course, the sample regression line$$y=58.813+0.2875 x$$was obtained. Here,$y=$ final student score at the end of the course$x=$ score on a diagnostic statistics test given at the beginning of the courseThe coefficient of determination was 0.1158 , and the estimated standard deviation of the estimator of the slope of the population regression line was 0.04566 .a. Interpret the slope of the sample regression line.b. Interpret the coefficient of determination.c. The information given allows the null hypothesis that the slope of the population regression line is 0 to be tested in two different ways against the alternative that it is positive. Carry out these tests and show that they reach the same conclusion.
Based on a sample of 30 observations, the population regression model$$y_i=\beta_0+\beta_1 x_i+\varepsilon_i$$was estimated. The least squares estimates obtained were as follows:$$b_0=10.1 \text { and } b_1=8.4$$The regression and error sums of squares were as follows:$$S S R=128 \text { and } S S E=\mathbf{2 8 6}$$a. Find and interpret the coefficient of determination.b. Test at the $10 \%$ significance level against a twosided alternative the null hypothesis that $\beta_1$ is 0 .c. Find$$\sum_{i=1}^{30}\left(x_i-\bar{x}\right)^2$$
Based on a sample of 25 observations, the population regression model$$y_i=\beta_0+\beta_1 x_1+\varepsilon_i$$was estimated. The least squares estimates obtained were as follows:$$b_0=15.6 \text { and } b_1=1.3$$The total and error sums of squares were as follows:$$S S T=268 \text { and } S S E=204$$a. Find and interpret the coefficient of determination.b. Test, against a two-sided alternative at the $5 \%$ significance level, the null hypothesis that the slope of the population regression line is 0 .c. Find a $95 \%$ confidence interval for $\boldsymbol{\beta}_1$.
An analyst believes that the only important determinant of banks' returns on assets $(Y$ ) is the ratio of loans to deposits $(X)$. For a random sample of 20 banks, the sample regression line$$y=0.97+0.47 x$$was obtained with coefficient of determination 0.720 .a. Find the sample correlation between returns on assets and the ratio of loans to deposits.b. Test against a two-sided alternative at the $5 \%$ significance level the null hypothesis of no linear association between the returns and the ratio.
If a regression of the yield per acre of corn on the quantity of fertilizer used is estimated using fertilizer quantities in the range typically used by farmers, the slope of the estimated regression line will certainly be positive. However, it is well known that, if an enormously high amount of fertilizer is used, corn yield will be very low. Discuss the benefits of applying regression analysis to a data set that includes a few cases of excessive fertilizer use combined with data from typical operations.
A college's economics department is attempting to determine if verbal or mathematical proficiency is more important for predicting academic success in the study of economics. The department faculty have decided to use the grade point average (GPA) in economics courses for graduates as a measure of success. Verbal proficiency is measured by the SAT verbal and the ACT English entrance examination test scores. Mathematical proficiency is measured by the SAT mathematics and the ACT mathematics entrance examination scores. The data for 112 students are available in a data file named Student GPA. The designation of the variable columns is presented in the Chapter 11 appendix. You should use your local statistical computer program to perform the analysis for this problem.a. Prepare a graphical plot of the economics GPA versus each of the two verbal proficiency scores and each of the two mathematical proficiency scores. Which variable is a better predictor? Note any unusual patterns in the data.b. Compute the linear model coefficients and the regression analysis statistios for the models that predict economics GPA as a function of each verbal and each mathematics score. Using both the SAT mathematics and verbal measures and the ACT mathematics and English measures, determine whether mathematical or verbal proficiency is the best predictor of economics GPA.c. Compare the descriptive statistics-mean, standard deviation, upper and lower quartiles, and range-for the predictor variables. Note the differences and indicate how these differences affect the capability of the linear model to predict.
The administrator of the National Highway Traffic Safety Administration (NHTSA) wants to know if the different types of vehicles in a state have a relationship to the highway death rate in the state. She has asked you to perform several regression analyses to determine if average vehicle weight, percentage of imported cars, percentage of light trucks, or average car age is related to crash deaths in automobiles and pickups. The data for the analyses are in the data file Vehicle Travel State. The variable descriptions and locations are contained in the Chapter 11 appendix.a. Prepare graphical plots of crash deaths versus each of the potential predictor variables. Note the relationship and any unusual patterns in the data points.b. Prepare a simple regression analysis of crash deaths on the potential predictor variables. Determine which, if any, of the regressions indicate a significant relationship.c. State the results of your analysis and rank the predictor variables in terms of their relationship to crash deaths.
The Department of Transportation wishes to know if states with a larger percentage of urban population have higher rates of automobile and pickup crash deaths. In addition, it wants to know if either the average speed on rural roads or the percentage of rural roads that are surfaced is related to crash death rates. Data for this study are included in the data file Vehicle Travel State.a. Prepare graphical plots of crash deaths versus each of the potential predictor variables. Note the relationship and any unusual patterns in the data points.b. Prepare a simple regression analysis of crash deaths on the potential predictor variables. Determine which, if any, of the regressions indicate a significant relationship.c. State the results of your analysis and rank the predictor variables in terms of their relationship to crash deaths.
An economist wishes to predict the market value of owner-occupied homes in small midwestern cities. She has collected a set of data from 45 small cities for a 2-year period and wants you to use these as the data source for the analysis. The data are stored in the file Citydatr. She wants you to develop two prediction equations: one that uses the size of the house as a predictor and a second that uses the tax rate as a predictor.a. Plot the market value of houses (hseval) versus the size of houses (sizense), and then versus the tax rates (taxrate). Note any unusual patterns in the data.b. Prepare regression analyses for the two predictor variables. Which variable is the stronger predictor of the value of houses?c. A business developer in a midwestern state has stated that local property tax rates in small townsneed to be lowered because if they are not, no one will purchase a house in these towns. Based on your analysis in this problem, evaluate the business developer's claim.
Stuart Wainwright, the vice president of pur- chasing for a large national retailer, has asked you to prepare an analysis of retail sales by state. He wants to know if either the percent of male unemployment or the per capita disposable income are related to per capita retail sales. Data for this study are stored in the data file Economic Activity, which is described in the data file catalog in the Chapter 11 appendix. Note that you may have to compute new variables using those variables in the data file.a. Prepare graphical plots and regression analyses to determine the relationships between per capita retail sales and unemployment and personal income. Compute $95 \%$ confidence intervals for the slope coefficients in each regression equation.b. What is the effect of a $$\$ 1,000$$ decrease in per capita income on per capita sales?c. For the per capita income regression equation what is the $95 \%$ confidence interval for retail sales at the mean per capita income and at $$\$ 1,000$$ above the mean per capita income?
A major national supplier of building materials for residential construction is concerned about total sales for next year. It is well known that the company's sales are directly related to the total national residential investment. Several New York bankers are predicting that interest rates will rise about two percentage points next year. You have been asked to develop a regression analysis that can be used to predict the effect of interest rate changes on residential investment. The time series data for this study are contained in the data file Macro2010, which is described in the Chapter 13 appendix.a. Develop two regression models to predict residential investment, using the prime interest rate for one and the federal funds interest rate for the other. Analyze the regression statistics and indicate which equation provides the best predictions.b. Determine the $95 \%$ confidence interval for the slope coefficient in both regression equations.c. Based on each model, predict the effect of a twopercentage-point increase in interest rates on residential investment.d. Using both models, compute $95 \%$ confidence intervals for the change in residential investment that results from a two-percentage-point increase in interest rates.
A prestigious national news service has gathered information on a number of nationally ranked private colleges; these data are contained in the data file Private Colleges. You have been asked to determine if the student/faculty ratio has an influence on the quality rating. Note that the smallest number indicates the highest rank. Prepare and analyze this question using simple regression and a scatter plot. Prepare a short discussion of your conclusion.
A prestigious national news service has gathered information on a number of nationally ranked private colleges; these data are contained in the data file Private Colleges. You have been asked to determine if the student/faculty ratio has an influence on the total annual cost after need-based financial aid. Prepare and analyze this question using simple regression and a scatter plot. Prepare a short discussion of your conclusion.
A prestigious national news service has gathered information on a number of nationally ranked private colleges; these data are contained in the data file Private Colleges. You have been asked to determine if the total cost after need-based aid has an influence on average debt. Prepare and analyze this question using simple regression and a scatter plot. Prepare a short discussion of your conclusion.
A prestigious national news service has gathered information on a number of nationally ranked private colleges; these data are contained in the data file Private Colleges. You have been asked to determine if the percentage of students admitted has an influence on the 4-year graduation rate. Prepare and analyze this question using simple regression and a scatter plot. Prepare a short discussion of your conclusion.
A prestigious national news service has gathered information on a number of nationally ranked private colleges; these data are contained in the data file Private Colleges. You have been asked to determine if the student faculty ratio has an influence on the 4-year graduation rate. Prepare and analyze this question using simple regression and a scatter plot. Prepare a short discussion of your conclusion.
You have been asked to study the relationship between median income and poverty rate at the county level. After some investigation you determine that the data file Food Nutrition Atlas includes both these measures for county-level data. Perform an appropriate analysis and report your conclusions. Your analysis should include a regression of median income on poverty level and an appropriate scatter plot. Additional analysis would also prove helpful.
The federal nutrition guidelines prepared by the Center for Nutrition Policy and Promotion of the U.S. Department of Agriculture stress the importance of eating substantial servings of fruits and vegetables to obtain a healthy diet. You have been asked to determine if the per capita consumption of fruits and vegetables at the county level are related to the percentage of obese adults in the county. Data for this study are contained in the data file Food Nutrition Atlas, whose variable descriptions are found in the Chapter 9 appendix.
The federal nutrition guidelines prepared by the Center for Nutrition Policy and Promotion of the U.S. Department of Agriculture stress the importance of eating substantial servings of fruits and vegetables to obtain a healthy diet. You have been asked to determine if the per capita consumption of fruits and vegetables at the county level is related to the percentage of adults with diabetes in the county. Data for this study are contained in the data file Food Nutrition Atlas, whose variable descriptions are found in the Chapter 9 appendix.
The federal nutrition guidelines prepared by the Center for Nutrition Policy and Promotion of the U.S. Department of Agriculture stress the importance of eating reduced amounts of meat to obtain a healthy diet. You have been asked to determine if the per capita consumption of meat at the county level are related to the percentage of obese adults in the county. Data for this study are contained in the data file Food Nutrition Atlas, whose variable descriptions are found in the Chapter 9 appendix.
The federal nutrition guidelines prepared by the Center for Nutrition Policy and Promotion of the U.S. Department of Agriculture stress the importance of eating reduced amounts of meat to obtain a healthy diet. You have been asked to determine if the per capita consumption of meat at the county level are related to the percentage of adults with diabetes in the county. Data for this study are contained in the data file Food Nutrition Atlas, whose variable descriptions are found in the Chapter 9 appendix.
There is a belief among many people that a healthy diet will cost more than a less healthy diet. Using research based on the available population survey data, can you condude that a healthy diet will in fact cost more than a less healthy diet? Using the daily cost and the measure of $\mathrm{HEI}$, provide evidence to either accept or reject this general belief. You will do the analysis based first on the data from the first interview, creating subsets of the data file using daycode $=1$, and a second time using data from the second interview, creating subsets of the data file using daycode $=2$. Note differences in the results between the first and second interviews.
A group of social workers who work with lowincome people have argued that the poverty income ratio is directly related to the quality of an individual person's diet. That is, people with higher ratios will be more likely to have higher-quality diets, and those with lower ratios will have lower-quality diets. Perform an appropriate analysis to determine if their claim is supported by evidence. You will do the analysis based first on the data from the first interview, creating subsets of the data file using daycode $=1$, and a second time using data from the second interview, creating subsets of the data file using daycode $=2$ Note differences in the results between the first and second interviews.
A number of nutritionists have argued that fastfood restaurants have a negative effect on nutrition quality. In this exercise you are asked to determine if there is evidence to conclude that increasing the number of meals at fast-food restaurants will have a negative effect on diet quality. In addition, you are asked to determine the effect of eating in fast-food restaurants has on the daily cost of food. You will do the analysis based first on the data from the first interview, creating subsets of the data file using daycode $=1$, and a second time using data from the second interview, creating subsets of the data file using daycode $=2$. Note differences in the results between the first and second interviews.
In recent news commentaries, it has been argued that the quality of family life has decayed in recent years. Arguments include statements that families do not share meals together. Because of busy schedules, families just go out to eat because there is limited time for food preparation. What is the relationship between the percent of calories consumed at home and the quality of diet, based on an appropriate analysis of the survey data? In addition, what is the effect of eating at home on daily food cost? You will do the analysis based first on the data from the first interview, creating subsets of the data file using daycode $=1$, and a second time using data from the second interview, creating subsets of the data file using daycode $=2$. Note differences in the results between the first and second interviews.
In recent news commentaries, it has been argued that the quality of family life has decayed in recent years. Arguments include statements that families do not share meals together. Because of busy schedules, families just go out to eat because there is limited time for food preparation. In addition, it is also argued that a meal that is carefully prepared at home using purchased food ingredients will provide better nutrition. What is the relationship between the percent of calories purchased at a food store for consumption at home and the quality of diet, based on an appropriate analysis of the survey data? Also, what is the effect of percent of food purchased at a store on the daily food cost? You will do the analysis based first on the data from the first interview, creating subsets of the data file using daycode $=1$, and a second time using data from the second interview, creating subsets of the data file using daycode $=2$. Note differences in the results between the first and second interviews.