• Home
  • Textbooks
  • Statistics for Business and Economics: Global Edition
  • Additional Topics in Regression Analysis

Statistics for Business and Economics: Global Edition

Newbold P., Carlson W.L., Thorne B.M.

Chapter 13

Additional Topics in Regression Analysis - all with Video Answers

Educators


Chapter Questions

02:21

Problem 1

Write the model specification and define the variables for a multiple regression model to predict college GPA as a function of entering SAT scores and the year in college: freshman, sophomore, junior, and senior.

Brooke Bussoletti
Brooke Bussoletti
Numerade Educator
01:21

Problem 2

Write the model specification and define the variables for a multiple regression model to predict wages in U.S. dollars as a function of years of experience and country of employment, indicated as Germany, Great Britain, Japan, United States, and Turkey.

Hossam Mohamed
Hossam Mohamed
Numerade Educator

Problem 3

Write the model specification and define the variables for a multiple regression model to predict the cost per unit produced as a function of factory type (indicated as classic technology, computer-controlled machines, and computer-controlled material handling), and as a function of country (indicated as Colombia, South Africa, and Japan).

Check back soon!

Problem 4

An economist wants to estimate a regression equation relating demand for a product $(Y)$ to its price $\left(X_1\right)$ and income $\left(X_2\right)$. It is to be based on 12 years of quarterly data. However, it is known that demand for this product is seasonal; that is, it is higher at certain times of the year than others.
a. One possibility for accounting for seasonality is to estimate the model
$$
\begin{aligned}
y= & \beta_0+\beta_1 x_1+\beta_2 x_2+\beta_3 x_3+\beta_4 x_4+\beta_5 x_5 \\
& +\beta_6 x_6+\varepsilon
\end{aligned}
$$
where $x_3, x_4, x_5$, and $x_6$ are dummy variable values, with
$x_3=1$ in first quarter of each year, 0 otherwise $x_4=1$ in second quarter of each year, 0 otherwise
$x_5=1$ in third quarter of each year, 0 otherwise
$x_6=1$ in fourth quarter of each year, 0 otherwise
Explain why this model cannot be estimated by least squares.
b. For a model that can be estimated is as follows:
$$
\begin{aligned}
y= & \beta_0+\beta_1 x_1+\beta_2 x_2+\beta_3 x_3+\beta_4 x_4 \\
& +\beta_5 x_5+\varepsilon
\end{aligned}
$$
interpret the coefficients on the dummy variables in the model.

Check back soon!

Problem 5

Sharon Parsons, president of Gourmet Box Mini Pizza, has asked for your assistance in developing a model that predicts the demand for the new snack lunch pizza named Pizza1. This product competes in a market with three other brands that are named B2, B3, and B4 for identification. At present the products are sold by three major distribution chains, identified as 1,2 , and 3 . These three chains have different market sizes, and, thus, sales for each distributor are likely to be different. The data file Market contains weekly data collected over the past 52 weeks from the three distribution chains. The variables in the data file are defined next.
Use multiple regression to develop a model that predicts the quantity of Pizza1 sold per week by each distributor. The model should contain only important predictor variables.
$$
\begin{array}{ll}
\hline {}{}{\text { Distributor }} & \begin{array}{l}
\text { Numerical identifier of the distributor } \\
1,2, \text { or } 3
\end{array} \\
\hline \text { Weeknum } & \begin{array}{l}
\text { Sequential number of the week in which } \\
\text { data were collected }
\end{array} \\
\text { Sales Pizza1 } & \begin{array}{l}
\text { Number of units of Pizza1 sold by the } \\
\text { distributor during the week }
\end{array} \\
\text { Price Pizza1 } & \begin{array}{l}
\text { Retail price for Pizza1 charged by the } \\
\text { distributor during that week }
\end{array} \\
\text { Promotion } & \begin{array}{l}
\text { Level of promotion for the week, designated } \\
\text { as 0, no promotion; 1, television ad; } 2 \text { store } \\
\text { display; } 3 \text { both television and store display }
\end{array} \\
\text { Sales B2 } & \begin{array}{l}
\text { Number of units of brand } 2 \text { sold by the } \\
\text { distributor during the week }
\end{array} \\
\text { Price B2 } & \begin{array}{l}
\text { Retail price for brand } 2 \text { charged by the } \\
\text { distributor during that week }
\end{array} \\
\text { Sales B3 } & \begin{array}{l}
\text { Number of units of brand } 3 \text { sold by the } \\
\text { distributor during the week }
\end{array} \\
\text { Price B3 } & \begin{array}{l}
\text { Retail price for brand } 3 \text { charged by the } \\
\text { distributor during that week }
\end{array} \\
\text { Sales B4 } & \begin{array}{l}
\text { Number of units of brand } 4 \text { sold by the } \\
\text { distributor during the week }
\end{array} \\
\text { Price B4 } & \begin{array}{l}
\text { Retail price of brand } 4 \text { charged by the } \\
\text { distributor during that week }
\end{array}
\end{array}
$$

Check back soon!
View

Problem 6

John Ramapujan is the plant manager for Kitchen Products, Inc. He has asked you to help identify worker factors that influence productivity. In particular, he is interested in gender differences, the effect of working on different shifts, and employee attitudes toward the present benefits plan provided by the company. As a first step in your project you have collected the time required to complete the assembly of a new coffee grinder for a number of workers in the plant. In addition you have identified the workers, by gender (1-male, 2 -female), shift (1-day, 2-afternoon, 3-night), and How satisfied are you with employee benefits?
1 - Very dissatisfied
2 - Somewhat dissatisfied
3 - No opinion
4- Somewhat satisfied
5 - Very satisfied
The data collected are a file named Completion Times. Prepare an appropriate analysis and write a short report on the conclusions from your analysis.

Victor Salazar
Victor Salazar
Numerade Educator
03:33

Problem 7

You have been asked to develop a multiple regression model to predict per capita sales of cold cereal in cities with populations over 100,000 . As a first step you hold a meeting with the key marketing managers that have experience with cereal sales. From this meeting you discover that per capita sales are expected to be influenced by the cereal price, price of competing cereals, mean per capita income, percentage of college graduates, mean annual temperature, and mean annual rainfall. You also learn that the linear relationship between price and per capita sales is expected to have a different slope for cities east of the Mississippi River. Per capita sales are expected to be higher in cities with high and low per capita income compared to cities with intermediate per capita income. Per capita sales are also expected to be different in the following four sectors of the country: Northwest, Southwest, Northeast, Southeast.
Prepare a model specification whose coefficients can be estimated using multiple regression. Define each variable completely and indicate the mathematical form of the model. Discuss your specification, indicate which variables you expect to be statistically significant, and explain the rationale for your expectation.

James Kiss
James Kiss
Numerade Educator
03:05

Problem 8

Maxine Makitright, president of Good Parts, Ltd., has asked you to develop a model that predicts the number of defective parts per 8-hour work shift in her factory. She believes that there are differences among the three daily shifts and among the four raw-material suppliers. In addition, higher production and a higher number of workers are thought to be related to increased number of defectives. Maxine visits the factory at various times, including all three shifts, to observe operations and to offer operating advice. She has provided you with a list of the shifts that she has visited and wants to know if the number of defectives increases or decreases when she visits the factory.
Prepare a written description of how you would develop a model to estimate and test for the various factors that might influence the number of defective parts produced per shift. Carefully define each coefficient in your model and define the test you would use. Indicate how you would collect the data and how you would define each variable used in the model. Discuss the interpretations that you would make from your model specification.

Nick Johnson
Nick Johnson
Numerade Educator
02:35

Problem 9

Custom Woodworking, Inc., has been in business for 40 years. The company produces high-quality custommade wooden furniture and very high quality interior cabinet and interior woodwork for expensive homes and offices. It has been very successful in large part because of the highly skilled craftworkers, who design and produce its products in consultation with customers. Many of the company's products have won national awards for quality design and artisanship. Each custom-made product is produced by a team of two or more craftworkers who first meet with the customer, prepare an initial design, review the design with the customer, and then build the product. Customers may also meet with the craftworkers at various times during the production.
The craftworkers are well educated and have developed excellent woodworking skills. Most have liberal arts degrees and have trained with skilled craftworkers. Employees are classified at three levels: 1, apprentice; 2, professional; and 3, master. Levels 2 and 3 pay higher wages, and workers typically move through the levels as they gain experience and skill. The company now has a diverse workforce, which includes white, black, and Latino workers and both men and women. When the business started 40 years ago, all workers were white males. About 20 years ago the company began to hire black and Latino craftworkers, and about 10 years ago they hired women craftworkers. The white male workers tend to be overrepresented in the higher job classifications because, in part, they have the most experience. At present, the workforce contains $40 \%$ white males, $30 \%$ black and Latino males, $15 \%$ white females, and $15 \%$ black and Latino females.
Recently, serious concerns have been expressed concerning wage discrimination. Specifically, it is alleged that women and nonwhite workers are not receiving fair compensation based on their experience. The company management claims that every person is paid fairly based on years of experience, job classification level, and individual ability. It claims that there are no differences in wages based on either race or gender in terms of either base wage or increment for each year of experience.
Explain how you would carry out an analysis to determine if management's claim is true. Show the details of your analysis and provide a clear rationale. Indicate the data that should be collected and the names and descriptions of the variables you will use in the analysis. Clearly indicate the statistical tests that would be used to determine the true situation and indicate the decision rules based on the hypothesis tests and results from the data.

Katelyn Chen
Katelyn Chen
Numerade Educator

Problem 10

You have been asked to serve as a consultant and expert witness for a wage-discrimination lawsuit. A group of Latino and black women have filed the suit against their company, Amalgamated Distributors, Inc. The women, who have between 5 and 25 years of service
with the company, allege that the average rate of their annual wage increase has been significantly less than that of a group of white males and a group of white females. The jobs for all three groups contain a variety of administrative, analytical, and managerial components. All the employees began with a bachelor's degree, and years of experience is an important factor for predicting job performance and worker productivity. You have been provided with the present monthly wages and the years of experience for all workers in the three groups. In addition, the data indicate those in all three groups who have obtained an MBA degree. Note that you do not perform any data analysis for this problem.
a. Develop a statistical model and analysis that can be used to analyze the data. Indicate hypothesis tests that can be used to provide strong evidence of wage discrimination if wage discrimination exists. The company has also hired a statistician as a consultant and expert witness. Describe your analysis completely and clearly.
b. Assume that your hypothesis tests result in strong evidence that supports your clients' claim. Briefly summarize the key points that you will make in your expert witness testimony to the court. The company's lawyer can be expected to cross-examine you with the help of a statistician who teaches statistics at a prestigious liberal arts college.

Check back soon!
03:06

Problem 11

Consider the following models estimated using regression analysis applied to time-series data. What is the long-term effect of a 1-unit increase in $x$ in period $t$ ?
a. $y_t=10+2 x_t+0.34 y_{t-1}$
b. $y_t=10+2.5 x_t+0.24 y_{t-1}$
c. $y_t=10+2 x_t+0.64 y_{t-1}$
d. $y_t=10+4.3 x_t+0.34 y_{t-1}$

Adriano Chikande
Adriano Chikande
Numerade Educator
01:42

Problem 12

A market researcher is interested in the average amount of money spent per year by college students on clothing. From 25 years of annual data, the following estimated regression was obtained through least squares:
$$
\hat{y}_t=50.72+\underset{(0.047)}{0.142 x_{1 t}}+\underset{(0.0021)}{0.027 x_{2 t}}+\underset{(0.136)}{0.432 y_{t-1}}
$$
where
$y=$ expenditure per student, in dollars, on clothes
$x_1=$ disposable income per student, in dollars, after the payment of tuition, fees, and room and board $x_2=$ index of advertising, aimed at the student market, on clothes

The numbers in parentheses below the coefficients are the coefficient standard errors.
a. Test, at the $5 \%$ level against the obvious one-sided alternative, the null hypothesis that, all else being equal, advertising does not affect expenditures on clothes in this market.
b. Find a $95 \%$ confidence interval for the coefficient on $x_1$ in the population regression.
c. With advertising held fixed, what would be the expected impact over time of a $$\$ 1$$ increase in disposable income per student on clothing expenditure?

Adriano Chikande
Adriano Chikande
Numerade Educator
02:12

Problem 13

Use the data from the Retail Sales file to estimate the regression model
$$
y_t=\beta_0+\beta_1 x_t+\gamma y_{t-1}+\varepsilon_t
$$
and test the null hypothesis that $\gamma=0$, where
$y_t=$ retail sales per household
$x_t=$ disposable income per household

Adriano Chikande
Adriano Chikande
Numerade Educator
01:32

Problem 14

The data file Money UK contains observations from the United Kingdom on the quantity of money in millions of pounds $(Y)$; income, in millions of pounds $\left(X_1\right)$; and the local authority interest rate $\left(X_2\right)$. Estimate the model (Mills 1978)
$$
y_t=\beta_0+\beta_1 x_{1 t}+\beta_2 x_{2 t}+\gamma y_{t-1}+\varepsilon_t
$$
and write a report on your findings.

Adriano Chikande
Adriano Chikande
Numerade Educator
01:34

Problem 15

The data file Pension Funds contains data on the market retum $(X)$ of stocks and the percentage $(Y)$ of portfolios in common stocks at market value at the end of the year for private pension funds. Estimate the model
$$
y_t=\beta_0+\beta_1 x_t+\gamma y_{t-1}+\varepsilon_t
$$
and write a report on your findings.

Adriano Chikande
Adriano Chikande
Numerade Educator
01:32

Problem 16

The data file Income Canada shows quarterly observations on income $(Y)$ and money supply $(X)$ in Canada. Estimate the model (Hsiao 1979)
$$
y_t=\beta_0+\beta_1 x_t+\gamma y_{t-1}+\varepsilon_t
$$
and write a report on your findings.

Adriano Chikande
Adriano Chikande
Numerade Educator
01:28

Problem 17

The data file Births Australia shows annual observations on the first confinement resulting in a live birth of the current marriage $(Y)$ and the number of first marriages (for females) in the previous year $(X)$ in Australia. Estimate the model (McDonald 1981)
$$
y_t=\beta_0+\beta_1 x_t+\gamma y_{t-1}+\varepsilon_t
$$
and write a report on your findings.

Adriano Chikande
Adriano Chikande
Numerade Educator
02:11

Problem 18

The data file Thailand Consumption shows 29 annual observations on private consumption $(Y)$ and disposable income $(X)$ in Thailand. Fit the regression model
$$
\log y_t=\beta_0+\beta_1 \log x_{1 t}+\gamma \log y_{t-1}+\varepsilon_t
$$
and write a report on your findings.

Adriano Chikande
Adriano Chikande
Numerade Educator
00:51

Problem 19

Suppose that the true linear model for a process was
$$
Y=\beta_0+\beta_1 X_1+\beta_2 X_2+\beta_3 X_3
$$
and you incorrectly estimated the model
$$
\gamma=\alpha_0+\alpha_1 X_2
$$
Interpret and contrast the coefficients for $X_2$ in the two models. Show the bias that results from using the second model.

Victor Salazar
Victor Salazar
Numerade Educator
00:51

Problem 20

Suppose that a regression relationship is given by the following:
$$
\gamma=\beta_0+\beta_1 X_1+\beta_2 X_2+\varepsilon
$$
If the simple linear regression of $Y$ on $X_1$ is estimated from a sample of $n$ observations, the resulting slope estimate is generally biased for $\beta_1$. However, in the special case where the sample correlation between $X_1$ and $X_2$ is 0 , this will not be so. In fact, in that case the same estimate results whether or not $X_2$ is included in the regression equation.
a. Explain verbally why this statement is true.
b. Show algebraically that this statement is true.

Victor Salazar
Victor Salazar
Numerade Educator
02:44

Problem 21

Transportation Research, Inc., has asked you to prepare some multiple regression equations to estimate the effect of variables on fuel economy. The data for this study are contained in the data file Motors, and the dependent variable is miles per gallon-milpgal-as established by the Department of Transportation certification.
a. Prepare a regression equation that uses vehicle horsepower-horspwer-and vehicle weight-weight-as independent variables. Interpret the coefficients.
b. Prepare a second biased regression with vehicle weight not included. What can you conclude about the coefficient of horsepower?

Adriano Chikande
Adriano Chikande
Numerade Educator

Problem 22

Use the data in the file Citydatr to estimate a regression equation that can be used to determine the marginal effect of the percent commercial property on the market value per owner-occupied residence (Hseval). Include the percent of owneroccupied residences (Homper), percent of industrial property (Indper), the median rooms per residence (sizehse), and per capita income (Incom 72) as additional predictor variables in your multiple regression equation. The variables are described in the Chapter 12 appendix. Indicate which of the variables are conditionally significant. Your final equation should include only significant variables. Run a second regression with median rooms per residence excluded. Interpret the new coefficient for percent commercial property that results from the second regression. Compare the two coefficients.

Check back soon!
View

Problem 23

In the regression model
$$
Y=\beta_0+\beta_1 X_1+\beta_2 X_2+\varepsilon
$$
the extent of any multicollinearity can be evaluated by finding the correlation between $X_1$ and $X_2$ in the sample. Explain why this is so.

Shu Naito
Shu Naito
Numerade Educator

Problem 24

An economist estimates the following regression model:
$$
y=\beta_0+\beta_1 x_1+\beta_2 x_2+\varepsilon
$$
The estimates of the parameters $\beta_1$ and $\beta_2$ are not very large compared with their respective standard errors. But the size of the coefficient of determination indicates quite a strong relationship between the dependent variable and the pair of independent variables. Having obtained these results, the economist strongly suspects the presence of multicollinearity. Since his chief interest is in the influence of $X_1$ on the dependent variable, he decides that he will avoid the problem of multicollinearity by regressing $Y$ on $X_1$ alone. Comment on this strategy.

Check back soon!
04:11

Problem 25

Based on data from 63 counties, the following model was estimated by least squares:
$$
\hat{y}=0.58-\underset{(0.019)}{0.052 x_1}-\underset{(0.042)}{0.005 x_2} \quad R^2=0.17
$$
where
$\hat{y}=$ growth rate in real gross domestic product
$x_1=$ real income per capita
$x_2=$ average tax rate, as a proportion of gross national product
The numbers below the coefficients are the coefficient standard errors. After the independent variable $X_1$, real income per capita, was dropped from the model, the regression of growth rate in real gross domestic product on $X_2$, average tax rate, was estimated. This yielded the following fitted model:
$$
\hat{y}=0.060-\underset{(0.034)}{0.074 x_2} \quad R^2=0.072
$$
Comment on this result.

Heather Duong
Heather Duong
Numerade Educator
02:12

Problem 26

In Chapter 11, the regression of retail sales per household on disposable income per household was estimated by least squares. The data are given in Table 11.1, and Table 11.2 shows the residuals and the predicted values of the dependent variable. Use the data file Retail Sales.
a. Graphically check for heteroscedasticity in the regression errors.
b. Check for heteroscedasticity by using a formal test.

Adriano Chikande
Adriano Chikande
Numerade Educator
01:32

Problem 27

Consider a regression model that uses 48 observations. Let $e_i$ denote the residuals from the fitted regression and $\hat{y}_i$ be the in-sample predicted values of the dependent variable. The least squares regression of $e_i^2$ on $\hat{y}_i$ has coefficient of determination 0.032 . What can you conclude from this finding?

Adriano Chikande
Adriano Chikande
Numerade Educator
03:01

Problem 28

The data file Economic Activity contains data for 50 states in the United States. Develop a multiple regression model to predict total retail sales for auto parts and dealers. Find two or three of the best predictor variables from those in the data file using the variable descriptions from the Chapter 11 appendix.
a. Compute the multiple regression model using the predictor variables selected.
b. Graphically check for heteroscedasticity in the regression errors.
c. Use a formal test to check for heteroscedasticity.

James Kiss
James Kiss
Numerade Educator
02:24

Problem 29

You have been asked by East Anglica Realty, Ltd., to provide a linear model that will estimate the selling price of homes as a function of family. There is particular concern for obtaining the most efficient estimate of the relationship between income and house price. East Anglica has collected data on their sales experience over the past 5 years, and the data are contained in the file East Anglica Realty, Ltd.
a. Estimate the regression of house price on family income.
b. Graphically check for heteroscedasticity.
c. Use a formal test of hypothesis to check for heteroscedasticity.
d. If you establish that there is heteroscedasticity in (b) and (c), perform another regression that corrects for heteroscedasticity.

Adriano Chikande
Adriano Chikande
Numerade Educator

Problem 30

Consider the following regression model:
$$
y_t=\beta_0+\beta_1 x_{1 t}+\beta_2 x_{2 t}+\cdots+\beta_K x_{K t}+\varepsilon_t
$$
Show that if
$$
\operatorname{Var}(\varepsilon)=K x_i^2 \quad(K>0)
$$
then
$$
\operatorname{Var}\left[\frac{\varepsilon_i}{x_i}\right]=K
$$
Discuss the possible relevance of this result in treating a form of heteroscedasticity.

Check back soon!
01:32

Problem 31

Refer to Exercise 13.14 and data file Money UK. Let $e_i$ denote the residuals from the fitted regression and $\hat{y}_i$ be the in-sample predicted values. The least squares regression of $e_i^2$ on $\hat{y}_i$ has coefficient of determination of 0.087 . What can you conclude from this finding?
Let $e_i$ denote the residuals from the fitted regression and $\hat{y}_i$ be the in-sample predicted values. Estimate the least squares regression of $e_i^2$ on $\hat{y}_i$ and compute the coefficient of determination. What can you conclude from this finding?

Adriano Chikande
Adriano Chikande
Numerade Educator

Problem 32

Suppose that a regression was run with three independent variables and 30 observations. The DurbinWatson statistic was 0.50 . Test the hypothesis that there was no autocorrelation. Compute an estimate of the autocorrelation coefficient if the evidence indicates that there was autocorrelation.
a. Repeat with the Durbin-Watson statistic equal to 0.80 .
b. Repeat with the Durbin-Watson statistic equal to 1.10 .
c. Repeat with the Durbin-Watson statistic equal to 1.25 .
d. Repeat with the Durbin-Watson statistic equal to 1.70 .

Check back soon!

Problem 33

Suppose that a regression was run with two independent variables and 28 observations. The Durbin-Watson statistic was 0.50 . Test the hypothesis that there was no autocorrelation. Compute an estimate of the autocorrelation coefficient if the evidence indicates that there was autocorrelation.
a. Repeat with the Durbin-Watson statistic equal to 0.80 .
b. Repeat with the Durbin-Watson statistic equal to 1.10 .
c. Repeat with the Durbin-Watson statistic equal to 1.25 .
d. Repeat with the Durbin-Watson statistic equal to 1.70 .

Check back soon!

Problem 34

In a regression based on 30 annual observations, U.S. farm income was related to four independent variables-grain exports, federal government subsidies, population, and a dummy variable for bad weather years. The model was fitted by least squares, resulting in a Durbin-Watson statistic of 1.29. The regression of $e_i^2$ on $\hat{y}_i$ yielded a coefficient of determination of 0.043 .
a. Test for heteroscedasticity.
b. Test for autocorrelated errors.

Check back soon!
02:11

Problem 35

The data file Money UK contains observations from the United Kingdom on the quantity of money in millions of pounds $(Y)$; income, in millions of pounds $\left(X_1\right)$; and the local authority interest rate $\left(X_2\right)$. Estimate the model (Mills 1978)
$$
y_t=\beta_0+\beta_1 x_{1 t}+\beta_2 x_{2 t}+\gamma y_{t-1}+\varepsilon_t
$$
and write a report on your findings.
What can be concluded from the Durbin-Watson statistic for the fitted regression?

Adriano Chikande
Adriano Chikande
Numerade Educator
02:11

Problem 36

The data file Thailand Consumption shows 29 annual observations on private consumption $(Y)$ and disposable income $(X)$ in Thailand. Fit the regression model
$$
\log y_t=\beta_0+\beta_1 \log x_{1 t}+\gamma \log y_{t-1}+\varepsilon_t
$$
and write a report on your findings.
Test the null hypothesis of no autocorrelated errors against the alternative of positive autocorrelation.

Adriano Chikande
Adriano Chikande
Numerade Educator
02:25

Problem 37

A factory operator hypothesized that his unit output costs (y) depend on wage rate $\left(x_1\right)$, other input costs $\left(x_2\right)$, overhead costs $\left(x_3\right)$, and advertising expenditures $\left(x_4\right)$. A series of 24 monthly observations was obtained, and a least squares estimate of the model yielded the following results:
$$
\begin{gathered}
\hat{y}_i=0.75+\underset{(0.07)}{0.24 x_{1 t}}+\underset{(0.12)}{0.56 x_{2 t}}-\underset{(0.23)}{0.32 x_{3 t}}+\underset{(0.5)}{0.23 x_{4 t}} \\
R^2=0.79 \quad d=0.85
\end{gathered}
$$
The figures in parentheses below the estimated coefficients are their estimated standard errors. What can you conclude from these results?

Akash M
Akash M
Numerade Educator
02:12

Problem 38

The data file Advertising Retail shows, for a consumer goods corporation, 22 consecutive years of data on sales $(y)$ and advertising $(x)$.
a. Estimate the regression:
$$
y_t=\beta_0+\beta_1 x_t+\varepsilon_t
$$
b. Check for autocorrelated errors in this model.
c. If necessary, re-estimate the model, allowing for autocorrelated errors.

Adriano Chikande
Adriano Chikande
Numerade Educator
03:48

Problem 39

The omission of an important independent variable from a time-series regression model can result in the appearance of autocorrelated errors. In Example 13.7 we estimated the model
$$
y_t=\beta_0+\beta_1 x_{1 t}+\varepsilon_t
$$
relating profit margin to net revenue per dollar for our savings and loan data. Carry out a Durbin-Watson test on the residuals from this model. What can you infer from the results?

Heather Duong
Heather Duong
Numerade Educator
02:43

Problem 40

Write brief reports, including examples, explaining the use of each of the following in specifying regression models:
a. Dummy variables
b. Lagged dependent variables
c. The logarithmic transformation

Sheryl Ezze
Sheryl Ezze
Numerade Educator
02:35

Problem 41

Consider the fitting of the following model:
$$
Y=\beta_0+\beta_1 X_1+\beta_2 X_2+\beta_3 X_3+\varepsilon
$$
where
$Y=$ tax revenues as a percentage of gross national product in a country
$X_1=$ exports as a percentage of gross national product in the country
$X_2=$ income per capita in the country
$X_3=$ dummy variable taking the value 1 if the country participates in some form of economic integration, 0 otherwise
This provides a means of allowing for the effects on tax revenue of participation in some form of economic integration. Another possibility would be to estimate the regression
$$
Y=\beta_0+\beta_1 X_1+\beta_2 X_2+\varepsilon
$$
separately for countries that did and did not participate in some form of economic integration. Explain how these approaches to the problem differ.

Nick Johnson
Nick Johnson
Numerade Educator
05:34

Problem 42

Discuss the following statement: In many practical regression problems, multicollinearity is so severe that it would be best to run separate simple linear regressions of the dependent variable on each independent variable.

Jennifer Stoner
Jennifer Stoner
Numerade Educator
01:28

Problem 43

Explain the nature of and the difficulties caused by each of the following:
a. Heteroscedasticity
b. Autocorrelated errors

Savannah Langenstein
Savannah Langenstein
Numerade Educator

Problem 44

The following model was fitted to data on 90 German chemical companies:
$$
\begin{aligned}
\hat{y}= & 0.819+\underset{(1.79)}{2.11 x_1}+\underset{(1.94)}{0.96 x_2}-\underset{(0.144)}{0.059 x_3}+\underset{(4.08)}{5.87 x_4} \\
& +\underset{\left(0.00226 x_5\right.}{0.0025)} \quad \bar{R}^2=.410
\end{aligned}
$$
where the numbers in parentheses are estimated coefficient standard errors and
$$
\begin{aligned}
y & =\text { share price } \\
x_1 & =\text { earnings per share } \\
x_2 & =\text { funds flow per share } \\
x_3 & =\text { dividends per share } \\
x_4 & =\text { book value per share } \\
x_5 & =\text { a measure of growth }
\end{aligned}
$$
a. Test at the $10 \%$ level the null hypothesis that the coefficient on $x_1$ is 0 in the population regression against the alternative that the true coefficient is positive.
b. Test at the $10 \%$ level the null hypothesis that the coefficient on $x_2$ is 0 in the population regression against the alternative that the true coefficient is positive.
c. The variable $X_2$ was dropped from the original model, and the regression of $Y$ on $\left(X_1, X_3, X_4, X_5\right)$ was estimated. The estimated coefficient on $X_1$ was 2.95 with standard error 0.63 . How can this result be reconciled with the conclusion of part a?

Check back soon!

Problem 45

The following model was fitted to data from 28 countries in 1989 in order to explain the market value of their debt at that time:
$$
\begin{aligned}
& \hat{y}=77.2-\underset{(8.0)}{9.6 x_1}-\underset{(2.73)}{17.2 x_2}-\underset{(0.056)}{0.15 x_3}+\underset{(1.0)}{2.2 x_4} \\
& R^2=0.84
\end{aligned}
$$
$(8.0)$
(273)
$(0056)$
$(1.0)$
$$
R^2=0.84
$$
where
$$
\begin{aligned}
& y= \text { secondary market price, in dollars, in } 1989 \\
& \text { of } \$ 100 \text { of the country's debt } \\
& x_1= 1 \text { if U.S. bank regulators have mandated } \\
& \text { write-down for the country's assets on books } \\
& \text { of U.S. banks, } 0 \text { otherwise } \\
& x_2= 1 \text { if the country suspended interest payments } \\
& \text { in } 1989,2 \text { if the country suspended interest } \\
& \text { payments before } 1989 \text { and was still in suspension, } \\
& \text { and } 0 \text { otherwise } \\
& x_3= \text { debt-to-gross-national-product ratio } \\
& x_4= \text { rate of real gross national product growth, } \\
& 1980-1985
\end{aligned}
$$
The numbers below the coefficients are the coefficient standard errors.
a. Interpret the estimated coefficient on $x_1$.
b. Test the null hypothesis that, all else being equal, debt-to-gross-national-product ratio does not linearly influence the market value of a country's debt against the alternative that the higher this ratio, the lower the value of the debt.
c. Interpret the coefficient of determination.
d. The specification of the dummy variable $x_2$ is unorthodox. An alternative would be to replace $x_2$ by the pair of variables $\left(x_5, x_6\right)$, defined as follows:
$x_5=1$ if the country suspended interest payments in 1989, 0 otherwise
$x_6=1$ if the country suspended interest payments before 1989 and was still in suspension, 0 otherwise
Compare the implications of these two alternative specifications.

Check back soon!
01:42

Problem 46

An attempt was made to construct a regression model explaining student scores in intermediate economics courses (Waldauer, Duggal, and Williams 1992). The population regression model assumed that
$Y=$ total student score in intermediate economics courses
$X_1=$ mathematics score on Scholastic Aptitude Test
$X_2=$ verbal score on Scholastic Aptitude Test
$X_3=$ grade in college algebra $(A=4, B=3, C=2$, $D=1$ )
$X_4=$ grade in college principles of economics course
$X_5=$ dummy variable taking the value 1 if the student is female and 0 if male
$X_6=$ dummy variable taking the value 1 if the instructor is male and 0 if female
$X_7=$ dummy variable taking the value 1 if the student and instructor are the same gender and 0 otherwise
This model was fitted to data on 262 students. Next we report $t$-ratios, so that $t_j$ is the ratio of the estimate of $\beta_j$ to its associated estimated standard error. These ratios are as follows:
$$
\begin{aligned}
& t_1=4.69, t_2=2.89, t_3=0.46, t_4=4.90, \\
& t_5=0.13, t_6=-1.08, t_7=0.88
\end{aligned}
$$
The objective of this study was to assess the impact of the gender of student and instructor on performance. Write a brief report outlining what has been learned about this issue.

Adriano Chikande
Adriano Chikande
Numerade Educator
07:39

Problem 47

The following regression was fitted by least squares to 32 annual observations on time-series data:
$$
\begin{aligned}
\log y_t= & 4.52-\underset{(0.28)}{0.62 \log x_{1 t}}+\underset{(0.38)}{0.92} \log x_{2 t}+\underset{(0.21)}{0.61 \log x_{3 t}} \\
& +\underset{(0.12)}{0.16 \log x_{4 t}+e_t} \quad \bar{R}^2=0.683 \quad d=0.61
\end{aligned}
$$
where
$y_t=$ quantity of U.S. wheat exported
$x_{1 t}=$ price of U.S. wheat on world market
$x_{2 t}=$ quantity of U.S. wheat harvested
$x_{3 t}=$ measure of income in countries importing U.S. wheat
$x_{4 \pm}=$ price of barley on world market
The numbers below the coefficients are the coefficient standard errors.
a. Interpret the estimated coefficient on $\log x_{1 t}$ in the context of the assumed model.
b. Test at the $5 \%$ level the null hypothesis that, all else being equal, income in importing countries has no effect on U.S. wheat exports against the alternative that higher income leads to higher expected exports. (Ignore, for now, the Durbin-Watson $d$ statistic.)
c. What null hypothesis can be tested by the $d$ statistic? Carry out this test for the present problem, using a $1 \%$ significance level.
d. In view of your finding in part c, comment on your conclusion in part $b$. How might you proceed to test the null hypothesis of part $b$ ?

Heather Duong
Heather Duong
Numerade Educator
04:46

Problem 48

The following regression was fitted by least squares to 30 annual observations on time-series data:
$$
\begin{aligned}
& \log y_t=4.31-\underset{(0.17)}{0.27} \log x_{1 t}+\underset{(0.21)}{0.53} \log x_{2 t} \\
& -\underset{(0.30)}{0.82} \log x_{3 t}+e_t \quad \bar{R}^2=0.615 \quad d=.49 \\
&
\end{aligned}
$$
where
$$
\begin{aligned}
y_t & =\text { number of business failures } \\
x_{1 t} & =\text { rate of unemployment } \\
x_{2 t} & =\text { short-term interest rate } \\
x_{3 t} & =\text { value of new business orders placed }
\end{aligned}
$$
The numbers below the coefficients are the coefficient standard errors.
a. Interpret the estimated coefficient on $\log x_{3 t}$ in the context of the assumed model.
b. What null hypothesis can be tested by the $d$ statistic? Carry out this test for the present problem using a $1 \%$ significance level.
c. Given your results in part b, is it possible to test, with the information given, the null hypothesis that, all else being equal, short-term interest rates do not influence business failures?
d. Estimate the correlation between adjacent error terms in the regression model.

Heather Duong
Heather Duong
Numerade Educator

Problem 49

A stockbroker is interested in the factors influencing the rate of return on the common stock of banks. For a sample of 30 banks, the following regression was estimated by least squares:
$$
\begin{aligned}
& \hat{y}=2.37+\underset{(0.39)}{0.84 x_1}+\underset{(0.12)}{0.15 x_2}-\underset{(0.09)}{0.13 x_3} \\
& +1.67 x_4 \quad R^2=0.317 \\
&
\end{aligned}
$$
where
$y=$ percentage rate of return on common stock of bank
$x_1=$ percentage rate of growth of bank's earnings
$x_2=$ percentage rate of growth of bank's assets
$x_3=$ loan losses as percentage of bank's assets
$x_4=1$ if bank head office is in New York City and 0 otherwise
The numbers below the coefficients are the coefficient standard errors.
a. Interpret the estimated coefficient on $x_4$.
b. Interpret the coefficient of determination, and use it to test the null hypothesis that, taken as a group, the four independent variables do not linearly influence the dependent variable.
c. Let $e_i$ denote the residuals from the fitted regression and $\hat{y}_i$ the in-sample predicted values of the dependent variable. The least squares regression of $e_i^2$ on $\hat{y}_i$ yielded coefficient of determination 0.082 . What can be concluded from this finding?

Check back soon!
01:42

Problem 50

A market researcher is interested in the average amount of money per year spent by students on entertainment. From 30 years of annual data, the following regression was estimated by least squares:
$$
\hat{y}_t=40.93+\underset{(0.106)}{0.253 x_t}+\underset{(0.134)}{0.546 y_{t-1}} \quad d=1.86
$$
where
$$
\begin{aligned}
y_t= & \text { expenditure per student, in dollars, on } \\
& \text { entertainment } \\
x_t= & \text { disposable income per student, in dollars, after } \\
& \text { payment of tuition, fees, and room and board }
\end{aligned}
$$
The numbers below the coefficients are the coefficient standard errors.
a. Find a $95 \%$ confidence interval for the coefficient on $x_t$ in the population regression.
b. What would be the expected impact over time of a $\$ 1$ increase in disposable income per student on entertainment expenditure?
c. Test the null hypothesis of no autocorrelation in the errors against the alternative of positive autocorrelation.

Adriano Chikande
Adriano Chikande
Numerade Educator
View

Problem 51

A local public utility would like to be able to predict a dwelling unit's average monthly electricity bill. The company statistician estimated by least squares the following regression model:
$$
y_t=\beta_0+\beta_1 x_{1 t}+\beta_2 x_{2 t}+\varepsilon_t
$$
where
$$
\begin{aligned}
y_t= & \text { average monthly electricity bill, in dollars } \\
x_{1 t}= & \text { average bimonthly automobile gasoline bill, } \\
& \text { in dollars } \\
x_{2 t}= & \text { number of rooms in dwelling unit }
\end{aligned}
$$
From a sample of 25 dwelling units, the statistician obtained the following output from the SAS program:
$$
\begin{array}{crcc}
\hline \text { Parameter } & \text { Estimate } & \begin{array}{c}
\text { Student's } t \text { for } H_0: \\
\text { parameter }=0
\end{array} & \begin{array}{c}
\text { Std. error } \\
\text { of estimate }
\end{array} \\
\hline \text { Intercept } & -10.8030 & & \\
x_1 & -0.0247 & -0.956 & 0.0259 \\
x_2 & 10.9409 & 18.517 & 0.5909 \\
\hline
\end{array}
$$
a. Interpret, in the context of the problem, the least squares estimate of $\beta_2$.
b. Test, against a two-sided alternative, the null hypothesis
$$
H_0: \beta_1=0
$$
c. The statistician is concerned about the possibility of multicollinearity. What information is needed to assess the potential severity of this problem?
d. It is suggested that household income is an important determinant of size of electricity bill. If this is so, what can you say about the regression estimated by the statistician?
e. Given the fitted model, the statistician obtains the predicted electricity bills, $\hat{y}_t$, and the residuals, $e_t$. He then regresses $\hat{e}_t^2$ on $\hat{y}_t$, finding that the regression has a coefficient of determination of 0.0470 . Interpret this finding.

Shu Naito
Shu Naito
Numerade Educator
02:11

Problem 52

The data file Indonesia Revenue show 15 annual observations from Indonesia on total government tax revenues other than from oil (y), national income $\left(x_1\right)$, and the value added by oil as a percentage of gross domestic product $\left(x_2\right)$. Estimate by least squares the following regression:
$$
\log y_t=\beta_0+\beta_1 \log x_{1 t}+\beta_2 \log x_{2 t}+\varepsilon_t
$$
Write a report summarizing your findings, including a test for autocorrelated errors.

Adriano Chikande
Adriano Chikande
Numerade Educator
02:11

Problem 53

The data file German Income shows 22 annual observations from the Federal Republic of Germany on percentage change in wages and salaries $(y)$, productivity growth $\left(x_1\right)$, and the rate of inflation $\left(x_2\right)$, as measured by the gross national product price deflator. Estimate by least squares the following regression:
$$
y_t=\beta_0+\beta_1 x_{1 \mathrm{t}}+\beta_2 x_{2 \mathrm{t}}+\varepsilon_t
$$
Write a report summarizing your findings, including a test for heteroscedasticity and a test for autocorrelated errors.

Adriano Chikande
Adriano Chikande
Numerade Educator
02:11

Problem 54

The data file Japan Imports shows 35 quarterly observations from Japan on quantity of imports (y), ratio of import prices to domestic prices $\left(x_1\right)$, and real gross national product $\left(x_2\right)$. Estimate by least squares the following regression:
$$
\log y_t=\beta_0+\beta_1 \log x_{1 t}+\beta_2 \log x_{2 t}+\gamma \log y_{t-1}+\varepsilon_t
$$
Write a report summarizing your findings, including a test for autocorrelated errors.

Adriano Chikande
Adriano Chikande
Numerade Educator
01:10

Problem 55

A study was conducted on the labor-hour costs of Federal Deposit Insurance Corporation (FDIC) audits of banks. Data were obtained on 91 such audits. Some of these were conducted by the FDIC alone and some jointly with state auditors. Auditors rated banks' management as good, satisfactory, fair, or unsatisfactory. The model estimated was $$
\begin{aligned}
& \log y=2.41+\underset{0.0477}{0.3674} \log x_1+\underset{(0.0628)}{0.2217 \log x_2} \\
& +0.0803 \log x_3-0.1755 x_4+0.2799 x_5 \\
& \begin{array}{lll}
(0.028) & (0.2905) & (0.1044)
\end{array} \\
& +\underset{(0.1657)}{0.5634 x_6}-\underset{(0.0787)}{0.2572 x_7}+e \quad R^2=0.766 \\
&
\end{aligned}
$$
where
$y=$ FDIC auditor labor-hours
$x_1=$ total assets of bank
$x_2=$ total number of offices in bank
$x_3=$ ratio of classified loans to total loans for bank
$x_4=1$ if management rating was "good," 0 otherwise
$x_5=1$ if management rating was "fair," 0 otherwise
$x_6=1$ if management rating was "unsatisfactory," 0 otherwise
$x_7=1$ if audit was conducted jointly with the state, 0 otherwise
The numbers in parentheses beneath coefficient estimates are the associated standard errors. Write a report on these results.

Dominador Tan
Dominador Tan
Numerade Educator

Problem 56

The data file Britain Sick Leave shows data from Great Britain on the days of sick leave per person $(Y)$, unemployment rate $\left(X_1\right)$, ratio of benefits to earnings $\left(X_2\right)$, and the real wage rate $\left(X_3\right)$. Estimate the model
$$
\log y_t=\beta_0+\beta_1 \log x_{1 t}+\beta_2 \log x_{2 t}+\beta_3 \log x_{3 t}+\varepsilon_t
$$
and write a report on your findings. Include in your analysis a check on the possibility of autocorrelated errors and, if necessary, a correction for this problem.

Check back soon!
04:46

Problem 57

The U.S. Department of Commerce has asked you to develop a regression model to predict quarterly investment in production and durable equipment. The suggested predictor variables include GDP, prime interest rate, per capita income lagged, federal government spending, and state and local government spending. The data for your analysis are found in the data file Macro2010, which is described in the data dictionary in the chapter appendix. Use data from the time period 1980.1 through 2010.4.
a. Estimate a regression model using only interest rate to predict the investment. Use the DurbinWatson statistic to test for autocorrelation.
b. Find the best multiple regression equation to predict investment using the predictor variables previously indicated. Use the Durbin-Watson statistic to test for autocorrelation.
c. What are the differences between the regression models in parts a and b in terms of goodness of fit, prediction capability, autocorrelation, and contributions to understanding the investment problem?

Heather Duong
Heather Duong
Numerade Educator

Problem 58

An economist has asked you to develop a regression model to predict consumption of service goods as a function of disposable personal income and other important variables. The data for your analysis are found in the data file Macro2010, which is described in the data dictionary in the chapter appendix. Use data from the period 1980.1 through 2010.4.
a. Estimate a regression model using only disposable personal income to predict consumption of service goods. Test for autocorrelation using the DurbinWatson statistic.
b. Estimate a multiple regression model using disposable personal income, total consumption lagged 1 period, and prime interest rate as additional predictors. Test for autocorrelation. Does this multiple regression model reduce the problem of autocorrelation?

Check back soon!

Problem 59

Jack Wong, a Tokyo investor, is considering plans to develop a primary steel plant in Japan. After reviewing the initial design proposal, he is concerned about the proposed mix of capital and labor. He has asked you to prepare several production functions using some historical data from the United States. The data file Metals contains 27 observations of the value-added output, labor input, and gross value of plant and equipment per factory.
a. Use multiple regression to estimate a linear production function with value-added output regressed on labor and capital.
b. Plot the residuals versus labor and equipment. Note any unusual patterns.
c. Use multiple regression with transformed variables to estimate a Cobb-Douglas production function of the form
$$
Y=\beta_0 L^{\beta_1} K^{\beta_2}
$$
where $y$ is the value added, $L$ is the labor input, and $K$ is the capital input.
d. Use multiple regression transformed variables to estimate a Cobb-Douglas production function with constant returns to scale. Note that this production function has the same form as the function estimated in part $c$, but it has the additional restriction that $\beta_1+\beta_2=1$. To develop the transformed regression model, substitute $\beta_2$ as a function of $\beta_1$ and convert to a regression format.
e. Compare the three production functions using residual plots and a standard error of the estimate that is expressed in the same scale. You will need to convert the predicted values from parts $\mathrm{c}$ and $\mathrm{d}$, which are in logarithms, back to the original units. Then you can subtract the predicted values from the original values of $Y$ to obtain the residuals. Use the residuals to compute comparable standard errors of the estimate.

Check back soon!
01:42

Problem 60

The administrator of a small city has asked you to identify variables that influence the mean market value of houses in small midwestern cities. You have obtained data from a number of small cities, which are stored in the data file Citydatr, with variables described in the Chapter 12 appendix. The candidate predictor variables are the median size of the house (sizehse), the property tax rate (taxrate; tax levy divided by total assessment), the total expenditures for city services (totexp), and the percent commercial property (comper).
a. Estimate the multiple regression model using all the indicated predictor variables. Select only statistically significant variables for your final equation.
b. An economist stated that since the data came from cities of different populations, your model is likely to contain heteroscedasticity. He argued that mean housing prices from larger cities would have a smaller variance because the number of houses used to compute the mean housing prices would be larger. Test for heteroscedasticity.
c. Estimate the multiple regression equation using weighted least squares with population as the weighting variable. Compare the coefficients for the weighted and unweighted multiple regression models.

Sheryl Ezze
Sheryl Ezze
Numerade Educator
06:25

Problem 61

The chief financial officer of a major service company has asked you to develop a regression model to predict consumption of service goods as a function of GDP and other important variables. The data for your analysis are found in the data file Macro2010, which is stored on your data disk and described in the data dictionary in the chapter appendix. Use data from the period 1980.1 through 2010.4.
a. Estimate a regression model using only GDP to predict consumption of service goods. Test for autocorrelation using the Durbin-Watson statistic.
b. Estimate a multiple regression model using GDP, total consumption lagged 1 period, imports or services, and prime interest rate as additional predictors. Test for autocorrelation. Does this multiple regression model reduce the problem of autocorrelation?

Gus Steppen
Gus Steppen
Numerade Educator

Problem 62

The marketing vice president of Consolidated Appliances has asked you to develop a regression model to predict consumption of durable goods as a function of disposable personal income and other important variables. The data for your analysis are found in the data file Macro2010, which is described in the data dictionary in the chapter appendix. Use data from the period 1976.1 through 2010.4.
a. Estimate a regression model using only disposable personal income to predict consumption of durable goods. Test for autocorrelation using the Durbin-Watson statistic.
b. Estimate a multiple regression model using disposable personal income, total consumption lagged 1 period, imports of goods, population, and prime interest rate as additional predictors. Test for autocorrelation. Does this multiple regression model reduce the problem of autocorrelation?

Check back soon!
05:36

Problem 63

You have been asked to develop a model using multiple regression that predicts the retail sale of beef using time-series data. The data file Beef Veal Consumption contains a number of variables related to the beef retail markets beginning in 1935 and extending through the present. The variables are described in the Chapter 13 appendix.
a. Prepare a model that includes a test and adjustment for serial correlation. Discuss your model and indicate important factors that predict beef sales.
b. Prepare a second analysis, but this time include only data beginning in the year 1980 .
c. Compare the two models estimates in a and b.

Jameson Kuper
Jameson Kuper
Numerade Educator

Problem 64

You have been asked to develop a model using multiple regression that predicts the retail sale of veal using time series data. The data file Beef Veal Consumption contains a number of variables related to the veal retail markets beginning in 1935 and extending through the present.
a. Prepare a model that includes a test and adjustment for serial correlation. Discuss your model and indicate important factors that predict beef sales.
b. Prepare a second analysis, but this time include only data beginning in the year 1980 .
c. Compare the two models estimates in $\mathrm{a}$ and $\mathrm{b}$.

Check back soon!

Problem 65

You have been asked to develop a model using multiple regression that predicts the retail sale of beef and veal combined using time series data. The data file Beef Veal Consumption contains a number of variables related to the beef and veal retail markets beginning in 1935 and extending through the present.
a. Prepare a model that includes a test and adjustment for serial correlation. Discuss your model and indicate important factors that predict beef sales.
b. Prepare a second analysis, but this time include only data beginning in the year 1980 .
c. Compare the two models estimates in $a$ and $b$.

Check back soon!
04:09

Problem 66

Health care cost is an increasingly important part of the United States economy. In this exercise you are to identify variables that are predictors for the cost of physician and clinical services, either individually or in combination. Use the data file Health Care Cost Analysis, which contains annual health care costs for the period 1960-2008. As a first step you are to explore the simple relationships between physician and clinical services cost and individual variables using a combination of simple correlations and graphical scatter plots. You should also examine the changes in cost of physicians and clinical services and other variables over time. Medical care costs are, of course, affected by various national policies and changes in health care providers and health insurance practice. Based on these analyses, develop a multiple regression model that predicts costs of physicians and clinical services. You will probably find that the model has errors that are serially correlated and this possibility should be tested for by using the Durbin-Watson test.
If serial correlation exists in your initial model then to adjust for serial correlation, you are to use the difference variables to estimate a model that predicts the change in physician and clinical services as a function of change in the predictor variables. Again, explore the simple relationship between the change in physician and clinical services and the change in the other predictor variables using correlations and scatter plots. Using these results, develop a multiple regression model using the changes in variables to predict the change in physician and clinical services costs.
Prepare a report that identifies variables that are related to cost of physicians and clinical services individually and in combination.

Heather Duong
Heather Duong
Numerade Educator
04:09

Problem 67

Health care cost is an increasingly important part of the U.S. economy. In this exercise you are to identify variables that are predictors for hospital cost, either individually or in combination. Use the data file Health Care Cost Analysis, which contains annual health care costs for the period 1960-2008. As a first step you are to explore the simple relationships between hospital cost and individual variables using a combination of simple correlations and graphical scatter plots. You should also examine the changes in hospital cost and other variables over time. Medical care costs are, of course, affected by various national policies and changes in health care providers and health insurance practice. Based on these analyses, develop a multiple regression model that predicts hospital cost. You will probably find that the model has errors that are serially correlated and this possibility should be tested for by using the DurbinWatson test.
If serial correlation exists in your initial model then use the difference variables to estimate a model that predicts the change as a function of change in the predictor variables. Again, explore the simple relationship between the change in hospital cost and the change in the other predictor variables using correlations and scatter plots. Using these results develop a multiple regression model using the changes in variables to predict the change in hospital care costs.
Prepare a report that identifies variables that are related to hospital cost individually and in combination.

Heather Duong
Heather Duong
Numerade Educator

Problem 68

Health care cost is an increasingly important part of the U.S. economy. In this exercise you are to identify variables that are predictors for drug cost, either individually or in combination. Use the data file Health Care Cost Analysis, which contains annual health care costs for the period 1960-2008. As a first step you are to explore the simple relationships between drug cost and individual variables using a combination of simple correlations and graphical scatter plots. You should also examine the changes in drug cost and other variables over time. Medical care costs are, of course, affected by various national policies and changes in health care providers and health insurance practice. Based on these analyses, develop a multiple regression model that predicts drug costs. You will probably find that the model has errors that are serially correlated and this possibility should be tested for by using the Durbin-Watson test.
If serial correlation exists in your initial model then use the difference variables to estimate a model that predicts the change in drug costs as a function of change in the predictor variables. Again, explore the simple relationship between the change in drug cost and the change in the other predictor variables using correlations and scatter plots. Using these results, develop a multiple regression model using the changes in variables to predict the change in drug cost.
Prepare a report that identifies variables that are related to drug cost individually and in combination.
$$
\begin{array}{ll}
\hline \text { C1 } & \text { Year } \\
\text { C2 } & \text { National Health Expenditures } \\
\text { C3 } & \text { Medicare } \\
\text { C4 } & \text { Hospital Care } \\
\text { C5 } & \text { Physician and Clinical Services } \\
\text { C6 } & \text { Prescription Drugs } \\
\text { C7 } & \text { Admin. \& Net Cost of Priv. Hlth } \\
\text { C8 } & \text { Income Low } 5 \text { th } \\
\text { C9 } & \text { Income Median } \\
\text { C10 } & \text { Income High } 5 \text { th } \\
\text { C11 } & \text { Income High } 5 \% \\
\text { C12 } & \text { Population } \\
\text { C13 } & \text { Unemployment } \\
\text { C14 } & \text { Percent } 65 \text { plus } \\
\text { C15 } & \text { Per age }<5 \\
\text { C16 } & \text { Lag Hosp care } \\
\text { C17 } & \text { Difference Hosp Care } \\
\text { C18 } & \text { Difference Physician } \\
\text { C19 } & \text { Difference Drugs } \\
\text { C20 } & \text { Difference Population } \\
\text { C21 } & \text { Difference } \%>65 \\
\text { C22 } & \text { Difference } \%<5 \\
\text { C23 } & \text { Difference Medicare Cost } \\
C 24 & \text { Difference Income }>5 \% \\
C 25 & \text { Difference Income Median } \\
C 26 & \text { Lag Diff } \% \text { Age }>65 \\
\hline
\end{array}
$$

Check back soon!