The goal in predictive analysis is to use training data to learn a model that can make predictions on new data. Answer the following questions. a. Suppose we increased the size of the training set. Would this likely improve or deteriorate the performance of the model on new data? Why? Increasing the size of the training set is likely to improve the model's performance on new data. Increasing the size of the training set tends to add more variability to the data. More variability tends to make it easier to detect which features are truly correlated with the target class and which features are not. b) Suppose we reduced the feature representation to include only the features with the highest mutual information with the target concept. Would this likely improve or deteriorate the performance of the model on new data? Why?
Added by Lori B.
Your feedback will help us improve your experience
Sri K and 85 other Intro Stats / AP Statistics educators are ready to help you.
Ask a new question
Labs
Want to see this concept in action?
Explore this concept interactively to see how it behaves as you change inputs.
Key Concepts
Recommended Videos
Training Data Data Point Feature A Feature B Target feature (Class) 1 2.5 3.4 1 2 1.4 -0.2 1 3 4.2 4.9 2 4 11.1 5.3 2 5 2.9 15.5 3 6 3.6 13.9 3 Testing Data testing data Feature A Feature B Target Feature (Class) 1 3.9 4.3 ? (Note: Round al numeric answers to 2 decimal places) The Euclidean distance of testing data 1 to data point 1 is . The Euclidean distance of testing data 1 to data point 2 is . The Euclidean distance of testing data 1 to data point 3 is . The Euclidean distance of testing data 1 to data point 4 is . The Euclidean distance of testing data 1 to data point 5 is . The Euclidean distance of testing data 1 to data point 6 is . The predicted target feature (class) when k = 1 is . The predicted target feature (class) when k = 3 is .
Madhur L.
1. What is the difference between prediction models and classification models? Give examples. 2. What accuracy measures are used to evaluate the predictive performance of a classification model? 3. What accuracy measures are used to evaluate the predictive performance of a prediction model? 4. How can the problem of overfitting be addressed? ' 5. When oversampling procedure is used? 6. A classification model was developed to predict whether an accident is deadly. In a validation dataset of 500 observations there were 450 who survived. Using this dataset for predictive assessment, the model classified 40 accidents as deadly. Compute the sensitivity and interpret. 7. What is the prediction error in a prediction model? How can predictive performance be evaluated?
Samriddhi S.
Recommended Textbooks
Elementary Statistics a Step by Step Approach
The Practice of Statistics for AP
Introductory Statistics
Transcript
Watch the video solution with this free unlock.
EMAIL
PASSWORD