Swinburne University of Technology Faculty of Health, Arts & Design STA30004 Practice Exam STA30004 Data Mining The actual exam has a similar format to this in that it also has four questions with a maximum mark of 100 of which 40 marks are for multiple choice questions. Hope this helps. You are allowed a one A4 set of notes for the exam. The exam is closed book. Question 1 (Multiple Choice: 2 marks per question) Please choose the answer that is MOST correct. Circle the bullet point for your chosen answer. a) Which ONE of the following methods can be used for undirected (unsupervised) data mining? · Regression trees · Self-Organising Maps · Gradient Boosting . Random Forests b) Which of the following descriptions of data mining is most INCORRECT? · It involves the exploration and analysis of large quantities of data. . It discovers meaningful rules and patterns that can be exploited to reveal actionable information. . It is a totally automatic process in that human involvement is negligible. · It helps you uncover what your data is trying to tell you. c) Results for a hypothetical marketing campaign targeting five towns is displayed below. Calculate the response rate, capture rate and lift for the suburb with the best results. Town Population A 200,000 B 300,000 C 100,000 D 200,000 E 200,000 Total 1000,000 Number of responses 1000 12000 1000 2000 4000 20,000 Response Rate (%) 0.5 4 1 1 2 2 Capture Rate (%) 5 tto 2.00 5 0.50 10 0.50 20 100 Lift 0.25 1.00 1.00 Which ONE of the following results are correct for the town with the best results? · Response rate= 1%, capture rate=20%, lift = 1 · Response rate= 4%, capture rate=tt0%, lift = 2 · Response rate= 1%, capture rate=10%, lift = 2 · Response rate= 2%, capture rate=20%, lift = 0.5 Practice Exam, Semester 2, 201ft Page 1 of 21
Swinburne University of Technology Faculty of Health, Arts & Design STA30004 Data Mining d) It is necessary to rescale the variables in a data set to Z-scores or to the [0-1] range in the following situation. . When the scales of measurement is the same and variances are similar for all the input variables. · Building a classification tree · Building a linear regression model · Creating a cluster analysis. Questions e, f, g and h relate to the following association results. The data relates to the demographic data and purchase behaviour of supermarket customers (gender, own/rent, payment by cash/card, product purchased). Note that there are only two possibilities for tenure (Homeowner and Rental). Chain Length Support[%]| Confidence[%] Transaction Count 1 2 25.73 2 2 25.43 3 2 25.13 4 2 23.72 48.57 5 2 22.42 45.34 6 2 21.52 42.07 7 8 2 20.92 42.83 2 20.52 42.01 9 2 19.92 39.41 10 2 19.92 40.78 11 2 19.62 40.16 12 2 19.62 38.81 13 2 19.62 38.81 14 2 19.42 38.42 15 2 19.12 39.14 16 2 18.22 35.62