Foundations of Statistics: Exam notes REMEMBER ITALICS FOR REPORTING STATS !!! Describing the distribution of a single variable? Exploring the relationship between two variables? Metric? Categorical? Both Metric? One Metric & One Categorical? Both Categorical? V Box Plots, Graph of means Histogram Box Plot Summary Statistics Percentiles Bar chart Pie chart Frequency Table Scatterplot Repeated Measures? Independent Groups? Crosstabulation 2 conditions? More than 2 conditions? 2 conditions? More than 2 conditions? One sample t-test Binomial test Pearson's r Linear Regression Paired Samples t-tests Not covered in this unit Independent samples t-tests Not covered in this unit Chi-Square Statistic 68-95-99.7 rule: used for normal (symmetric or bell curved) distributions. 68% of values lie within 1 SD of the mean. 95% fall within 2 SD's. 99.7% fall within 3 SD's of the mean. Confounding variable - a variable in the study (apart from the IV) that increases variation in the DV but also varies systematically with the IV. To remove it use randomisation. MODE - most common category MEAN - average MEDIAN - middle number when listed in order - 50th Percentile Directional hypothesis = predicting 1 group does/has more/less than another group. Non-directional hypothesis = groups make a different number of VARIABLE. Repeated measures design = when a variable is repeatedly measured on the same participants. Controls nuisance variables effectively. Nuisance variable - a variable in the study (apart from the IV) that increases variation in the DV. Odds Ratio - used for Retrospective studies Eg: The odds of someone in the exercise program developing diabetes were 0.02041 to 1 ( 3/147 Those that did DIVIDE those that didn't) The odds of someone not in the exercise program developing diabetes were 0.04167 to 1 (6/144) Odds ratio = . 02041/.04167 = 0.490 Therefore, The odds of someone in the exercise program developing gestational diabetes are .490 times the odds of someone not in the program developing gestational diabetes. p-value determines the significance of results. If p =. 000 write p <. 001 Significant if less than 0.050 Population - the entire collection of sampling units that data can be sampled from. Relative Risk - used in Experimental & Observational studies - explores whether an IV reduces the risk of the DV. IV referred to as RISK FACTOR. DV referred to as THE OUTCOME. Eg: Proportion of people taking asprin who had a stroke = 15/150 = . 10 Proportion of people taking placebo who had a stroke = 20/100 = . 20 Relative risk = . 10/.20 = 0.50 Therefore, participants who took the asprin were half as likely to suffer from a stroke as those who took the placebo. Regression equation Y=a+bxX Y=DV, a=vertical interept, b=regression coefficient 'slope', X=IV. Sample - is a smaller subset of a population. Sample sizes - small (25) will have more variation in sample proportions, larger (100) cluster closer to the proportion in population. Should be at least 30 to approximate normal distribution. Sampling methods - Stratified random sampling = Dividing the sampling between each