00:01
Hello viewers, in this problem we have a table that gives us three values of x, two values of y, and we have to explain the idea behind the simpsom's paradox and construct in part a the frequency marginal distribution of this table.
00:15
Now, simpsons paradox in statistics is an effect that occurs when the marginal association between two categorical variables is qualitatively different from the partial association between the same two variables after controlling.
00:30
For one or more variables.
00:32
And a marginal distribution of a variable is defined as the frequency distribution of either row or column variable for the desired contingency table.
00:42
This table is useful to eradicate the influence of either row or column variable for the desired contingency table.
00:49
Now for this we will add individual values of each variable and make a table for the marginal distribution.
00:58
Hence this is the table for the frequency marginal distribution.
01:01
Where all the rows and columns are individually added.
01:07
That is, 20 plus 30 will give us 50 and 20 plus 25 plus 30 will give us 75.
01:14
And this process is repeated for all the three rows and columns.
01:19
Now, in part b we have to construct a relative frequency marginal distribution.
01:24
A relative frequency is defined as the number of times the value occurs divided by the total number of observations in the data system.
01:32
Set.
01:34
So, if we denote it as rf, that is relative frequency, it will be frequency divided by the total frequency.
01:52
The relative frequency marginal distribution for the row variable is obtained by calculating or dividing the total or the row total for each value by the contingency table...