Business Intelligence- ITAP 3382 (Homework #1)
Due Date: March 5, 2025
Q1 (20 marks):
a. What is the purpose of data normalization? Why is it needed?
b. What are the major differences between K-Means and Hierarchical Agglomerative Clustering approaches? Which algorithm is computationally more expensive? Why?
c. How do we determine optimal number of clusters in K-Means algorithm?
d. Based on the following dissimilarity matrix:
a. Which two data objects are most similar?
b. Which two data objects are most dissimilar?
c. Which data object is most similar to data object X4? Why?
\begin{tabular}{|l|l|l|l|l|l|}
\hline
& X1 & X2 & X3 & X4 & X5 & X6 \\
\hline
X1 & 0 & & & & & \\
\hline
X2 & 2 & 0 & & & & \\
\hline
X3 & 4 & 9 & 0 & & & \\
\hline
X4 & 1 & 7 & 1.1 & 0 & & \\
\hline
X5 & 9 & 5 & 2.5 & 2.9 & 0 & \\
\hline
X6 & 0.5 & 1 & 0.8 & 0.9 & 2 & 0 \\
\hline
\end{tabular}