Consider the 2-band QMF bank shown in Figure 2.35. In this figure, $x(n)$ denotes speech frames of 256 samples and $\hat{x}(n)$ denotes the synthesized speech frames.
a. Design the transfer functions, $F_0(z)$ and $F_1(z)$, such that aliasing is cancelled. Also calculate the overall delay of the QMF bank.
b. Select an arbitrary voiced speech frame from $\mathrm{Ch} 2$ speech.wav. Give time-domain and frequency-domain plots of $x_{d 0}(n)$ and $x_{d 1}(n)$ for that particular frame. Comment on the frequency-domain plots with regard to low-pass/high-pass band-splitting.
Figure 2.35. A two-band QMF bank.
FIGURE CANT COPY
Figure 2.36. Speech synthesis from a select number (subset) of FFT components.
FIGURE CANT COPY
Table 2.2. Signal-to-noise ratio (SNR) and MOS values.
$$
\begin{array}{|c|c|c|}
\hline \begin{array}{l}
\text { Number of FFT } \\
\text { components, } L
\end{array} & \mathrm{SNR}_{\text {overall }} & \begin{array}{l}
\text { Subjective evaluation MOS (mean } \\
\text { opinion score) in a scale of } \\
1-5 \text { for the entire speech record }
\end{array} \\
\hline \begin{array}{l}
16 \\
128
\end{array} & & \\
\hline
\end{array}
$$
c. Repeat step (b) for $x_1(n)$ and $x_{d 1}(n)$ in order to compare the signals before and after the downsampling stage.
d. Calculate the SNR between the input speech record, $x(n)$, and the synthesized speech record, $\hat{x}(n)$. Use the following equation to compute the SNR,
$$
\mathrm{SNR}=10 \log _{10}\left(\frac{\sum_n x^2(n)}{\sum_n(x(n)-\hat{x}(n))^2}\right)(\mathrm{dB})
$$
Listen to the synthesized speech record and comment on its quality.
e. Choose a low-pass $F_0(z)$ and a high-pass $F_1(z)$, such that aliasing occurs. Compute the SNR. Listen to the synthesized speech and describe its perceptual quality relative to the output speech in step(d). (Hint: Use first-order IIR filters.)