Question

Consider the 2-band QMF bank shown in Figure 2.35. In this figure, $x(n)$ denotes speech frames of 256 samples and $\hat{x}(n)$ denotes the synthesized speech frames. a. Design the transfer functions, $F_0(z)$ and $F_1(z)$, such that aliasing is cancelled. Also calculate the overall delay of the QMF bank. b. Select an arbitrary voiced speech frame from $\mathrm{Ch} 2$ speech.wav. Give time-domain and frequency-domain plots of $x_{d 0}(n)$ and $x_{d 1}(n)$ for that particular frame. Comment on the frequency-domain plots with regard to low-pass/high-pass band-splitting. Figure 2.35. A two-band QMF bank. FIGURE CANT COPY Figure 2.36. Speech synthesis from a select number (subset) of FFT components. FIGURE CANT COPY Table 2.2. Signal-to-noise ratio (SNR) and MOS values. $$ \begin{array}{|c|c|c|} \hline \begin{array}{l} \text { Number of FFT } \\ \text { components, } L \end{array} & \mathrm{SNR}_{\text {overall }} & \begin{array}{l} \text { Subjective evaluation MOS (mean } \\ \text { opinion score) in a scale of } \\ 1-5 \text { for the entire speech record } \end{array} \\ \hline \begin{array}{l} 16 \\ 128 \end{array} & & \\ \hline \end{array} $$ c. Repeat step (b) for $x_1(n)$ and $x_{d 1}(n)$ in order to compare the signals before and after the downsampling stage. d. Calculate the SNR between the input speech record, $x(n)$, and the synthesized speech record, $\hat{x}(n)$. Use the following equation to compute the SNR, $$ \mathrm{SNR}=10 \log _{10}\left(\frac{\sum_n x^2(n)}{\sum_n(x(n)-\hat{x}(n))^2}\right)(\mathrm{dB}) $$ Listen to the synthesized speech record and comment on its quality. e. Choose a low-pass $F_0(z)$ and a high-pass $F_1(z)$, such that aliasing occurs. Compute the SNR. Listen to the synthesized speech and describe its perceptual quality relative to the output speech in step(d). (Hint: Use first-order IIR filters.)

   Consider the 2-band QMF bank shown in Figure 2.35. In this figure, $x(n)$ denotes speech frames of 256 samples and $\hat{x}(n)$ denotes the synthesized speech frames.
a. Design the transfer functions, $F_0(z)$ and $F_1(z)$, such that aliasing is cancelled. Also calculate the overall delay of the QMF bank.
b. Select an arbitrary voiced speech frame from $\mathrm{Ch} 2$ speech.wav. Give time-domain and frequency-domain plots of $x_{d 0}(n)$ and $x_{d 1}(n)$ for that particular frame. Comment on the frequency-domain plots with regard to low-pass/high-pass band-splitting.
Figure 2.35. A two-band QMF bank.
FIGURE CANT COPY
Figure 2.36. Speech synthesis from a select number (subset) of FFT components.
FIGURE CANT COPY
Table 2.2. Signal-to-noise ratio (SNR) and MOS values.
$$
\begin{array}{|c|c|c|}
\hline \begin{array}{l}
\text { Number of FFT } \\
\text { components, } L
\end{array} & \mathrm{SNR}_{\text {overall }} & \begin{array}{l}
\text { Subjective evaluation MOS (mean } \\
\text { opinion score) in a scale of } \\
1-5 \text { for the entire speech record }
\end{array} \\
\hline \begin{array}{l}
16 \\
128
\end{array} & & \\
\hline
\end{array}
$$
c. Repeat step (b) for $x_1(n)$ and $x_{d 1}(n)$ in order to compare the signals before and after the downsampling stage.
d. Calculate the SNR between the input speech record, $x(n)$, and the synthesized speech record, $\hat{x}(n)$. Use the following equation to compute the SNR,
$$
\mathrm{SNR}=10 \log _{10}\left(\frac{\sum_n x^2(n)}{\sum_n(x(n)-\hat{x}(n))^2}\right)(\mathrm{dB})
$$

Listen to the synthesized speech record and comment on its quality.
e. Choose a low-pass $F_0(z)$ and a high-pass $F_1(z)$, such that aliasing occurs. Compute the SNR. Listen to the synthesized speech and describe its perceptual quality relative to the output speech in step(d). (Hint: Use first-order IIR filters.)
Show more…
Audio Signal Processing and Coding
Audio Signal Processing and Coding
Andreas Spanias, Ted… 1st Edition
Chapter 2, Problem 24 ↓

Instant Answer

verified

Step 1

To design the transfer functions \( F_0(z) \) and \( F_1(z) \) such that aliasing is cancelled, we need to ensure that the sum of the polyphase components of the analysis and synthesis filters results in a perfect reconstruction. For a 2-band QMF bank, the  Show more…

Show all steps

lock
AceChat toggle button
Close icon
Ace pointing down

Please give Ace some feedback

Your feedback will help us improve your experience

Thumb up icon Thumb down icon
Thanks for your feedback!
Profile picture
Consider the 2-band QMF bank shown in Figure 2.35. In this figure, $x(n)$ denotes speech frames of 256 samples and $\hat{x}(n)$ denotes the synthesized speech frames. a. Design the transfer functions, $F_0(z)$ and $F_1(z)$, such that aliasing is cancelled. Also calculate the overall delay of the QMF bank. b. Select an arbitrary voiced speech frame from $\mathrm{Ch} 2$ speech.wav. Give time-domain and frequency-domain plots of $x_{d 0}(n)$ and $x_{d 1}(n)$ for that particular frame. Comment on the frequency-domain plots with regard to low-pass/high-pass band-splitting. Figure 2.35. A two-band QMF bank. FIGURE CANT COPY Figure 2.36. Speech synthesis from a select number (subset) of FFT components. FIGURE CANT COPY Table 2.2. Signal-to-noise ratio (SNR) and MOS values. $$ \begin{array}{|c|c|c|} \hline \begin{array}{l} \text { Number of FFT } \\ \text { components, } L \end{array} & \mathrm{SNR}_{\text {overall }} & \begin{array}{l} \text { Subjective evaluation MOS (mean } \\ \text { opinion score) in a scale of } \\ 1-5 \text { for the entire speech record } \end{array} \\ \hline \begin{array}{l} 16 \\ 128 \end{array} & & \\ \hline \end{array} $$ c. Repeat step (b) for $x_1(n)$ and $x_{d 1}(n)$ in order to compare the signals before and after the downsampling stage. d. Calculate the SNR between the input speech record, $x(n)$, and the synthesized speech record, $\hat{x}(n)$. Use the following equation to compute the SNR, $$ \mathrm{SNR}=10 \log _{10}\left(\frac{\sum_n x^2(n)}{\sum_n(x(n)-\hat{x}(n))^2}\right)(\mathrm{dB}) $$ Listen to the synthesized speech record and comment on its quality. e. Choose a low-pass $F_0(z)$ and a high-pass $F_1(z)$, such that aliasing occurs. Compute the SNR. Listen to the synthesized speech and describe its perceptual quality relative to the output speech in step(d). (Hint: Use first-order IIR filters.)
Close icon
Play audio
Feedback
Powered by NumerAI
Need help? Use Ace
Ace is your personal tutor. It breaks down any question with clear steps so you can learn.
Start Using Ace
Ace is your personal tutor for learning
Step-by-step explanations
Instant summaries
Summarize YouTube videos
Understand textbook images or PDFs
Study tools like quizzes and flashcards
Listen to your notes as a podcast
Continue solving this problem
Create a free account to:
  • View full step-by-step solution
  • Ask follow-up questions with Ace AI
  • Save progress and study later
Continue Free
Numerade

Get step-by-step video solution
from top educators

Continue with Clever
or



By creating an account, you agree to the Terms of Service and Privacy Policy
Already have an account? Log In

A free answer
just for you

Watch the video solution with this free unlock.

Numerade

Log in to watch this video
...and 100,000,000 more!


EMAIL

PASSWORD

OR
Continue with Clever