Open reading frames are regions of DNA that encode at least one AUG start codon and have a long stretch of codons specifying amino acids before any stop codons (UAA, UGA, UAG) are encountered. Recall that a segment of double-stranded DNA has six possible reading frames, three in the top strand and three in the bottom strand.
Consider the double-stranded DNA sequence below:
5' CAATGGCTAGGTACTATGTATGAGATCATGATCTTTACAAATCCGAG 3'
3' GTTACCGATCCATGATACATACTCTAGTACTAGAAATGTTTAGGCTC 5'
(a) (6 pts) Convert the three reading frames of the top strand into RNA sequences by choosing in turn, the first, second, and third nucleotides as the first base in a codon. Write the RNA sequences in the 5' --> 3' direction with a space between each codon. Underline the start codons and double underline the stop codons. Number the reading frames as 1, 2, and 3. Reading frame 1 starts with the first nucleotide, reading frame 2 starts with the second nucleotide, and reading frame 3 starts with the third nucleotide in each strand.
(b) (4 pts) You are trying to identify the gene for a large protein (with many amino acids). Could any of these reading frames encode this protein? Can you eliminate any of the possible reading frames above from coding for this protein? If so, which one(s) and why?