1. Markov Decision Processes (5 points)
For Thanksgiving, you are trying to prepare a turkey. To prepare the turkey, one can either decrease or increase the temperature of the oven, resulting in the turkey being undercooked, cooked, or burnt. The Markov Decision Process for this scenario is provided in Figure 1.
The transition probabilities for decreasing and increasing the temperature are provided in blue and red, respectively. The rewards for the transitions are underlined and green.
Using value iteration on the given MDP, compute $V_1(cooked)$. Write down the intermediate steps, including the Q-values for the state-action pairs.
+1
0.3
-10
0.2
Decrease
+1
Increase
1.0
+1
0.7
+4
0.4
0.6 +1
0.8
+2
Undercooked
Cooked
Burnt
Figure 1: Markov Decision Process