Toggle Navigation
Games
Blog
Class PIN
Join for Free
Sign in
Toggle Navigation
Games
PIN
Join for Free
Blog
Pricing
Contact us
Help center
Sign in
Study
Review Final Exam
0
%
0
0
0
Back
Restart
If γ = 0, the agent cares about:
Immediate reward only
Oops!
Okay!
MDP stands for:
Markov Decision Process
Oops!
Okay!
What does transition function T(s,a,s') represent?
Probability of moving to state s′ after action a in state s.
Oops!
Okay!
When do we stop Value Iteration?
When values stop changing significantly.
Oops!
Okay!
How is a centroid updated in K-means?
By taking the mean of all assigned points.
Oops!
Okay!
What does γ (gamma) represent?
Discount factor
Oops!
Okay!
Why can K-means give different results on the same dataset?
Random centroid initialization.
Oops!
Okay!
Policy iteration consists of:
Policy evaluation + policy improvement
Oops!
Okay!
What is the main goal of PCA?
PCA reduces dimensions while preserving maximum variance.
Oops!
Okay!
K-means is what type of learning?
Unsupervised
Oops!
Okay!
What happens first in K-means?
Assign points
Oops!
Okay!
If K increases, WCSS(cost function) usually:
Decreases
Oops!
Okay!
In PCA, what does the first principal component maximize?
Variance
Oops!
Okay!
When does Policy Iteration stop?
When policy no longer changes.
Oops!
Okay!
If Q(s,a) values are: Left = 3 Right = 7 Up = 5 What action will policy choose?
Right
Oops!
Okay!
What does K represent in K-means?
Oops!
Okay!
If living reward is negative, what behavior is encouraged?
Faster finishing.
Oops!
Okay!
Which are components of an MDP?
States, Actions, Rewards
Oops!
Okay!
What is a terminal state?
A state where the episode ends.
Oops!
Okay!
Your experience on this site will be improved by allowing cookies.
Allow cookies