Reinforcement Learning & Decision Making Flashcards
7 cards from real AML practice questions. Tap to flip, then mark Knew It or Still Learning — missed cards come back until you master them.
Read the first 7 Reinforcement Learning & Decision Making flashcards as text
In reinforcement learning, what does the agent seek to maximize over time?
Answer: Cumulative discounted reward
An RL agent maximizes cumulative discounted reward over time, balancing immediate and future gains according to the discount factor.
Which of the following is NOT a standard component of a Markov Decision Process (MDP)?
Answer: Neural network architecture
An MDP is formally defined by states, actions, transition probabilities, rewards, and a discount factor — neural network architecture is an implementation choice, not part of the MDP definition.
What effect does setting the discount factor (γ) close to 0 have on a reinforcement learning agent?
Answer: The agent becomes myopic, prioritizing only immediate rewards
A discount factor near 0 causes the agent to heavily discount future rewards, making it focus almost entirely on immediate rewards (myopic behavior).
What is the key distinction between model-based and model-free reinforcement learning?
Answer: Model-based RL learns or uses an environment model for planning; model-free learns directly from interactions
Model-based RL builds or leverages an explicit model of environment dynamics for planning, while model-free RL learns policies or value functions directly from experience without modeling transitions.
What is the exploration-exploitation dilemma in reinforcement learning?
Answer: Balancing trying new actions to gather information against using known high-reward actions
The exploration-exploitation dilemma involves balancing the need to explore unknown actions (to discover better rewards) against exploiting already-known high-reward actions.
What does a stochastic policy output for a given state in reinforcement learning?
Answer: A probability distribution over all possible actions
A stochastic policy maps each state to a probability distribution over actions, enabling exploration by sampling from this distribution during training.
What does the Bellman equation fundamentally express in reinforcement learning?
Answer: The recursive relationship between a state's value and the values of its successor states
The Bellman equation expresses value recursively: the value of a state equals the immediate reward plus the discounted value of successor states, enabling dynamic programming solutions.