โ† All Artificial Intelligence Flashcard Decks

Reinforcement Learning Flashcards

6 cards from real Artificial Intelligence practice questions. Tap to flip, then mark Knew It or Still Learning โ€” missed cards come back until you master them.

Read the first 6 Reinforcement Learning flashcards as text
  1. What is Proximal Policy Optimization (PPO) designed to solve compared to earlier policy gradient methods?

    Answer: Preventing excessively large policy updates that destabilize training by clipping the objective function

    PPO clips the policy update ratio to restrict changes within a small range, ensuring stable incremental improvement without the complexity of TRPO's hard constraint.

  2. What is an 'actor-critic' architecture in reinforcement learning?

    Answer: A design combining a policy network (actor) that selects actions and a value network (critic) that evaluates them to reduce variance

    Actor-critic methods simultaneously learn a policy (actor) and a value function (critic), using the critic's value estimates to reduce the high variance of pure policy gradient updates.

  3. What was the significance of AlphaGo defeating the world Go champion in 2016?

    Answer: It demonstrated that deep RL combined with Monte Carlo Tree Search could master a game long considered too complex for AI due to its enormous state space

    AlphaGo's victory showed that deep RL combined with MCTS could outperform top human players in Go, a game with more possible positions than atoms in the observable universe.

  4. What is 'reward shaping' in reinforcement learning?

    Answer: Adding supplementary reward signals to guide the agent toward desired behaviors when the natural reward is sparse or delayed

    Reward shaping augments the environment's sparse reward with additional dense rewards that reflect progress toward the goal, speeding up learning without changing the optimal policy.

  5. What is 'multi-agent reinforcement learning' (MARL)?

    Answer: A setting with multiple agents that simultaneously learn in a shared environment, potentially cooperating or competing

    MARL studies how multiple agents learn simultaneously in environments where their actions affect each other, encompassing cooperative, competitive, and mixed settings.

  6. What distinguishes 'on-policy' from 'off-policy' reinforcement learning?

    Answer: On-policy algorithms learn about the policy being used to collect data; off-policy algorithms can learn from data generated by a different (older) policy

    On-policy methods (e.g., SARSA, PPO) update the policy using only data collected by the current policy, while off-policy methods (e.g., Q-learning, DQN) can reuse older experience stored in a replay buffer.