Incident and Event Response Flashcards
7 cards from real AWS practice questions. Tap to flip, then mark Knew It or Still Learning โ missed cards come back until you master them.
Read the first 7 Incident and Event Response flashcards as text
Which AWS service provides a personalized view of the health of AWS services and sends notifications about events that might impact your specific resources?
Answer: AWS Health Dashboard
AWS Health Dashboard provides account-specific health events and notifications for resources in your account, distinguishing it from the general public service health page.
Which AWS Systems Manager capability provides a central location where engineers can view, investigate, and resolve operational work items called OpsItems?
Answer: Systems Manager OpsCenter
OpsCenter aggregates and standardizes OpsItems from various AWS services, giving operations teams a unified view for investigation and resolution.
An automated remediation Lambda function needs to stop a non-compliant EC2 instance when a configuration change is detected. Which service should trigger this action automatically?
Answer: AWS Config automatic remediation
AWS Config evaluates resource configurations against rules and can trigger automatic remediation actions, including invoking SSM Automation runbooks or Lambda functions on non-compliant resources.
During a major incident, your team needs to run a pre-approved multi-step operational workflow across several AWS services. Which Systems Manager feature should you use?
Answer: Systems Manager Automation
Systems Manager Automation executes runbooks (SSM Automation documents) that define multi-step workflows, API calls, and conditional branching for incident remediation.
Which AWS service provides on-call schedules, escalation plans, response plans, and post-incident analysis capabilities for DevOps incident management?
Answer: AWS Systems Manager Incident Manager
AWS Systems Manager Incident Manager manages the full incident lifecycle with on-call management, escalation plans, runbook execution, and post-incident analysis templates.
What is the primary purpose of a Dead Letter Queue (DLQ) in an event-driven incident response architecture?
Answer: To capture and retain messages that fail processing for later analysis
DLQs capture messages that cannot be successfully processed after the maximum number of retries, preserving them for debugging and reprocessing rather than losing them.
Which Amazon EventBridge feature allows your team to capture events during an incident and replay them against remediation logic for post-incident validation?
Answer: EventBridge Event Archive and Replay
EventBridge Event Archive captures a stream of events matching specified criteria, and Replay replays those archived events against an event bus, enabling post-incident testing and investigation.