Cost Optimization and Cloud Resource Management Flashcards
6 cards from real SRE practice questions. Tap to flip, then mark Knew It or Still Learning — missed cards come back until you master them.
Read the first 6 Cost Optimization and Cloud Resource Management flashcards as text
What is 'cloud cost optimization' in the SRE context, and why is it part of SRE's responsibilities?
Answer: Cloud cost optimization ensures that infrastructure resources are right-sized and efficiently utilized, supporting reliability goals by preventing resource waste that could be invested in redundancy, monitoring, and DR capabilities
SREs own infrastructure efficiency as part of service ownership — over-provisioned services waste budget that could fund reliability improvements, while under-provisioned services create reliability risks. Cost and reliability are two dimensions of the same resource allocation decision.
What is 'right-sizing' in cloud infrastructure, and what data is needed to do it effectively?
Answer: Right-sizing matches instance types and sizes to actual resource consumption by analyzing CPU, memory, and I/O utilization metrics over time, selecting the smallest instance that provides adequate headroom for peak load and scaling events
Right-sizing requires historical utilization data to identify instances that are significantly over-provisioned (average CPU <20%, memory <30%) and to determine the appropriate size that provides performance headroom without excessive waste.
What are 'reserved instances' (RIs) or 'committed use discounts' (CUDs), and when should SREs recommend them?
Answer: Reserved instances are pre-purchased cloud compute commitments (typically 1-3 years) that provide 30-70% cost savings over on-demand pricing in exchange for committing to a minimum usage level — appropriate for stable, predictable baseline workloads
RIs provide significant cost savings (often 40-60%) for stable workloads in exchange for a usage commitment. They are appropriate for services with predictable, stable resource needs — not for development environments or highly variable workloads.
What are 'spot instances' (AWS) or 'preemptible VMs' (GCP), and what type of workloads are they MOST suitable for?
Answer: Spot/preemptible instances are deeply discounted (70-90%) cloud compute that can be reclaimed by the provider with short notice; they are most suitable for fault-tolerant batch workloads like data processing, rendering, and CI/CD workers that can be safely interrupted and restarted
The 2-minute reclamation notice means spot/preemptible instances are unsuitable for stateful production services or real-time user-facing workloads. They excel at fault-tolerant batch jobs that checkpoint progress and can restart from a checkpoint when preempted.
What is 'resource tagging strategy' in cloud cost management, and why is it important for SRE teams?
Answer: Resource tags (key-value pairs on cloud resources) enable cost allocation by service, team, environment, and owner — allowing SREs to identify cost anomalies, enforce accountability, and make data-driven right-sizing decisions per service
Without consistent tagging, cloud costs appear as undifferentiated infrastructure spend. Tagging enables per-service, per-team cost visibility — essential for identifying which services are expensive, which teams are over-provisioned, and where optimization has the most impact.
What is 'idle resource detection,' and what are common examples of cloud waste it addresses?
Answer: Idle resource detection identifies cloud resources that are running but not being used, such as stopped EC2 instances with attached EBS volumes, unattached Elastic IPs, oversized databases with near-zero query load, and dev environments running 24/7
Cloud environments accumulate waste over time as services are deprecated but not fully decommissioned, development environments are left running, and snapshots/volumes accumulate without cleanup policies.