SRE Foundation Certification Exam — Questions and Answers
Question 1: A monitoring system sends 500 alerts in a single day during a major outage, most of which are symptoms of one root cause. What monitoring design pattern would BEST reduce this alert storm?
- Increase alert thresholds globally by 50% to reduce sensitivity
- Route all alerts to email instead of paging to reduce interrupt load
- Implement an alert blackout window during known outage periods
- Alert on symptoms that matter to users (e.g., elevated error rate) rather than on every downstream effect and system metric (Correct answer)
Correct answer: Alert on symptoms that matter to users (e.g., elevated error rate) rather than on every downstream effect and system metric
Symptom-based alerting (alert on high user-visible error rates) rather than cause-based alerting (alert on every internal metric deviation) dramatically reduces alert storms by targeting the root signal rather than all downstream effects.
Question 2: In the context of SRE certification, what is the most important consideration when implementing capacity planning & scaling?
- Ensuring alignment with established standards, stakeholder needs, and best practices (Correct answer)
- Minimizing documentation to save time
- Completing implementation as quickly as possible regardless of quality
- Delegating all responsibilities to junior staff
Correct answer: Ensuring alignment with established standards, stakeholder needs, and best practices
When implementing capacity planning & scaling, SRE professionals must ensure alignment with industry standards and stakeholder needs. Hasty implementation without proper planning often leads to compliance issues and suboptimal outcomes.
Question 3: What is 'toil' in SRE terminology as it relates to on-call work?
- Repetitive, manual, and automatable operational work that scales with service growth (Correct answer)
- Writing documentation and runbooks
- Any work that improves system reliability
- Performing capacity planning exercises
Correct answer: Repetitive, manual, and automatable operational work that scales with service growth
Toil is repetitive manual work tied to running a production service that provides no enduring value and grows with the system — a primary target for automation.
Question 4: Which runbook section helps an on-call engineer determine whether an alert represents a real user-impacting issue?
- Rollback procedures
- Change history
- Impact assessment / severity criteria (Correct answer)
- On-call contact list
Correct answer: Impact assessment / severity criteria
Impact assessment sections define thresholds and signals that distinguish genuine user impact from benign anomalies, guiding severity classification.
Question 5: Why is post-incident review important?
- To punish the responsible person.
- To find bugs in unrelated systems.
- To analyze the root cause and prevent recurrence (Correct answer)
- To ensure compliance with HR rules.
Correct answer: To analyze the root cause and prevent recurrence
Post-incident reviews, often called blameless postmortems, are crucial for learning from failures and improving system resilience. Their purpose is to thoroughly analyze the root cause of an incident, identify contributing factors, and implement preventative measures to avoid similar issues in the future. This process fosters a culture of continuous improvement rather than assigning blame.
Question 6: Which of the following best describes 'toil' in the SRE context?
- Customer support tickets that require engineering intervention to resolve
- Manual, repetitive, automatable operational work that scales linearly with service load and does not produce lasting improvement (Correct answer)
- Technical debt accumulated from architectural shortcuts taken during initial development
- Any work that is performed outside of business hours, including on-call rotations
Correct answer: Manual, repetitive, automatable operational work that scales linearly with service load and does not produce lasting improvement
Toil is the specific category of work that is manual, repetitive, tactical (no enduring improvement), reactive, and scales proportionally with service growth — the opposite of engineering work that reduces future burden.
Question 7: What does 'cardinality' mean in the context of metrics and observability?
- The compression ratio of stored metrics
- The retention period for time-series data
- The number of unique label/tag value combinations for a metric (Correct answer)
- The frequency at which metrics are sampled
Correct answer: The number of unique label/tag value combinations for a metric
Cardinality refers to the number of unique combinations of label values, and high cardinality can strain time-series databases.
Question 8: What is the purpose of a 'blameless post-mortem' in on-call culture?
- To file a formal complaint against a vendor
- To identify which engineer caused the incident so they can be disciplined
- To calculate the financial cost of the incident
- To analyze what went wrong and improve systems without attributing personal fault (Correct answer)
Correct answer: To analyze what went wrong and improve systems without attributing personal fault
Blameless post-mortems focus on systemic causes and process improvements rather than individual blame, fostering a culture where engineers report problems honestly.
Question 9: What is 'N+1 query problem' in application performance, and how does it impact scalability?
- The N+1 problem is a load balancer misconfiguration that sends one extra request to a backend for every N normal requests as a health check
- The N+1 problem occurs when an application makes one query to retrieve N records, then makes N additional individual queries for related data — resulting in N+1 total queries that scale linearly with data size (Correct answer)
- The N+1 problem occurs when autoscaling adds N+1 instances instead of the required N, wasting one instance of capacity
- The N+1 problem refers to requiring N+1 database servers instead of N for redundancy, increasing infrastructure cost proportionally
Correct answer: The N+1 problem occurs when an application makes one query to retrieve N records, then makes N additional individual queries for related data — resulting in N+1 total queries that scale linearly with data size
N+1 queries are a common ORM-related anti-pattern: fetch 100 users (1 query), then fetch each user's profile individually (100 more queries) = 101 total queries instead of 2. This makes the service linearly slower as data volume grows.
Question 10: In the context of release engineering, what is 'environment parity'?
- Keeping development, staging, and production environments as similar as possible to reduce 'works on my machine' issues (Correct answer)
- Using the same CI server for all environments
- Ensuring all environments have equal compute resources
- Matching the number of environments to the number of teams
Correct answer: Keeping development, staging, and production environments as similar as possible to reduce 'works on my machine' issues
Environment parity minimizes differences between environments so that behavior in staging reliably predicts production behavior.
Question 11: What is 'context-sensitive escalation' and how does it improve on simple time-based escalation?
- Context-sensitive escalation uses the on-call engineer's location (context) to determine who should be paged next, always routing to the nearest team member
- Context-sensitive escalation considers the severity, service type, and current conditions (e.g., high-traffic event, recent deployment) to escalate differently depending on context — a high-traffic event may trigger faster escalation than a normal business day (Correct answer)
- Context-sensitive escalation delays all escalations by a configurable amount based on the time of day, reducing middle-of-the-night pages
- Context-sensitive escalation requires the on-call engineer to explain the incident context before escalation is approved, ensuring only valid escalations are made
Correct answer: Context-sensitive escalation considers the severity, service type, and current conditions (e.g., high-traffic event, recent deployment) to escalate differently depending on context — a high-traffic event may trigger faster escalation than a normal business day
Simple time-based escalation ('escalate if unresolved after 30 minutes') is inflexible. Context-sensitive escalation adjusts escalation urgency based on incident severity, business context (peak traffic periods), recent changes, and affected user tier.
Question 12: What is 'horizontal vs. vertical scaling,' and when is each approach MOST appropriate?
- Horizontal and vertical scaling are equivalent in cost and complexity; the choice is arbitrary
- Horizontal scaling is always superior to vertical scaling for all types of services because modern cloud infrastructure makes it costless
- Vertical scaling is always preferred first because adding instances introduces coordination complexity and network overhead
- Horizontal scaling adds more instances of a service; vertical scaling increases the resources (CPU, memory) of existing instances. Horizontal scaling is preferred for stateless services and enables near-infinite capacity; vertical scaling is simpler but has hardware limits and causes downtime during upgrades (Correct answer)
Correct answer: Horizontal scaling adds more instances of a service; vertical scaling increases the resources (CPU, memory) of existing instances. Horizontal scaling is preferred for stateless services and enables near-infinite capacity; vertical scaling is simpler but has hardware limits and causes downtime during upgrades
Horizontal scaling (scale out) adds instances and distributes load — works well for stateless services but requires load balancing and may introduce distributed system complexity. Vertical scaling (scale up) is simpler but hits hardware ceilings and typically requires a restart.
Question 13: What is the purpose of a change review board (CRB)?
- To handle customer service.
- To approve or deny changes based on impact analysis (Correct answer)
- To delay all changes.
- To automate testing.
Correct answer: To approve or deny changes based on impact analysis
A Change Review Board (CRB) is a formal group responsible for evaluating proposed changes to IT systems or services. Its primary purpose is to assess the potential impact and risks of these changes, ensuring they align with business objectives and minimize disruption. By approving or denying changes based on a thorough impact analysis, the CRB helps maintain system stability and service quality.
Question 14: What is a Kubernetes NetworkPolicy, and what is the default behavior if no NetworkPolicy exists in a namespace?
- NetworkPolicies control which nodes pods can be scheduled on based on network topology
- NetworkPolicies are firewall rules for pods; by default (no NetworkPolicy), all pods can communicate freely with all other pods in the cluster — the default is allow-all (Correct answer)
- NetworkPolicies only apply to traffic entering the cluster from external sources, not pod-to-pod traffic
- NetworkPolicies are the default configuration — if none exist, all pod-to-pod communication is blocked
Correct answer: NetworkPolicies are firewall rules for pods; by default (no NetworkPolicy), all pods can communicate freely with all other pods in the cluster — the default is allow-all
By default, Kubernetes allows all pod-to-pod communication across all namespaces (open network). NetworkPolicies selectively restrict this by defining ingress/egress rules — but require a CNI plugin that supports them (e.g., Calico, Cilium) to take effect.
Question 15: What does Amdahl's Law imply for SRE capacity planning when scaling parallel systems?
- Network bandwidth is always the primary bottleneck
- Doubling CPU cores always halves processing time
- Performance scales linearly with added resources
- The speedup from parallelization is limited by the sequential (non-parallelizable) fraction of the workload (Correct answer)
Correct answer: The speedup from parallelization is limited by the sequential (non-parallelizable) fraction of the workload
Amdahl's Law states that sequential portions of a workload cap the maximum speedup achievable through parallelization, setting an upper bound on scaling benefits.
Question 16: Which capacity planning concept describes the point where adding more resources produces diminishing or negative returns due to coordination overhead?
- Saturation point
- Universal Scalability Law (USL) retrograde region (Correct answer)
- Coherency penalty
- Amdahl ceiling
Correct answer: Universal Scalability Law (USL) retrograde region
The Universal Scalability Law models throughput vs. concurrency and predicts a retrograde region where coordination costs cause throughput to decrease with more resources.
Question 17: What distinguishes a 'playbook' from a 'runbook' in some SRE organizations?
- Playbooks provide high-level strategy and decision trees; runbooks provide specific step-by-step procedures (Correct answer)
- There is no meaningful distinction between the two terms
- Playbooks are for developers; runbooks are for operators
- Playbooks are automated; runbooks are always manual
Correct answer: Playbooks provide high-level strategy and decision trees; runbooks provide specific step-by-step procedures
Some teams use 'playbook' to mean a higher-level guide with decision logic, while 'runbook' refers to a specific ordered set of executable steps for a known failure mode.
Question 18: What is the relationship between service reliability and feature velocity in the SRE model?
- Reliability always takes priority because outages cost more than delayed features
- Reliability and feature velocity are fundamentally opposed, and organizations must permanently choose one as their primary goal
- Feature velocity always takes priority because unreleased features generate no business value
- They are balanced through the error budget: when the budget is healthy, teams can move fast; when it is exhausted, reliability work takes priority over features (Correct answer)
Correct answer: They are balanced through the error budget: when the budget is healthy, teams can move fast; when it is exhausted, reliability work takes priority over features
The error budget is the mechanism that balances these competing priorities dynamically. Healthy budget = release freely; exhausted budget = pause and fix. Neither always wins — the budget decides.
Question 19: What is the purpose of a 'capacity model' in SRE practice?
- Tracking incident resolution times
- Documenting on-call rotation schedules
- Mapping resource consumption to traffic levels to forecast infrastructure needs (Correct answer)
- Defining SLOs for each service tier
Correct answer: Mapping resource consumption to traffic levels to forecast infrastructure needs
A capacity model quantifies the relationship between traffic/usage and resource consumption, enabling accurate infrastructure forecasting as load grows.
Question 20: Which Git workflow practice helps avoid long-lived branches and supports frequent integration?
- Gitflow with release branches
- Trunk-based development (Correct answer)
- Fork-and-merge strategy
- Monorepo with per-team branches
Correct answer: Trunk-based development
Trunk-based development requires developers to commit to the main branch frequently, reducing merge conflicts and integration risk.
Question 21: What is the 'reliability hierarchy' in SRE, and what does it imply about incident priorities?
- Customer satisfaction → Features → Performance → Availability — reliability priorities should follow business value
- SLA → SLO → SLI → Error Budget — each layer constrains the one below it in importance
- Hardware → Network → OS → Application → User — incidents should always be debugged from the bottom layer up
- Monitoring → Incident Response → Postmortem → Automation → Capacity Planning — each layer must be solid before the next adds value; an incident cannot be well-managed if monitoring is blind (Correct answer)
Correct answer: Monitoring → Incident Response → Postmortem → Automation → Capacity Planning — each layer must be solid before the next adds value; an incident cannot be well-managed if monitoring is blind
The SRE reliability pyramid establishes that without solid monitoring, incident response is blind; without good incident response, postmortems lack data; without postmortems, automation lacks direction — each level depends on the foundation below it.
Question 22: What is a 'post-deployment smoke test' in a CI/CD pipeline?
- A manual walkthrough by the QA team
- A security scan of the deployment environment
- A minimal set of critical checks run immediately after deployment to verify the application is functional (Correct answer)
- A performance load test run weeks after deployment
Correct answer: A minimal set of critical checks run immediately after deployment to verify the application is functional
Smoke tests are lightweight, fast checks that verify core functionality is working right after deployment before deeper validation.
Question 23: Why should SLOs be set based on what users actually need rather than on what the system can currently achieve?
- System-capability-based SLOs are always lower than user-need SLOs, leading to SLA breaches
- User-need-based SLOs are required by industry compliance frameworks like ISO 27001
- Setting SLOs based on current capability locks in the status quo and may over-invest in reliability that users don't value, while user-need-based SLOs create the right reliability incentives (Correct answer)
- User surveys are the only valid input for SLO setting because engineers cannot estimate reliability
Correct answer: Setting SLOs based on current capability locks in the status quo and may over-invest in reliability that users don't value, while user-need-based SLOs create the right reliability incentives
SLOs calibrated to current capability lock in technical debt and over-engineering simultaneously. User-need-based SLOs create a clear target: meet the minimum reliability users require, invest the rest in innovation.
Question 24: Which of the following best describes 'RED' method metrics in SRE observability?
- Rate, Errors, Duration (Correct answer)
- Requests, Events, Delays
- Response, Error, Dependency
- Reliability, Efficiency, Durability
Correct answer: Rate, Errors, Duration
The RED method focuses on Rate (requests per second), Errors (failed requests), and Duration (distribution of request latencies) for services.
Question 25: When an SRE team says 'toil has a 50% cap,' what happens if the team consistently exceeds this cap?
- Product teams should be asked to slow feature development
- The cap should be raised to match actual workload
- Engineers should work overtime to complete both toil and project work
- The team should escalate to management to either increase automation investment or reduce service scope/SLO commitments (Correct answer)
Correct answer: The team should escalate to management to either increase automation investment or reduce service scope/SLO commitments
Exceeding the toil cap is a signal that the team is under-resourced for automation investment or overextended in service ownership — both require management action.
Question 26: What is a common drawback of not using automation in infrastructure management?
- Increased human errors and inconsistent environments (Correct answer)
- Improved error handling.
- Greater scalability.
- Reduced configuration visibility.
Correct answer: Increased human errors and inconsistent environments
Without automation, infrastructure management relies heavily on manual processes, which are prone to human error, such as typos or missed steps. This often leads to 'configuration drift,' where environments that should be identical become inconsistent over time. Such inconsistencies make troubleshooting difficult, increase deployment times, and reduce overall system reliability, leading to increased operational burden and instability.
Question 27: What does 'mean time to restore' (MTTR) measure in the context of deployments?
- Average time to recover service after a deployment-caused outage (Correct answer)
- Time between successive deployments
- Average time for a CI build to complete
- Average time to deploy a new release
Correct answer: Average time to recover service after a deployment-caused outage
MTTR measures how quickly a team can restore service after an incident, including those caused by bad deployments.
Question 28: What is 'on-call shadowing' and why is it a recommended practice for new team members?
- On-call shadowing is a security audit practice where senior engineers secretly monitor on-call engineers' actions
- On-call shadowing pairs new team members with experienced on-call engineers, allowing them to observe incident response and use runbooks in real situations before taking primary on-call responsibility (Correct answer)
- On-call shadowing requires new team members to work back-to-back on-call shifts to rapidly build experience
- On-call shadowing means the on-call engineer should always have a senior engineer available on video call during their shift
Correct answer: On-call shadowing pairs new team members with experienced on-call engineers, allowing them to observe incident response and use runbooks in real situations before taking primary on-call responsibility
Shadowing provides safe, low-stakes learning — new team members see real incidents handled, understand how runbooks are used in practice, and build confidence before they hold primary on-call responsibility.
Question 29: What is 'autoscaling,' and what metric is MOST appropriate to trigger autoscaling for a web API service?
- Autoscaling is only appropriate for stateless services; stateful services must be manually scaled to avoid data consistency issues
- Autoscaling should always use CPU utilization as the trigger because it is the most accurate proxy for user-facing load across all service types
- Autoscaling based on memory utilization is most appropriate because memory is the limiting resource for most web applications
- Autoscaling adjusts the number of running instances based on demand; for a web API, scaling based on requests per second (RPS) or concurrent connections is more appropriate than CPU utilization alone, as some APIs are I/O-bound rather than CPU-bound (Correct answer)
Correct answer: Autoscaling adjusts the number of running instances based on demand; for a web API, scaling based on requests per second (RPS) or concurrent connections is more appropriate than CPU utilization alone, as some APIs are I/O-bound rather than CPU-bound
CPU-only autoscaling fails for I/O-bound services where CPU stays low even under heavy load because threads are waiting on database or network I/O. Request rate or concurrent connection metrics better represent actual service demand.
Question 30: What is 'right-sizing' in cloud capacity management?
- Matching instance or resource size to actual workload requirements to minimize waste (Correct answer)
- Scaling all services to the same instance type for consistency
- Provisioning at the maximum possible size for safety
- Choosing the cheapest available instance type
Correct answer: Matching instance or resource size to actual workload requirements to minimize waste
Right-sizing analyzes actual resource consumption and selects the instance type or size that meets performance requirements without over-provisioning.
Question 31: What does 'mean time between failures' (MTBF) measure, and why is it potentially misleading as a primary reliability metric?
- MTBF measures the average time between incidents; it is misleading because two services with the same MTBF but very different incident durations (MTTR) will have vastly different availability (Correct answer)
- MTBF measures the frequency of deployment changes; it is misleading because not all changes cause failures
- MTBF measures the time to deploy a fix after a failure is detected; it is misleading because it doesn't capture detection time
- MTBF measures the predicted hardware lifespan; it is misleading for software services because software doesn't wear out
Correct answer: MTBF measures the average time between incidents; it is misleading because two services with the same MTBF but very different incident durations (MTTR) will have vastly different availability
MTBF alone doesn't capture availability. A service with MTBF of 30 days but MTTR of 1 day has very different availability than one with MTBF of 30 days but MTTR of 1 minute. Availability = MTBF / (MTBF + MTTR).
Question 32: Which tool is useful for tracking incidents and changes?
- Email only.
- Jira or ServiceNow for traceability (Correct answer)
- YouTube.
- MS Paint.
Correct answer: Jira or ServiceNow for traceability
Tools like Jira and ServiceNow are invaluable for tracking incidents and changes due to their comprehensive capabilities. They provide centralized platforms for logging, categorizing, assigning, and monitoring the entire lifecycle of incidents, problems, and changes. This ensures clear traceability, accountability, and provides historical data essential for analysis and continuous improvement in IT operations.
Question 33: What is 'cloud waste from over-provisioned Kubernetes resources' and how is it detected and reduced?
- Over-provisioned Kubernetes resources are detected by counting the number of pods running — too many pods indicates over-provisioning
- Kubernetes clusters always use 100% of provisioned capacity because the scheduler fills all available space — over-provisioning is not possible in Kubernetes
- Pods with resource requests significantly above their actual consumption waste cluster capacity by reserving CPU and memory that goes unused; detected by comparing requested vs. actual usage via metrics, and reduced by Vertical Pod Autoscaler (VPA) or manual request right-sizing (Correct answer)
- Kubernetes resource waste is eliminated by setting all pod resource requests to zero, letting the scheduler allocate resources dynamically
Correct answer: Pods with resource requests significantly above their actual consumption waste cluster capacity by reserving CPU and memory that goes unused; detected by comparing requested vs. actual usage via metrics, and reduced by Vertical Pod Autoscaler (VPA) or manual request right-sizing
Pod resource requests determine cluster node provisioning. If pods request 4 vCPU but average 0.2 vCPU, nodes are provisioned for the 4 vCPU request — wasting 95% of allocated (and paid-for) compute.
Question 34: An SRE notices that every deployment requires manual SSH into servers to update config files. What practice should they implement?
- Add a reminder step in the deployment checklist
- Delegate manual config updates to a dedicated ops team
- Configuration as code — version-control config and apply it automatically as part of the deployment pipeline (Correct answer)
- Increase the deployment frequency to reduce config drift
Correct answer: Configuration as code — version-control config and apply it automatically as part of the deployment pipeline
Configuration as code ensures config changes are versioned, reviewed, and deployed automatically alongside application changes.
Question 35: Which scaling pattern is most appropriate when a system's bottleneck is stateful session data?
- Vertical scaling of the web tier only
- Adding more read replicas to the database
- Increasing CDN cache TTLs
- Horizontal scaling with sticky sessions or a shared session store (Correct answer)
Correct answer: Horizontal scaling with sticky sessions or a shared session store
Stateful sessions require either sticky sessions to route users to the same instance or a shared session store (like Redis) to allow any instance to serve any user.
Question 36: Which logging anti-pattern causes the most problems during high-severity incidents?
- Using a centralized log aggregation system
- Including timestamps in every log line
- Using structured JSON format
- Logging at DEBUG level in production, causing I/O saturation under load (Correct answer)
Correct answer: Logging at DEBUG level in production, causing I/O saturation under load
Verbose DEBUG logging in production can saturate disk I/O and CPU under load, worsening incidents while making logs harder to search.
Question 37: What is the benefit of using 'alert ownership' metadata (such as team labels) in an alerting system?
- It ensures alerts are routed to the correct on-call team with the knowledge to resolve the issue (Correct answer)
- It satisfies regulatory requirements for incident attribution in compliance-heavy industries
- It allows alert thresholds to be dynamically adjusted based on each team's error budget balance
- It enables automatic SLO calculation by attributing alert frequency to specific engineering teams
Correct answer: It ensures alerts are routed to the correct on-call team with the knowledge to resolve the issue
Alert ownership metadata ensures pages are routed to the team responsible for that service, reducing response time and preventing confusion about who should be investigating a given alert.
Question 38: Which of the following best describes a key competency required for on-call practices & runbooks in SRE practice?
- Reliance on a single methodology for all situations
- The ability to work independently without any oversight
- Strong analytical skills combined with effective communication and ethical judgment (Correct answer)
- Memorization of all relevant regulations without understanding context
Correct answer: Strong analytical skills combined with effective communication and ethical judgment
SRE professionals working in on-call practices & runbooks need analytical skills to assess situations, communication skills to convey findings, and ethical judgment to make sound decisions.
Question 39: Why does Google's SRE model argue that having a separate SRE team, rather than embedding reliability work in development teams, is beneficial?
- Having a separate team reduces the headcount needed in the development organization
- Regulatory requirements mandate that operations and development teams be separated
- SREs have unique technical skills that cannot be learned by developers
- A dedicated SRE team creates structural incentives to prioritize reliability — developers are motivated to ship features, while SREs are specifically accountable for stability and operational excellence (Correct answer)
Correct answer: A dedicated SRE team creates structural incentives to prioritize reliability — developers are motivated to ship features, while SREs are specifically accountable for stability and operational excellence
Google's SRE model creates a team whose primary incentive and success metric is reliability, counterbalancing development teams whose primary incentive is feature velocity. The structural separation aligns incentives with the organization's reliability goals.
Question 40: An SLA promises customers 99.5% monthly availability with a 10% service credit for each additional 0.5% of downtime. The service experienced 5 hours of downtime in a 30-day month. Was the SLA breached?
- Yes — 99.5% of 43,200 minutes allows only 216 minutes (3.6 hours) of downtime; 5 hours (300 minutes) exceeds this (Correct answer)
- No — SLA calculations exclude weekends and after-hours periods by default
- Yes — any downtime in a month automatically breaches a 99.5% SLA
- No — 5 hours is within the 99.5% threshold for a 30-day month
Correct answer: Yes — 99.5% of 43,200 minutes allows only 216 minutes (3.6 hours) of downtime; 5 hours (300 minutes) exceeds this
30 days × 24 hours × 60 minutes = 43,200 minutes. 0.5% × 43,200 = 216 minutes (3.6 hours) allowed. 5 hours = 300 minutes, which exceeds 216 minutes, so the SLA is breached.
SRE Foundation Certification Exam
The SRE Foundation Certification Exam validates an individual's understanding of the core principles and practices of Site Reliability Engineering (SRE).
Exam Rules
- You can skip questions and return to them later
- Flag questions for review before submitting
- No feedback shown until you submit the entire exam
- Unanswered questions count as wrong — answer everything
- 10 pretest questions are mixed in and don't affect your score
- Timer auto-submits when time runs out
- Your progress is auto-saved every 30 seconds