← All Cloud Engineer Flashcard Decks

Case Studies & Practical Application Flashcards

7 cards from real Cloud Engineer practice questions. Tap to flip, then mark Knew It or Still Learning — missed cards come back until you master them.

Read the first 7 Case Studies & Practical Application flashcards as text
  1. A company's GCP project is unexpectedly billed $50,000 for BigQuery queries run by a data science team. No budget alerts fired. What two controls should be implemented immediately?

    Answer: Set BigQuery custom cost controls (maximum bytes billed per query) and create a GCP Budget with Pub/Sub alert actions

    Maximum bytes billed per query prevents runaway scans at execution time, while Pub/Sub-triggered budget alerts can automatically restrict access or notify ops when thresholds are crossed.

  2. An AWS Lambda function writes processed records to DynamoDB. Under load, the function receives ProvisionedThroughputExceededException errors. The DynamoDB table uses a single partition key with high cardinality. What is the issue?

    Answer: Lambda concurrency is sending too many requests to a DynamoDB hot partition due to non-uniform key access patterns

    Even with high-cardinality keys, if certain key values are accessed far more frequently, those partitions become hot and exceed their individual throughput limits regardless of total table capacity.

  3. A company runs a critical Azure SQL Database. During a DR test, they restore a geo-redundant backup to the secondary region and find data is 2 hours old. Their RPO requirement is 15 minutes. What change is needed?

    Answer: Enable Active Geo-Replication or Auto-Failover Groups, which provide near-real-time replication with RPO of ~5 seconds

    Geo-redundant backups have an RPO of 1-2 hours; Active Geo-Replication continuously replicates transactions to a secondary database, achieving RPO of seconds to meet the 15-minute SLA.

  4. A team uses Terraform to manage GCP infrastructure. After a colleague manually modified a Cloud SQL instance's tier in the GCP Console, the next `terraform plan` shows no changes. What explains this?

    Answer: The Terraform state file still reflects the old configuration; running `terraform refresh` updates state to match real infrastructure

    Terraform compares its state file against the declared configuration, not live infrastructure; `terraform refresh` (or `terraform plan -refresh=true`) syncs state with actual resource attributes, revealing the drift.

  5. A containerized app on AWS ECS Fargate needs to access a secret stored in AWS Secrets Manager. The developer hardcodes the secret in the Docker image as an environment variable for simplicity. What is the security risk and correct approach?

    Answer: Container images are often stored in ECR and inspectable; use ECS task definition secrets injection with IAM role-based access instead

    Environment variables baked into Docker images are visible in image layers, ECR image metadata, and ECS task definitions; secrets injection from Secrets Manager via IAM ensures secrets are fetched at runtime and never stored in the image.

  6. A startup's MongoDB Atlas cluster on GCP starts receiving read timeouts during business hours. Atlas charts show the primary node's opcounters are normal but query executor scanned 10M documents per query. What should the engineer do?

    Answer: Identify the queries using Atlas Performance Advisor and create compound indexes on the high-cardinality filter fields

    Scanning 10M documents per query is a collection scan; the correct fix is indexing the fields used in query filters, which reduces the scanned document count from millions to the result set size.

  7. A company migrates a stateful legacy app to Azure that stores session data in local files. After deploying to Azure App Service with 3 instances, users are randomly logged out. What is the root cause?

    Answer: Azure App Service load balances across instances, and session files on one instance are not available to the others

    Local file-based sessions are instance-local; when a user's subsequent request hits a different instance, that instance has no session file, appearing as a logout — the fix is centralized session storage like Azure Redis Cache.