PMMockr

QuestionsRoot Cause AnalysisGoogle

A key metric for Google Cloud spiked unexpectedly. How do you determine if it is healthy

Problem Statement Description

Product context: Google is a consumer technology, ads, AI, and cloud company; its products include Search, YouTube, Android, Maps, Gmail, Chrome, Google Play, Workspace, and Google Cloud.

Google Cloud has observed an unexpected spike in a key product or business metric. The spike may represent a positive development, such as genuine customer adoption, increased workload usage, successful go-to-market activity, or growth from new internet users and emerging digital businesses. It may also reflect an unhealthy issue, such as instrumentation errors, abusive activity, pricing anomalies, duplicate events, service retries, or short-lived usage that does not translate into durable value.

Your task is to walk through how you would determine whether the spike is healthy. Treat this as a root-cause analysis discussion for a large-scale cloud platform where metrics may span enterprise customers, developers, APIs, infrastructure usage, billing events, regions, products, and customer segments. The interviewer is looking for how you structure ambiguity, validate the data, segment the change, and distinguish sustainable growth from noise or risk.

You do not need to know the exact internal Google Cloud metric. You should define what information you would request, how you would inspect the spike, what hypotheses you would test, and how your findings would inform next steps for product, engineering, sales, trust and safety, or operations teams.

The experience should consider:

- How you would frame the anomaly: metric definition, baseline, expected seasonality, magnitude, timing, and whether the spike is absolute, percentage-based, or cohort-specific.

- How you would verify instrumentation health, including logging changes, pipeline delays, duplicate events, schema changes, dashboard bugs, or newly added traffic sources.

- Which segmentations matter for Google Cloud, such as product area, API, region, customer type, account age, workload type, acquisition channel, billing status, and new versus existing customers.

- How you would distinguish healthy adoption from unhealthy behavior, including trial abuse, bot activity, repeated retries, misconfigured workloads, one-time migrations, or anomalous enterprise usage.

- What supporting metrics you would inspect, such as retention, paid conversion, revenue quality, error rates, quota usage, support tickets, latency, churn risk, and customer satisfaction.

- How you would gather evidence across data, customer feedback, sales context, incident reports, and engineering changes without jumping to a conclusion.

- What short-term mitigations or monitoring you would consider if the spike creates operational, financial, privacy, security, or customer-trust risk.

- How you would prevent recurrence through better alerting, metric ownership, anomaly detection, runbooks, and post-incident learning.

The goal is to demonstrate a clear, structured RCA approach that can determine whether the spike reflects real, valuable Google Cloud growth or an issue that requires investigation, mitigation, or correction.

What this question tests

Practise this question under interview conditions. Answer it out loud against a timer with an AI interviewer that asks follow-ups, then review the scored report.

Start a timed mock interview

Related Root Cause Analysis questions

All Root Cause Analysis questions · Product manager interview questions by skill area