PMMockr

QuestionsRoot Cause AnalysisGoogle

A key metric for Google Cloud spiked unexpectedly. How do you determine if it is healthy

Problem Statement Description

Product context: Google is a consumer technology, ads, AI, and cloud company; its products include Search, YouTube, Android, Maps, Gmail, Chrome, Google Play, Workspace, and Google Cloud.

You are the PM responsible for a Google Cloud product area, and an important business or product health metric has spiked unexpectedly. The metric could relate to usage, sign-ups, API calls, compute consumption, revenue, errors, support contacts, or another KPI that leadership monitors closely. Your task is to determine whether the spike represents healthy growth, a measurement issue, abnormal customer behavior, abuse, or a product/operational problem.

This is an RCA-style interview question. You should frame how you would investigate the anomaly across Google Cloud’s ecosystem of enterprise customers, developers, workloads, regions, channels, and customer segments, including users building for emerging or newly connected internet populations. The focus is not to jump to a conclusion, but to show a structured way to validate the metric, isolate the source of the spike, evaluate business and user impact, and decide what action is needed.

The experience should consider:

- How you would define the metric precisely, including numerator, denominator, time window, expected baseline, and what “spike” means.

- Initial instrumentation checks, such as logging changes, pipeline delays, duplicate events, schema changes, alerting thresholds, or dashboard bugs.

- Segmentation by product, region, customer type, workload, acquisition channel, pricing plan, platform, and cohort.

- How to distinguish healthy adoption from unhealthy signals such as fraud, bot activity, accidental customer behavior, quota misuse, outages, or cost surprises.

- Supporting evidence you would seek from adjacent metrics, such as retention, conversion, revenue, latency, error rates, support tickets, quota usage, billing disputes, and customer feedback.

- How you would prioritize hypotheses and identify the fastest checks to confirm or rule them out.

- Mitigation options if the spike is harmful, including customer communication, quota controls, rollback, alert tuning, or escalation to engineering, support, sales, security, or finance.

- Longer-term prevention, including improved monitoring, anomaly detection, ownership, documentation, and post-incident learning.

Your goal is to describe a clear, product-led investigation plan that helps Google Cloud decide whether the spike is a positive signal to amplify, a neutral reporting artifact to correct, or a risk that requires immediate operational response.

What this question tests

Practise this question under interview conditions. Answer it out loud against a timer with an AI interviewer that asks follow-ups, then review the scored report.

Start a timed mock interview

Related Root Cause Analysis questions

All Root Cause Analysis questions · Product manager interview questions by skill area