Engineering note · 2018 · Kubernetes / Prometheus
Detecting Kubernetes API latency
An archived account of investigating alert rules for sudden API server latency. The source describes the reasoning, query choices, and a way to reproduce the condition.
The problem
A large query against Kubernetes objects could increase API server response time for other users. An alert based on a long summary window could respond too slowly to a sudden spike.
Song’s contribution
In the 2018 note, I worked through the metric choice for this alert. I compared a summary metric with a histogram bucket query, looked into the metric definitions, and documented a way to generate many objects to exercise the API server.
Approach and evidence
The first query used apiserver_request_latencies_summary. The note explains why its long sliding window made it a poor fit for a responsive alert. The revised approach used apiserver_request_latencies_bucket with histogram_quantile and a five-minute rate window.
The public article contains the query details and the test setup. It does not provide a measured production result or a claim about alert deployment.