Skip to main content
This page covers the Pinot components themselves, then the infrastructure underneath them. Most of these alerts are handled by StarTree end to end.
Critical alerts page StarTree on-call 24/7. Warning alerts reach StarTree’s alert channels without paging anyone. Thresholds are defaults; StarTree may tune them for your environment. See the overview for how to read an entry.

Cluster health

Pinot controllers, brokers, servers, minions and ZooKeeper: restarts, memory, CPU and coordination.

Decision tree

Follow the tree from the symptom to the alert most likely to explain it. Use the zoom controls at the top right of the diagram to enlarge it.

Critical alerts

NoControllerLeaderAlert

PinotControllerHighThreadsBlocked

PinotHighGcPercent

PinotComponentHighCPUUsage

Often transient: this alert usually clears on its own.

HighServerPodsRestarts

HighNonServerPodsRestarts

KubernetesPinotContainerOOMKilled

KubernetesPinotPodCrashLooping

PinotPodWaiting

PinotPodStuckInContainerCreating

PinotPodImagePullErrors

HighAvailablePinotStatefulsetReplicasMismatch

ZookeeperErrorCount

ZookeeperLeaderElectionTooLong

ZookeeperIncorrectNumberOfLeaders

ZookeeperIncorrectQuorumSize

ZookeeperMetricsMissing

Warning alerts

Infrastructure StarTree manages

StarTree handles these alerts end to end. They’re listed so you know what is watched; none of them needs action from you.

Metrics

The series below are available in Grafana through the Prometheus data source. For how names, suffixes, and labels are built, see Reading Pinot metrics.

Metrics worth watching

These are the ones to reach for first; they are not the complete set.

Metrics behind each alert

The Prometheus series each alert rule evaluates, for the Pinot component alerts.