Critical alerts page StarTree on-call 24/7. Warning alerts reach StarTree’s alert channels without paging anyone. Thresholds are defaults; StarTree may tune them for your environment. See the overview for how to read an entry.
Cluster health
Pinot controllers, brokers, servers, minions and ZooKeeper: restarts, memory, CPU and coordination.Decision tree
Follow the tree from the symptom to the alert most likely to explain it. Use the zoom controls at the top right of the diagram to enlarge it.Critical alerts
NoControllerLeaderAlert
PinotControllerHighThreadsBlocked
PinotHighGcPercent
PinotComponentHighCPUUsage
Often transient: this alert usually clears on its own.
HighServerPodsRestarts
HighNonServerPodsRestarts
KubernetesPinotContainerOOMKilled
KubernetesPinotPodCrashLooping
PinotPodWaiting
PinotPodStuckInContainerCreating
PinotPodImagePullErrors
HighAvailablePinotStatefulsetReplicasMismatch
ZookeeperErrorCount
ZookeeperLeaderElectionTooLong
ZookeeperIncorrectNumberOfLeaders
ZookeeperIncorrectQuorumSize
ZookeeperMetricsMissing
Warning alerts
Infrastructure StarTree manages
StarTree handles these alerts end to end. They’re listed so you know what is watched; none of them needs action from you.Kubernetes nodes and network (7 critical, 5 warning)
Kubernetes nodes and network (7 critical, 5 warning)
Kubernetes pods and volumes (6 critical, 5 warning)
Kubernetes pods and volumes (6 critical, 5 warning)
Platform services (16 critical, 5 warning)
Platform services (16 critical, 5 warning)
API gateway (Traefik) (5 critical, 2 warning)
API gateway (Traefik) (5 critical, 2 warning)
Certificates (3 critical, 2 warning)
Certificates (3 critical, 2 warning)
Monitoring stack (11 critical, 1 warning)
Monitoring stack (11 critical, 1 warning)
ThirdEye (9 critical, 3 warning)
ThirdEye (9 critical, 3 warning)
Metrics
The series below are available in Grafana through the Prometheus data source. For how names, suffixes, and labels are built, see Reading Pinot metrics.Metrics worth watching
These are the ones to reach for first; they are not the complete set.Metrics behind each alert
The Prometheus series each alert rule evaluates, for the Pinot component alerts.Related
- Alerts and metrics overview: severity, the symptom router, the alert index, and how metric names are built.
- Cluster operations troubleshooting
- Grafana dashboards

