Skip to main content
Freshness and completeness of data coming from Kafka, Kinesis and other streams. These page first because stale data is usually what users notice.
Critical alerts page StarTree on-call 24/7. Warning alerts reach StarTree’s alert channels without paging anyone. Thresholds are defaults; StarTree may tune them for your environment. See the overview for how to read an entry.

Decision tree

Follow the tree from the symptom to the alert most likely to explain it. Use the zoom controls at the top right of the diagram to enlarge it.

Critical alerts

RealtimeIngestionStopped

Often transient: this alert usually clears on its own.

HighRealtimeIngestionOffsetLag

NoRealtimeIngestionPerServer

Often transient: this alert usually clears on its own.

PartitionConsumerCreateExceptions

Usually your fix: the cause is most often in your stream, source system, or table config.

RealtimeConsumptionExceptions

RealtimeIngestionHighRowsWithErrorsRate

Usually your fix: the cause is most often in your stream, source system, or table config.

StreamDataLoss

Usually your fix: the cause is most often in your stream, source system, or table config.

Warning alerts

Metrics

The series below are available in Grafana through the Prometheus data source. For how names, suffixes, and labels are built, see Reading Pinot metrics.

Metrics worth watching

These are the ones to reach for first; they are not the complete set.

Metrics behind each alert

The Prometheus series each alert rule evaluates.