Skip to main content
Whether queries succeed and return on time. Error-rate alerts count only Pinot-side errors, not bad SQL.
Critical alerts page StarTree on-call 24/7. Warning alerts reach StarTree’s alert channels without paging anyone. Thresholds are defaults; StarTree may tune them for your environment. See the overview for how to read an entry.

Decision tree

Follow the tree from the symptom to the alert most likely to explain it. Use the zoom controls at the top right of the diagram to enlarge it.

Critical alerts

HighQueryCriticalErrorRate

HighSystemQueryErrorRate

HighQueryTimeoutPercentage

QueryQuotaAlert

Usually your fix: the cause is most often in your stream, source system, or table config.

LowBrokerAvailability

NoBrokerAvailable

BrokerDataTableDeserializationExceptions

NettyResponseFetchExceptions

Warning alerts

Metrics

The series below are available in Grafana through the Prometheus data source. For how names, suffixes, and labels are built, see Reading Pinot metrics.

Metrics worth watching

These are the ones to reach for first; they are not the complete set. For per-query rather than aggregate analysis, use system_query_log — metrics tell you the shape of the problem, the query log tells you which query.

Metrics behind each alert

The Prometheus series each alert rule evaluates.