Critical alerts page StarTree on-call 24/7. Warning alerts reach StarTree’s alert channels without paging anyone. Thresholds on these pages are defaults; StarTree may tune them for your environment.
Start from the symptom
New data is late or missing
Real-time ingestion stopped, lagging, or dropping rows. 9 alerts.
Queries fail, time out, or get rejected
Error rates, timeouts, query quota, broker availability. 12 alerts.
Batch data or a scheduled task didn't land
File, SQL, segment import, and external table sync tasks, and minion capacity. 20 alerts.
Results are incomplete, inconsistent, or near a limit
Serving replicas, segments in error, disk, key, and storage limits. 11 alerts.
Upsert results or storage look wrong
Primary-key growth, compaction, snapshots, replica consistency. 7 alerts.
Components restart, run hot, or infrastructure is unhealthy
Out-of-memory kills, CPU, GC, ZooKeeper, Kubernetes, platform services. 100 alerts.
How to read an alert entry
Each critical alert on the category pages has its own entry:
Some entries carry a note. Often transient means the alert usually clears on its own. Usually your fix means the cause is most often in your stream, source system, or table config. Warning alerts are listed in a table on each page, and every page ends with the metrics each alert evaluates.
Grafana dashboards named on these pages are described in Grafana dashboards. API paths are Pinot controller endpoints, available from Cluster Manager → Swagger API in the Data Portal.
Alert index
All 159 alerts, 119 critical and 40 warning, in alphabetical order.Reading Pinot metrics
Your cluster’s Pinot metrics are available in Grafana, through the Prometheus data source. Each alert page lists the metrics worth watching for its area and the series each alert evaluates. This section explains how those names are built, so you can find the one you need rather than guess.Naming convention
Metric names follow a consistent shape:
So a real-time ingestion delay gauge on a server becomes
pinot_server_realtimeIngestionDelayMs_Value, and a broker query counter becomes pinot_broker_queries_Count.
A few metrics do not follow the pattern. RocksDB metrics, used by off-heap upsert, appear as pinot_server_rocksdb_ticker_<name> and pinot_server_rocksdb_histogram_<name>_<stat>. JVM metrics such as jvm_gc_collection_seconds_sum and jvm_threads_state keep their standard JVM names and carry a component label instead of a pinot_ prefix.
Suffixes
Not every statistic is available for every metric. If a suffix such as
_Mean or _FifteenMinuteRate returns nothing, use one from the table above instead — _50thPercentile for a typical value, or a rate over _Count for throughput. Confirm what a metric actually offers by running the bare name in Explore and reading the series that come back.
Labels
Metrics are labelled so you can slice them without name explosion:
Some metrics also carry pod-level labels such as
pod or kubernetes_pod_name. These identify the component instance rather than a table, so resource questions (“which server is hot?”) are answered by grouping on those, and workload questions (“which table is slow?”) by grouping on table.
The
table label always carries the type suffix. table="orders_REALTIME" and table="orders_OFFLINE" are separate series for a hybrid table, so a per-table chart of a hybrid table needs both, or a table=~"orders_.*" matcher.Finding a metric that is not listed
The lists above are curated, not exhaustive — Pinot exports far more than this. To find something specific:1
Browse in Grafana Explore
Open Grafana, choose Explore, and select the Prometheus data source. Type
pinot_server_ in the metric field and the metric browser will complete against everything currently present in your cluster. This is the fastest way to check whether a metric exists at all.2
Search by fragment
Metric names are camelCase inside a snake_case wrapper, so search on a distinctive fragment —
Upsert, Ingestion, Segment — rather than a whole name. Prometheus’s metric browser matches on substring.3
Check the labels before charting
Run the bare metric name first and look at the label set on the returned series. Whether a metric is per-table, per-partition, or per-pod determines whether you need
sum by (table) or avg by (pod) — and getting this wrong is the most common cause of a chart that looks alarming and means nothing.4
Copy from an existing panel
Every panel on the built-in dashboards exposes its query. Open the panel menu, choose Explore, and you get a working query with the right labels and aggregation to modify.
Related
- Health checks and Table health and email alerts: Data Portal checks and email alerts, separate from these Prometheus alerts.
- Grafana dashboards: the built-in dashboards over these metrics.
- Accessing logs: when a metric tells you something happened and you need the log line.
- Query Logger: per-query records rather than aggregates.
- Troubleshooting: start here: step-by-step triage when something is wrong.

