Skip to main content
StarTree runs 159 Prometheus alerts on every StarTree Cloud environment. They’re split across five pages by what you’d notice, plus one page for the platform underneath. Use this page to find the right one, look an alert up by name in the index below, or learn how the Pinot metrics behind them are named.
Critical alerts page StarTree on-call 24/7. Warning alerts reach StarTree’s alert channels without paging anyone. Thresholds on these pages are defaults; StarTree may tune them for your environment.

Start from the symptom

New data is late or missing

Real-time ingestion stopped, lagging, or dropping rows. 9 alerts.

Queries fail, time out, or get rejected

Error rates, timeouts, query quota, broker availability. 12 alerts.

Batch data or a scheduled task didn't land

File, SQL, segment import, and external table sync tasks, and minion capacity. 20 alerts.

Results are incomplete, inconsistent, or near a limit

Serving replicas, segments in error, disk, key, and storage limits. 11 alerts.

Upsert results or storage look wrong

Primary-key growth, compaction, snapshots, replica consistency. 7 alerts.

Components restart, run hot, or infrastructure is unhealthy

Out-of-memory kills, CPU, GC, ZooKeeper, Kubernetes, platform services. 100 alerts.

How to read an alert entry

Each critical alert on the category pages has its own entry: Some entries carry a note. Often transient means the alert usually clears on its own. Usually your fix means the cause is most often in your stream, source system, or table config. Warning alerts are listed in a table on each page, and every page ends with the metrics each alert evaluates. Grafana dashboards named on these pages are described in Grafana dashboards. API paths are Pinot controller endpoints, available from Cluster Manager → Swagger API in the Data Portal.

Alert index

All 159 alerts, 119 critical and 40 warning, in alphabetical order.

Reading Pinot metrics

Your cluster’s Pinot metrics are available in Grafana, through the Prometheus data source. Each alert page lists the metrics worth watching for its area and the series each alert evaluates. This section explains how those names are built, so you can find the one you need rather than guess.
A metric that has never been recorded does not exist. If no value has been published — a table that has never been queried, a task type that has never run — the metric is absent rather than present with a zero. An empty chart can mean “nothing happened,” not “the metric is wrong.”

Naming convention

Metric names follow a consistent shape:
So a real-time ingestion delay gauge on a server becomes pinot_server_realtimeIngestionDelayMs_Value, and a broker query counter becomes pinot_broker_queries_Count. A few metrics do not follow the pattern. RocksDB metrics, used by off-heap upsert, appear as pinot_server_rocksdb_ticker_<name> and pinot_server_rocksdb_histogram_<name>_<stat>. JVM metrics such as jvm_gc_collection_seconds_sum and jvm_threads_state keep their standard JVM names and carry a component label instead of a pinot_ prefix.

Suffixes

Not every statistic is available for every metric. If a suffix such as _Mean or _FifteenMinuteRate returns nothing, use one from the table above instead — _50thPercentile for a typical value, or a rate over _Count for throughput. Confirm what a metric actually offers by running the bare name in Explore and reading the series that come back.

Labels

Metrics are labelled so you can slice them without name explosion: Some metrics also carry pod-level labels such as pod or kubernetes_pod_name. These identify the component instance rather than a table, so resource questions (“which server is hot?”) are answered by grouping on those, and workload questions (“which table is slow?”) by grouping on table.
The table label always carries the type suffix. table="orders_REALTIME" and table="orders_OFFLINE" are separate series for a hybrid table, so a per-table chart of a hybrid table needs both, or a table=~"orders_.*" matcher.

Finding a metric that is not listed

The lists above are curated, not exhaustive — Pinot exports far more than this. To find something specific:
1

Browse in Grafana Explore

Open Grafana, choose Explore, and select the Prometheus data source. Type pinot_server_ in the metric field and the metric browser will complete against everything currently present in your cluster. This is the fastest way to check whether a metric exists at all.
2

Search by fragment

Metric names are camelCase inside a snake_case wrapper, so search on a distinctive fragment — Upsert, Ingestion, Segment — rather than a whole name. Prometheus’s metric browser matches on substring.
3

Check the labels before charting

Run the bare metric name first and look at the label set on the returned series. Whether a metric is per-table, per-partition, or per-pod determines whether you need sum by (table) or avg by (pod) — and getting this wrong is the most common cause of a chart that looks alarming and means nothing.
4

Copy from an existing panel

Every panel on the built-in dashboards exposes its query. Open the panel menu, choose Explore, and you get a working query with the right labels and aggregation to modify.