> ## Documentation Index
> Fetch the complete documentation index at: https://docs.startree.ai/llms.txt
> Use this file to discover all available pages before exploring further.

# Query alerts

> Alerts that watch query errors, timeouts, quota rejections, and broker availability: when each fires, what it means, and what you or StarTree can do.

Whether queries succeed and return on time. Error-rate alerts count only Pinot-side errors, not bad SQL.

<Info>
  **Critical** alerts page StarTree on-call 24/7. **Warning** alerts reach StarTree's alert channels without paging anyone. Thresholds are defaults; StarTree may tune them for your environment. See the [overview](/corecapabilities/observability/alerts/overview) for how to read an entry.
</Info>

## Decision tree

Follow the tree from the symptom to the alert most likely to explain it. Use the zoom controls at the top right of the diagram to enlarge it.

```mermaid placement="top-right" actions={true} theme={null}
flowchart TD
  A["Queries fail or are slow"] --> B{"What does the response say?"}
  B -- "429 / quota exceeded" --> B1["QueryQuotaAlert<br/>raise quota.maxQueriesPerSecond"]
  B -- "Timeout (250, 240, 400)" --> C{"Server CPU or GC high?"}
  C -- "Yes" --> C1["See Platform health<br/>scale or spread load"]
  C -- "No" --> C2["HighQueryTimeoutPercentage<br/>expensive query: rewrite or index"]
  B -- "Segments missing or unavailable (235, 305)" --> D["See Data availability"]
  B -- "Server not responding (427)" --> C1
  B -- "Connection refused / 5xx" --> E["LowBrokerAvailability<br/>NoBrokerAvailable"]
  B -- "Other internal errors" --> F["HighSystemQueryErrorRate<br/>Broker dashboard: Query Error Code Rate"]
```

## Critical alerts

### `HighQueryCriticalErrorRate`

| Field | Detail |
| - | - |
| **Fires when** | Over 2% of queries hit a critical error for 5 min, with more than 2 queries/s. |
| **What it means** | Queries are failing right now. |
| **Check first** | • Pinot - Broker → **Critical Error One Minute Rate**<br />• **Query Error Code Rate** to see which code |
| **What you can do** | If one query shape dominates, rewrite it |
| **What StarTree does** | • Restarts a faulty server<br />• Scales the cluster |

### `HighSystemQueryErrorRate`

| Field | Detail |
| - | - |
| **Fires when** | Pinot-side errors exceed 5% of queries for 15 min, with more than 10 queries/s. |
| **What it means** | Errors come from Pinot, not from your SQL. |
| **Check first** | • Pinot - Broker → **Query Error Code Rate**<br />• 235 / 305 segment missing · 427 server not responding · 250 execution timeout · 240 scheduling timeout · 400 broker timeout · 720 planning |
| **What you can do** | 400 or 720: simplify the query (for example, large IN lists) |
| **What StarTree does** | • Compares ideal state with external view<br />• Fixes the server behind the error code |

### `HighQueryTimeoutPercentage`

| Field | Detail |
| - | - |
| **Fires when** | Over 50% of a table's queries time out, for 5 min. |
| **What it means** | Timeouts are the main failure for that table. |
| **Check first** | • Execution timeouts with no server-not-responding errors: an expensive query<br />• Both rising together: servers are saturated<br />• Logging → **Slow Queries Drilldown** |
| **What you can do** | • Rewrite heavy queries or add indexes<br />• Set `timeoutMs` on the table or per query (broker default 10 s) |
| **What StarTree does** | Adds capacity when servers are saturated |
| **See also** | [Memory-throttled scheduler: troubleshooting](/corecapabilities/query_data/advanced_operations/workload-isolation/memory-throttled-scheduler#troubleshooting) |

### `QueryQuotaAlert`

<Tip>Usually your fix: the cause is most often in your stream, source system, or table config.</Tip>

| Field | Detail |
| - | - |
| **Fires when** | A table rejects queries over its QPS quota for 1 hour. |
| **What it means** | Pinot returns 429 for some queries. It's a quota, not an outage. |
| **Check first** | • Pinot - Broker → **Query Quota Exceeding**<br />• The quota is split across brokers and enforced per second, so short bursts can trip it |
| **What you can do** | • Raise `quota.maxQueriesPerSecond`<br />• Smooth out client bursts |

### `LowBrokerAvailability`

| Field | Detail |
| - | - |
| **Fires when** | Fewer than half the brokers are ready, for 5 min. |
| **What it means** | Query capacity is reduced. Brokers often fail from heavy queries. |
| **Check first** | Pinot - Broker → **Pod Restart Count**, **JVM Used Memory** |
| **What you can do** | • Use `DISTINCTCOUNTHLL` instead of exact distinct counts on large data<br />• Cap very large result sizes |
| **What StarTree does** | • Restores brokers<br />• Raises broker memory |

### `NoBrokerAvailable`

| Field | Detail |
| - | - |
| **Fires when** | No broker pods exist, for 5 min. |
| **What it means** | Every query fails. Ingestion continues. |
| **What you can do** | Nothing. StarTree handles this one. |
| **What StarTree does** | Restores brokers immediately |

### `BrokerDataTableDeserializationExceptions`

| Field | Detail |
| - | - |
| **Fires when** | Brokers fail to read server responses, for 5 min. |
| **What it means** | Some queries fail while merging results. |
| **What you can do** | Nothing. StarTree handles this one. |
| **What StarTree does** | Finds the server and version mismatch behind it |

### `NettyResponseFetchExceptions`

| Field | Detail |
| - | - |
| **Fires when** | Brokers fail to fetch responses from servers, for 15 min. |
| **What it means** | Some queries fail or return partial results. |
| **What you can do** | Nothing. StarTree handles this one. |
| **What StarTree does** | Checks server health and network |

## Warning alerts

| Alert | Fires when | What it means | What you can do |
| - | - | - | - |
| `QueriesFailingWithUnavailableSegment` | Over 1% of queries hit unavailable segments, for 15 min. | Results are partial. See Data availability. | Nothing needed |
| `BrokerQueryExecutionLatency` | Broker p95 query time is over 5 s, for 10 min. | Queries are slow across the broker. | Look at Logging → Slow Queries Drilldown |
| `BrokerResponseAlert` | Over 50 queries answered with only some servers in 5 min. | Those results are partial. | Nothing needed |
| `BrokerNoServerFoundAlert` | Over 50 queries found no server for a segment in 5 min. | Those queries miss data. | Nothing needed |

## Metrics

The series below are available in [Grafana](/corecapabilities/observability/grafana) through the Prometheus data source. For how names, suffixes, and labels are built, see [Reading Pinot metrics](/corecapabilities/observability/alerts/overview#reading-pinot-metrics).

### Metrics worth watching

These are the ones to reach for first; they are not the complete set.

| Metric | What it tells you |
| - | - |
| `pinot_broker_queries_Count` | Single-stage query volume. The denominator for most error-rate calculations. |
| `pinot_broker_multiStageQueries_Count`, `pinot_broker_multiStageQueriesGlobal_Count` | Multi-stage engine query volume, per table and cluster-wide. |
| `pinot_broker_queryTotalTimeMs_95thPercentile` | Broker-side P95 latency, per table. |
| `pinot_broker_queryErrorBrokerTimeout_OneMinuteRate` and siblings | Error rate split by cause — `ExecutionTimeout`, `QuerySchedulingTimeout`, `ServerNotResponding`, `ServerSegmentMissing`, `BrokerSegmentUnavailable`, `QueryPlanning`, `SqlRuntime`, `Internal`, `MergeResponse`, `QueryCancellation`, `ServerShuttingDown`, `BrokerRequestSend`. Splitting by cause is what turns "queries are failing" into a diagnosis. |
| `pinot_broker_queryCriticalError_OneMinuteRate` | Errors classed as critical. |
| `pinot_broker_brokerResponsesWithTimeouts_Count` | Responses that timed out. |
| `pinot_broker_brokerResponsesWithUnavailableSegments_Count` | Responses returned with segments missing — results may be incomplete. |
| `pinot_broker_brokerResponsesWithPartialServersResponded_Count` | Responses built without every server answering. |
| `pinot_broker_noServerFoundExceptions_Count` | No server could serve the table. Routing or availability, not query shape. |
| `pinot_broker_queryQuotaExceeded_Count` | Queries rejected by quota. |
| `pinot_server_schedulingTimeoutExceptions_Count` | Queries that timed out waiting to be scheduled rather than while executing — the distinction between "too busy" and "too slow". |
| `pinot_server_numMissingSegments_Count` | Segments a server was asked for and did not have. |

For per-query rather than aggregate analysis, use [`system_query_log`](/corecapabilities/query_data/advanced_operations/query-logger) — metrics tell you the shape of the problem, the query log tells you which query.

### Metrics behind each alert

The Prometheus series each alert rule evaluates.

| Alert | Metrics in the rule |
| - | - |
| `HighQueryCriticalErrorRate` | `pinot_broker_queryCriticalError_OneMinuteRate`, `pinot_broker_queries_Count`, `pinot_broker_multiStageQueriesGlobal_Count` |
| `HighSystemQueryErrorRate` | `pinot_broker_queryErrorSqlRuntime_OneMinuteRate`, `pinot_broker_queryErrorInternal_OneMinuteRate`, `pinot_broker_queryErrorQuerySchedulingTimeout_OneMinuteRate`, `pinot_broker_queryErrorExecutionTimeout_OneMinuteRate`, `pinot_broker_queryErrorBrokerTimeout_OneMinuteRate`, `pinot_broker_queryErrorServerSegmentMissing_OneMinuteRate`, `pinot_broker_queryErrorBrokerSegmentUnavailable_OneMinuteRate`, `pinot_broker_queryErrorServerNotResponding_OneMinuteRate`, `pinot_broker_queryErrorBrokerRequestSend_OneMinuteRate`, `pinot_broker_queryErrorMergeResponse_OneMinuteRate`, `pinot_broker_queryErrorQueryCancellation_OneMinuteRate`, `pinot_broker_queryErrorServerShuttingDown_OneMinuteRate`, `pinot_broker_queryErrorQueryPlanning_OneMinuteRate`, `pinot_broker_queries_Count`, `pinot_broker_multiStageQueriesGlobal_Count` |
| `HighQueryTimeoutPercentage` | `pinot_broker_brokerResponsesWithTimeouts_Count`, `pinot_broker_queries_Count` |
| `QueryQuotaAlert` | `pinot_broker_queryQuotaExceeded_Count` |
| `LowBrokerAvailability` | `kube_statefulset_status_replicas_ready`, `kube_statefulset_status_replicas` |
| `NoBrokerAvailable` | `kube_pod_info` |
| `BrokerDataTableDeserializationExceptions` | `pinot_broker_exceptions_dataTableDeserialization_Count` |
| `NettyResponseFetchExceptions` | `pinot_broker_exceptions_responseFetch_Count` |
| `QueriesFailingWithUnavailableSegment` | `pinot_broker_brokerResponsesWithUnavailableSegments_Count`, `pinot_broker_queries_Count`, `pinot_broker_multiStageQueries_Count` |
| `BrokerQueryExecutionLatency` | `pinot_broker_queryTotalTimeMs_95thPercentile` |
| `BrokerResponseAlert` | `pinot_broker_brokerResponsesWithPartialServersResponded_Count` |
| `BrokerNoServerFoundAlert` | `pinot_broker_noServerFoundExceptions_Count` |

## Related

* [Alerts and metrics overview](/corecapabilities/observability/alerts/overview): severity, the symptom router, the alert index, and how metric names are built.
* [Troubleshooting by feature: Queries, indexes and performance](/corecapabilities/observability/troubleshooting-by-feature#queries-indexes-and-performance)
* [Query troubleshooting](/corecapabilities/query_data/troubleshooting)
* [Error codes](/corecapabilities/observability/error-codes)


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.