> ## Documentation Index
> Fetch the complete documentation index at: https://docs.startree.ai/llms.txt
> Use this file to discover all available pages before exploring further.

# Platform health alerts

> Alerts that watch Pinot components, ZooKeeper, Kubernetes, platform services, certificates, and monitoring, most of which StarTree handles end to end.

This page covers the Pinot components themselves, then the infrastructure underneath them. Most of these alerts are handled by StarTree end to end.

<Info>
  **Critical** alerts page StarTree on-call 24/7. **Warning** alerts reach StarTree's alert channels without paging anyone. Thresholds are defaults; StarTree may tune them for your environment. See the [overview](/corecapabilities/observability/alerts/overview) for how to read an entry.
</Info>

## Cluster health

Pinot controllers, brokers, servers, minions and ZooKeeper: restarts, memory, CPU and coordination.

### Decision tree

Follow the tree from the symptom to the alert most likely to explain it. Use the zoom controls at the top right of the diagram to enlarge it.

```mermaid placement="top-right" actions={true} theme={null}
flowchart TD
  A["Components restart or run hot"] --> B{"Last termination reason?"}
  B -- "OOMKilled" --> B1["KubernetesPinotContainerOOMKilled<br/>heavy queries or task waves"]
  B -- "Crash on start" --> B2["PinotPodWaiting<br/>config, keystore or deep store access"]
  B -- "Liveness kill" --> B3["PinotHighGcPercent<br/>long GC pauses"]
  A --> C{"CPU above 85% of request<br/>for 30 min?"}
  C -- "With query errors" --> C1["Capacity problem<br/>scale or limit queries"]
  C -- "Alone" --> C2["Often an ingestion or<br/>reload peak"]
  A --> D{"Coordination?"}
  D -- "No lead controller" --> D1["NoControllerLeaderAlert"]
  D -- "ZooKeeper quorum or leader" --> D2["Zookeeper* alerts"]
```

### Critical alerts

#### `NoControllerLeaderAlert`

| Field | Detail |
| - | - |
| **Fires when** | No controller is leader, for 10 min. |
| **What it means** | Table changes, segment commits and tasks stall. |
| **Check first** | Pinot Controllers → **Pinot Controller Leader** |
| **What you can do** | Nothing. StarTree handles this one. |
| **What StarTree does** | Restores leadership |

#### `PinotControllerHighThreadsBlocked`

| Field | Detail |
| - | - |
| **Fires when** | Blocked controller threads exceed half the runnable ones, for 10 min. |
| **What it means** | The controller is contended; operations slow down. |
| **What you can do** | Nothing. StarTree handles this one. |
| **What StarTree does** | Takes thread dumps and restarts |

#### `PinotHighGcPercent`

| Field | Detail |
| - | - |
| **Fires when** | A Pinot or ZooKeeper JVM spends over 50% of time in GC, for 15 min. |
| **What it means** | Memory pressure; queries and ingestion slow down. |
| **Check first** | JVM dashboard |
| **What you can do** | Fix queries that build large results |
| **What StarTree does** | Adds memory or finds the leak |

#### `PinotComponentHighCPUUsage`

<Note>Often transient: this alert usually clears on its own.</Note>

| Field | Detail |
| - | - |
| **Fires when** | A server, broker or controller uses over 85% of its CPU request for 30 min. Servers still starting are skipped. |
| **What it means** | Sustained load. Paired with query errors, it's a capacity problem. |
| **Check first** | Pinot Servers → **CPU % Usage** |
| **What you can do** | Move heavy reloads or backfills off peak hours |
| **What StarTree does** | Scales or rebalances |

#### `HighServerPodsRestarts`

| Field | Detail |
| - | - |
| **Fires when** | A server restarts more than twice in 30 min. |
| **What it means** | Replicas drop and ingestion can stop. |
| **Check first** | Logging / Kubernetes Events |
| **What you can do** | Nothing. StarTree handles this one. |
| **What StarTree does** | Finds the OOM or crash cause |

#### `HighNonServerPodsRestarts`

| Field | Detail |
| - | - |
| **Fires when** | A broker, controller or ZooKeeper pod restarts more than twice in 15 min. |
| **What it means** | Queries or coordination are disrupted. |
| **Check first** | Logging / Kubernetes Events |
| **What you can do** | Nothing. StarTree handles this one. |
| **What StarTree does** | Finds the OOM or crash cause |

#### `KubernetesPinotContainerOOMKilled`

| Field | Detail |
| - | - |
| **Fires when** | A Pinot container was killed for running out of memory. |
| **What it means** | Often heavy queries, task waves or on-heap upsert metadata. |
| **Check first** | Logging / Kubernetes Events |
| **What you can do** | Avoid exact distinct counts and very large LIMITs |
| **What StarTree does** | Raises memory or moves upsert metadata to RocksDB |

#### `KubernetesPinotPodCrashLooping`

| Field | Detail |
| - | - |
| **Fires when** | A Pinot pod restarts more than 5 times in 10 min. |
| **What it means** | The pod can't stay up. |
| **What you can do** | Nothing. StarTree handles this one. |
| **What StarTree does** | Fixes the crash cause |

#### `PinotPodWaiting`

| Field | Detail |
| - | - |
| **Fires when** | A Pinot pod is in CrashLoopBackOff for 15 min. |
| **What it means** | Usually bad config, keystore or deep store access. |
| **What you can do** | Check any credentials you supplied for deep store |
| **What StarTree does** | Fixes startup |

#### `PinotPodStuckInContainerCreating`

| Field | Detail |
| - | - |
| **Fires when** | A Pinot pod is stuck creating its container for 15 min. |
| **What it means** | Usually a volume or node problem. |
| **What you can do** | Nothing. StarTree handles this one. |
| **What StarTree does** | Fixes scheduling |

#### `PinotPodImagePullErrors`

| Field | Detail |
| - | - |
| **Fires when** | A Pinot pod can't pull its image for 15 min. |
| **What it means** | The pod can't start. |
| **What you can do** | Nothing. StarTree handles this one. |
| **What StarTree does** | Fixes the image reference or registry |

#### `HighAvailablePinotStatefulsetReplicasMismatch`

| Field | Detail |
| - | - |
| **Fires when** | A Pinot component with 2+ replicas has fewer than 2 ready, for 1 hour. |
| **What it means** | High availability is lost. |
| **What you can do** | Nothing. StarTree handles this one. |
| **What StarTree does** | Restores replicas |

#### `ZookeeperErrorCount`

| Field | Detail |
| - | - |
| **Fires when** | ZooKeeper reports an unrecoverable error. |
| **What it means** | A ZooKeeper thread died. |
| **Check first** | Pinot ZooKeeper dashboard |
| **What you can do** | Nothing. StarTree handles this one. |
| **What StarTree does** | Checks quorum and restarts the member |

#### `ZookeeperLeaderElectionTooLong`

| Field | Detail |
| - | - |
| **Fires when** | ZooKeeper keeps electing a leader, for 15 min. |
| **What it means** | Coordination is unstable. |
| **What you can do** | Nothing. StarTree handles this one. |
| **What StarTree does** | Restores quorum |

#### `ZookeeperIncorrectNumberOfLeaders`

| Field | Detail |
| - | - |
| **Fires when** | ZooKeeper doesn't have exactly one leader, for 10 min. |
| **What it means** | Coordination is unstable. |
| **What you can do** | Nothing. StarTree handles this one. |
| **What StarTree does** | Restores quorum |

#### `ZookeeperIncorrectQuorumSize`

| Field | Detail |
| - | - |
| **Fires when** | ZooKeeper quorum drops below 3, for 10 min. |
| **What it means** | Fault tolerance is reduced. |
| **What you can do** | Nothing. StarTree handles this one. |
| **What StarTree does** | Restores members |

#### `ZookeeperMetricsMissing`

| Field | Detail |
| - | - |
| **Fires when** | ZooKeeper leader metrics are missing for 20 min. |
| **What it means** | A monitoring gap, not necessarily an outage. |
| **What you can do** | Nothing. StarTree handles this one. |
| **What StarTree does** | Restores metrics |

### Warning alerts

| Alert | Fires when | What it means | What you can do |
| - | - | - | - |
| `ZookeeperLeaderUnavailable` | ZooKeeper reports time without a leader, for 10 min. | Brief coordination gaps. | Nothing needed |
| `PinotStatefulsetReplicasMismatch` | A Pinot component has fewer ready replicas than desired, for 1 hour. | Capacity is reduced. | Nothing needed |
| `ContainerOOMKills` | A container was OOM-killed in the last 24 hours. | Repeated memory pressure. | Nothing needed |

## Infrastructure StarTree manages

StarTree handles these alerts end to end. They're listed so you know what is watched; none of them needs action from you.

<AccordionGroup>
  <Accordion title="Kubernetes nodes and network (7 critical, 5 warning)">
    | Alert | Severity | Fires when |
    | - | - | - |
    | `NodeDataDiskUsageCritical` | Critical | Node data disk over 90%, 5 min |
    | `NodeConntrackExhausted` | Critical | Node connection tracking exhausted; new connections drop |
    | `KubernetesNodeNotReady` | Critical | A node not Ready, 5 min |
    | `KubernetesNodeMemoryPressure` | Critical | Node memory pressure, 15 min |
    | `KubernetesNodeDiskPressure` | Critical | Node disk pressure, 15 min |
    | `KubernetesNodeNetworkUnavailable` | Critical | Node network unavailable, 5 min |
    | `NodeConntrackLow` | Critical | Connection tracking allowance running low, 5 min |
    | `NodeDataDiskUsageHigh` | Warning | Node data disk over 80%, 15 min |
    | `KubernetesNodeOutOfPodCapacity` | Warning | Node over 90% of pod capacity, 15 min |
    | `NodeNetworkBandwidthInThrottled` | Warning | Inbound bandwidth throttled, 5 min |
    | `NodeNetworkBandwidthOutThrottled` | Warning | Outbound bandwidth throttled, 5 min |
    | `NodeNetworkPPSThrottled` | Warning | Packets-per-second throttled, 5 min |
  </Accordion>

  <Accordion title="Kubernetes pods and volumes (6 critical, 5 warning)">
    | Alert | Severity | Fires when |
    | - | - | - |
    | `KubernetesPodPending` | Critical | A pod pending, 10 min |
    | `KubernetesPodFailed` | Critical | Over 50 failed instances of a pod, 10 min |
    | `KubernetesMinionPodNotHealthy` | Critical | A minion pod not running, 25 min |
    | `KubernetesPodCrashLooping` | Critical | A platform pod restarts 5+ times in 10 min |
    | `KubernetesContainerOOMKilled` | Critical | A platform container OOM-killed |
    | `PodVolumeCrossed90Percent` | Critical | A platform volume over 90%, 5 min |
    | `PodVolumeCrossed80Percent` | Warning | A platform volume over 80%, 5 min |
    | `KubernetesPodNotHealthy` | Warning | A pod Pending, Unknown or Failed, 10 min |
    | `KubernetesContainerTerminated` | Warning | A container terminated with an error |
    | `KubernetesReplicasetReplicasMismatch` | Warning | ReplicaSet not fully ready, 10 min |
    | `KubernetesStatefulsetUpdateNotRolledOut` | Warning | StatefulSet update not rolled out, 10 min |
  </Accordion>

  <Accordion title="Platform services (16 critical, 5 warning)">
    | Alert | Severity | Fires when |
    | - | - | - |
    | `PlatformAuthP95Latency` | Critical | Sign-in p95 over 100 ms, 15 min |
    | `PlatformAuthZP95Latency` | Critical | Authorization p95 over 100 ms, 15 min |
    | `PlatformAuthInfoP95Latency` | Critical | Auth info API p95 over 100 ms, 15 min |
    | `PlatformAuthAPIP95Latency` | Critical | Other auth APIs p95 over 200 ms, 15 min |
    | `AuthServiceHighHttp5xxErrorRate` | Critical | Auth service 5xx over 10%, 15 min |
    | `RBACServiceHighHttp5xxErrorRate` | Critical | RBAC service 5xx over 10%, 15 min |
    | `PlatformStatefulsetReplicasMismatch` | Critical | Platform StatefulSet not fully ready, 15 min |
    | `PlatformPodWaiting` | Critical | Platform pod stuck waiting, 15 min |
    | `MariadbDown` | Critical | Metadata database down, 5 min |
    | `MariadbHAReplicationIOThreadStopped` | Critical | Database replication I/O thread stopped |
    | `MariadbHAReplicationSQLThreadStopped` | Critical | Database replication SQL thread stopped |
    | `MariadbHANotAvailable` | Critical | Database cluster below 2 nodes, 5 min |
    | `MariadbHAReplicationLagHigh` | Critical | Database replication lag high, 10 min |
    | `MariadbHAReplicationRecvQueueSize` | Critical | Database replication write latency high, 10 min |
    | `MariadbHAReplicationSendQueueSize` | Critical | Database replication send latency high, 10 min |
    | `MariadbConnectionsExhausted` | Critical | Database connections nearly exhausted, 5 min |
    | `PlatformDeploymentReplicasMismatch` | Warning | Platform deployment not fully ready, 15 min |
    | `PlatformFrontendErrorLogs` | Warning | Rising error logs in the frontend |
    | `PlatformAuthErrorLogs` | Warning | Rising error logs in auth |
    | `OperatorDisabled` | Warning | Operator disabled, 48 h |
    | `OperatorOffline` | Warning | Operator offline, 48 h |
  </Accordion>

  <Accordion title="API gateway (Traefik) (5 critical, 2 warning)">
    | Alert | Severity | Fires when |
    | - | - | - |
    | `TraefikDeploymentReplicasMismatch` | Critical | Gateway replicas not all available, 5 min |
    | `TraefikServiceConnectionsAbnormalGrowth` | Critical | Open connections jump 4× above the 5-min average and over 500 |
    | `TraefikHighHttp499PinotErrorRateService` | Critical | Over 20% of Pinot requests closed by the client (499), 5 min |
    | `TraefikPinotBrokerHttp5xxBurnRateHigh` | Critical | Broker 5xx burning the 30-day error budget 6× too fast |
    | `TraefikPinotBrokerHttp403BurnRateHigh` | Critical | Broker 403s burning the 30-day error budget 6× too fast |
    | `TraefikHighHttp4xxErrorRateService` | Warning | Over 5% 4xx on a service, 1 min |
    | `TraefikHighHttp5xxErrorRateService` | Warning | Over 5% 5xx on a service, 1 min |
  </Accordion>

  <Accordion title="Certificates (3 critical, 2 warning)">
    | Alert | Severity | Fires when |
    | - | - | - |
    | `CertificateExpiration` | Critical | A platform certificate expires in under 4 days |
    | `PinotCertificateExpiration` | Critical | A Pinot endpoint certificate expires in under 4 days |
    | `X509ExporterReadErrors` | Critical | Certificate exporter can't read certificates, 15 min |
    | `CertificateRenewal` | Warning | A platform certificate expires in under 7 days |
    | `PinotCertificateRenewal` | Warning | A Pinot certificate expires in under 7 days |
  </Accordion>

  <Accordion title="Monitoring stack (11 critical, 1 warning)">
    | Alert | Severity | Fires when |
    | - | - | - |
    | `PrometheusPodVolumeCrossed75Percent` | Critical | Prometheus volume over 75%, 5 min |
    | `PrometheusContainerHighMemoryUsage` | Critical | Prometheus over 80% of its memory limit, 10 min |
    | `PrometheusContainerOOMKilled` | Critical | Prometheus OOM-killed |
    | `TSDBDataInsufficient` | Critical | Less than 15 days of metrics retained |
    | `PrometheusTsdbCompactionsFailed` | Critical | Metrics storage compaction failed in 2 h |
    | `PrometheusTsdbWalCorruptions` | Critical | Write-ahead log corrupted in 2 h |
    | `PrometheusTsdbHeadTruncationsFailed` | Critical | Head truncation failed in 2 h |
    | `PrometheusTsdbWalTruncationsFailed` | Critical | WAL truncation failed in 2 h |
    | `PrometheusTsdbReloadFailures` | Critical | Storage reload failed in 2 h |
    | `PrometheusTsdbCheckpointCreationFailures` | Critical | Checkpoint creation failed in 2 h |
    | `PrometheusTsdbCheckpointDeletionFailures` | Critical | Checkpoint deletion failed in 2 h |
    | `PrometheusPodVolumeCrossed65Percent` | Warning | Prometheus volume over 65%, 5 min |
  </Accordion>

  <Accordion title="ThirdEye (9 critical, 3 warning)">
    | Alert | Severity | Fires when |
    | - | - | - |
    | `DetectionTaskHealth` | Critical | Over 20% of detection tasks fail over 2 h |
    | `ThirdeyeNotificationFailure` | Critical | A notification failed to send in 15 min |
    | `ThirdeyePodWaiting` | Critical | A ThirdEye pod stuck waiting, 15 min |
    | `MySQLStorageExhaustion` | Critical | ThirdEye database volume over 90%, 5 min |
    | `ThirdeyeCertificateExpiration` | Critical | A ThirdEye certificate expires in under 4 days |
    | `ThirdeyeMultipleNamespaceConfigurations` | Critical | Conflicting namespace configuration in the database |
    | `ThirdeyePodVolumeCrossed90Percent` | Critical | A ThirdEye volume over 90%, 5 min |
    | `KubernetesThirdeyeContainerOOMKilled` | Critical | A ThirdEye container OOM-killed |
    | `KubernetesThirdeyePodCrashLooping` | Critical | A ThirdEye pod restarts 5+ times in 10 min |
    | `MySQLStorageUsage` | Warning | ThirdEye database volume over 80%, 5 min |
    | `ThirdeyeCertificateRenewal` | Warning | A ThirdEye certificate expires in under 7 days |
    | `ThirdeyePodVolumeCrossed80Percent` | Warning | A ThirdEye volume over 80%, 5 min |
  </Accordion>
</AccordionGroup>

## Metrics

The series below are available in [Grafana](/corecapabilities/observability/grafana) through the Prometheus data source. For how names, suffixes, and labels are built, see [Reading Pinot metrics](/corecapabilities/observability/alerts/overview#reading-pinot-metrics).

### Metrics worth watching

These are the ones to reach for first; they are not the complete set.

| Metric | What it tells you |
| - | - |
| `pinot_controller_pinotControllerLeader_Value` | `1` on the leader. Summing to zero across the cluster means no leader. |
| `jvm_gc_collection_seconds_sum` | GC time, by `component` and `gc`. Rate of this over wall-clock time is the GC percentage. |
| `jvm_threads_state` | Thread counts by `state` and `component`. Blocked-versus-runnable is how a wedged controller shows up. |

### Metrics behind each alert

The Prometheus series each alert rule evaluates, for the Pinot component alerts.

| Alert | Metrics in the rule |
| - | - |
| `NoControllerLeaderAlert` | `pinot_controller_pinotControllerLeader_Value` |
| `PinotControllerHighThreadsBlocked` | `jvm_threads_state` |
| `PinotHighGcPercent` | `jvm_gc_collection_seconds_sum` |
| `PinotComponentHighCPUUsage` | `container_cpu_user_seconds_total`, `container_cpu_system_seconds_total`, `kube_pod_container_resource_requests`, `pinot_server_startupStatusCheckInProgress_Value`, `kube_pod_status_phase` |
| `HighServerPodsRestarts` | `kube_pod_container_status_restarts_total` |
| `HighNonServerPodsRestarts` | `kube_pod_container_status_restarts_total` |
| `KubernetesPinotContainerOOMKilled` | `kube_pod_container_status_restarts_total`, `kube_pod_container_status_last_terminated_reason` |
| `KubernetesPinotPodCrashLooping` | `kube_pod_container_status_restarts_total` |
| `PinotPodWaiting` | `kube_pod_container_status_waiting_reason` |
| `PinotPodStuckInContainerCreating` | `kube_pod_container_status_waiting_reason` |
| `PinotPodImagePullErrors` | `kube_pod_container_status_waiting_reason` |
| `HighAvailablePinotStatefulsetReplicasMismatch` | `kube_statefulset_status_replicas_ready`, `kube_statefulset_status_replicas` |
| `ZookeeperErrorCount` | `unrecoverable_error_count` |
| `ZookeeperLeaderElectionTooLong` | `election_time_count` |
| `ZookeeperIncorrectNumberOfLeaders` | `leader_uptime` |
| `ZookeeperIncorrectQuorumSize` | `quorum_size` |
| `ZookeeperMetricsMissing` | `leader_unavailable_time_sum` |
| `ZookeeperLeaderUnavailable` | `leader_unavailable_time_sum` |
| `PinotStatefulsetReplicasMismatch` | `kube_statefulset_status_replicas_ready`, `kube_statefulset_status_replicas` |
| `ContainerOOMKills` | `kube_pod_container_status_restarts_total`, `kube_pod_container_status_last_terminated_reason` |

## Related

* [Alerts and metrics overview](/corecapabilities/observability/alerts/overview): severity, the symptom router, the alert index, and how metric names are built.
* [Cluster operations troubleshooting](/corecapabilities/cluster-operations/troubleshooting)
* [Grafana dashboards](/corecapabilities/observability/grafana)


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.