> ## Documentation Index
> Fetch the complete documentation index at: https://docs.startree.ai/llms.txt
> Use this file to discover all available pages before exploring further.

# Data availability and storage alerts

> Alerts that watch serving replicas, segments in error, and disk, key, and storage limits: when each fires, what it means, and what you or StarTree can do.

Whether every segment has enough serving replicas, and whether a table is close to a disk, key or segment limit.

<Info>
  **Critical** alerts page StarTree on-call 24/7. **Warning** alerts reach StarTree's alert channels without paging anyone. Thresholds are defaults; StarTree may tune them for your environment. See the [overview](/corecapabilities/observability/alerts/overview) for how to read an entry.
</Info>

## Decision tree

Follow the tree from the symptom to the alert most likely to explain it. Use the zoom controls at the top right of the diagram to enlarge it.

```mermaid placement="top-right" actions={true} theme={null}
flowchart TD
  A["Results incomplete or<br/>change between runs"] --> B{"Segments in ERROR?<br/>GET /tables/{t}/externalview"}
  B -- "Yes" --> B1["SegmentError<br/>POST /segments/{t}/reset?errorSegmentsOnly=true"]
  B -- "No" --> C{"Restart, rebalance or<br/>scale-out in progress?"}
  C -- "Yes" --> C1["Replicas are reloading<br/>SegmentReplicasCriticallyLowForHATable clears on its own"]
  C -- "No" --> D{"Servers restarting or OOM?"}
  D -- "Yes" --> D1["See Platform health"]
  D -- "No" --> E{"Ingestion or uploads blocked?"}
  E -- "Yes" --> E1["ResourceUtilizationLimitExceeded<br/>StorageQuotaUtilizationAlert"]
  E -- "No" --> F["Upsert table? See Upsert tables"]
```

## Critical alerts

### `SegmentReplicasCriticallyLowForHATable`

<Note>Often transient: this alert usually clears on its own.</Note>

| Field | Detail |
| - | - |
| **Fires when** | A table with replication 2+ has 34% or fewer replicas serving, for 10 min. |
| **What it means** | Queries may return partial results. Usually clears once replicas finish loading. |
| **Check first** | • Pinot Controllers → **Percent Of Replicas**, **Segments In Error State**<br />• Compare `GET /tables/{table}/externalview` with `GET /tables/{table}/idealstate` |
| **What you can do** | • Use replication 3 for critical tables<br />• Reduce minion task parallelism if it overlaps |
| **What StarTree does** | • Resets segments in ERROR<br />• Investigates servers that keep restarting |

### `ResourceUtilizationLimitExceeded`

| Field | Detail |
| - | - |
| **Fires when** | A table hits a server disk, primary-key or segment-count limit, for 10 min. |
| **What it means** | Real-time consumption pauses and new tasks aren't generated for that table. |
| **Check first** | • `GET /tables/{table}/pauseStatus` shows the reason<br />• Total Provisioned Capacity dashboard |
| **What you can do** | • Shorten retention, roll up or merge segments<br />• Set an upsert `metadataTTL` where supported |
| **What StarTree does** | • Adds disk or servers<br />• Consumption resumes on its own once below the limit |
| **See also** | [Segment count limits: monitoring and troubleshooting](/corecapabilities/manage-data/recipes/segment-count-limits#monitoring-and-troubleshooting) |

### `PinotZnodeSizeTooCloseToZkBufferLimit`

| Field | Detail |
| - | - |
| **Fires when** | A table's ideal state or external view is within 10 KB of ZooKeeper's size limit. |
| **What it means** | The table can stop accepting new segments. |
| **Check first** | Pinot Controllers → **Idealstate Znode Size** |
| **What you can do** | Reduce segment count with merge-rollup or retention |
| **What StarTree does** | Raises the ZooKeeper buffer if needed |
| **See also** | [Segment count limits: monitoring and troubleshooting](/corecapabilities/manage-data/recipes/segment-count-limits#monitoring-and-troubleshooting) |

### `TierStorageServerPreloadIndexReachingLimit`

| Field | Detail |
| - | - |
| **Fires when** | A server's tiered-storage preload cache is over 90% reserved, for 5 min. |
| **What it means** | When full, reads fall back to object storage: slower, still correct. |
| **What you can do** | Narrow the preloaded index keys |
| **What StarTree does** | Raises the preload budget |
| **See also** | [Tiered storage setup: monitoring metrics](/corecapabilities/manage-data/set-up-tiered-storage/setup) |

### `PinotPodVolumeCrossed90Percent`

| Field | Detail |
| - | - |
| **Fires when** | A Pinot volume is over 90% full, for 5 min. |
| **What it means** | Ingestion can stop and queries slow down. |
| **Check first** | Kubernetes / Storage / Volumes dashboard |
| **What you can do** | Shorten retention |
| **What StarTree does** | Grows the volume or adds servers |

### `SystemQueryLogStorageQuotaUtilization`

| Field | Detail |
| - | - |
| **Fires when** | The `system_query_log` table reaches 85% of its storage quota, for 10 min. |
| **What it means** | At 100%, new query-log segments are rejected until the quota rises. |
| **What you can do** | Nothing. StarTree handles this one. |
| **What StarTree does** | Expands the quota |

## Warning alerts

| Alert | Fires when | What it means | What you can do |
| - | - | - | - |
| `SegmentReplicasUnavailable` | 67% or fewer replicas serving, for 10 min. | Early warning before the critical alert. | Nothing needed |
| `SegmentError` | A table has segments in ERROR for 30 min. | Results can differ between runs. | Nothing needed |
| `ConfiguredRFLessThanThreshold` | A table has replication below 2 on a multi-server cluster, for 1 hour. | Any server restart takes part of the table offline. | Raise replication, or accept the risk |
| `StorageQuotaUtilizationAlert` | A table uses over 80% of its storage quota, for 10 min. | Uploads are rejected at 100%. | Raise `quota.storage` or shorten retention |
| `PinotPodVolumeCrossed80Percent` | A Pinot volume is over 80% full, for 5 min. | Early warning before 90%. | Nothing needed |

## Metrics

The series below are available in [Grafana](/corecapabilities/observability/grafana) through the Prometheus data source. For how names, suffixes, and labels are built, see [Reading Pinot metrics](/corecapabilities/observability/alerts/overview#reading-pinot-metrics).

### Metrics worth watching

These are the ones to reach for first; they are not the complete set.

| Metric | What it tells you |
| - | - |
| `pinot_controller_percentOfReplicas_Value` | Share of segment replicas available, per table. Non-lead controllers emit a negative placeholder, so filter `>= 0`. |
| `pinot_controller_replicationFromConfig_Value` | Configured replication factor. |
| `pinot_controller_segmentsInErrorState_Value` | Segments stuck in error. |
| `pinot_controller_tableStorageQuotaUtilization_Value` | Percentage of the table's storage quota in use. |
| `pinot_controller_idealstateZnodeByteSize_Value`, `pinot_controller_externalviewZnodeByteSize_Value`, `pinot_controller_propertystoreSegmentChildrenByteSize_Value` | ZooKeeper znode sizes — the constraint that eventually bounds segment count. |
| `pinot_server_preloadCacheLayerReservedSizeBytes_Value`, `pinot_server_preloadCacheLayerMaxSizeBytes_Value` | Tiered-storage preload index headroom. Chart the ratio. |

### Metrics behind each alert

The Prometheus series each alert rule evaluates.

| Alert | Metrics in the rule |
| - | - |
| `SegmentReplicasCriticallyLowForHATable` | `pinot_broker_percentOfReplicas_Value`, `pinot_controller_percentOfReplicas_Value`, `pinot_controller_replicationFromConfig_Value`, `pinot_controller_numberOfReplicas_Value` |
| `ResourceUtilizationLimitExceeded` | `pinot_controller_resourceUtilizationLimitExceeded_Value` |
| `PinotZnodeSizeTooCloseToZkBufferLimit` | `pinot_controller_idealstateZnodeByteSize_Value`, `pinot_controller_externalviewZnodeByteSize_Value`, `pinot_controller_propertystoreSegmentChildrenByteSize_Value` |
| `TierStorageServerPreloadIndexReachingLimit` | `pinot_server_preloadCacheLayerReservedSizeBytes_Value`, `pinot_server_preloadCacheLayerMaxSizeBytes_Value` |
| `PinotPodVolumeCrossed90Percent` | `kubelet_volume_stats_used_bytes`, `kubelet_volume_stats_capacity_bytes` |
| `SystemQueryLogStorageQuotaUtilization` | `pinot_controller_tableStorageQuotaUtilization_Value` |
| `SegmentReplicasUnavailable` | `pinot_controller_percentOfReplicas_Value` |
| `SegmentError` | `pinot_controller_segmentsInErrorState_Value` |
| `ConfiguredRFLessThanThreshold` | `pinot_controller_replicationFromConfig_Value`, `kube_statefulset_status_replicas` |
| `StorageQuotaUtilizationAlert` | `pinot_controller_tableStorageQuotaUtilization_Value` |
| `PinotPodVolumeCrossed80Percent` | `kubelet_volume_stats_used_bytes`, `kubelet_volume_stats_capacity_bytes` |

## Related

* [Alerts and metrics overview](/corecapabilities/observability/alerts/overview): severity, the symptom router, the alert index, and how metric names are built.
* [Troubleshooting by feature: Segments and storage](/corecapabilities/observability/troubleshooting-by-feature#segments-and-storage)
* [Health checks](/corecapabilities/observability/health-checks)
* [Table health and email alerts](/corecapabilities/observability/table-health-and-alerting)


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.