> ## Documentation Index
> Fetch the complete documentation index at: https://docs.startree.ai/llms.txt
> Use this file to discover all available pages before exploring further.

# Upsert table alerts

> Alerts that watch upsert and dedup tables: primary-key growth, compaction, snapshots, RocksDB, and replica consistency.

Primary-key growth, compaction and snapshots for upsert and dedup tables.

<Info>
  **Critical** alerts page StarTree on-call 24/7. **Warning** alerts reach StarTree's alert channels without paging anyone. Thresholds are defaults; StarTree may tune them for your environment. See the [overview](/corecapabilities/observability/alerts/overview) for how to read an entry.
</Info>

## Decision tree

Follow the tree from the symptom to the alert most likely to explain it. Use the zoom controls at the top right of the diagram to enlarge it.

```mermaid placement="top-right" actions={true} theme={null}
flowchart TD
  A["Upsert table looks wrong"] --> B{"Same query, different<br/>results per replica?"}
  B -- "Yes" --> B1["UpsertInconsistentReplicas<br/>enforceConsumptionInOrder"]
  B -- "No" --> C{"Storage growing<br/>faster than keys?"}
  C -- "Yes" --> C1["HighTotalDocsToPKCountRatio<br/>compaction is behind"]
  C -- "No" --> D{"Key count near the<br/>per-server limit?"}
  D -- "Yes" --> D1["HighTotalNumberOfPrimaryKeysPerServer<br/>more servers, or metadataTTL where supported"]
  D -- "No" --> E["SnapshotCreationFailed<br/>slow restarts, no query impact"]
```

## Critical alerts

### `HighTotalNumberOfPrimaryKeysPerServer`

| Field | Detail |
| - | - |
| **Fires when** | A server holds over 2 billion primary keys, for 30 min. |
| **What it means** | Early warning before the key limit pauses ingestion. |
| **Check first** | Pinot Upsert → **Upsert Table Primary Key Count (per server)** |
| **What you can do** | Set `metadataTTL` on the upsert config, except on tables that run `SegmentRefreshTask`. Ask StarTree support before changing it on a live table |
| **What StarTree does** | Spreads partitions across more servers |
| **See also** | [Upsert operations guide: supported configurations](/corecapabilities/manage-data/upsert-operations-guide#supported-configurations-at-a-glance) |

### `HighTotalDocsToPKCountRatio`

| Field | Detail |
| - | - |
| **Fires when** | An upsert table holds over 10 documents per primary key, for 1 hour. |
| **What it means** | Compaction is behind, so storage and cost grow. |
| **Check first** | A daily sawtooth is normal. A floor that keeps rising means compaction is failing |
| **What you can do** | Run compaction more often |
| **What StarTree does** | Fixes failing compaction tasks |
| **See also** | [Upsert operations guide: troubleshooting](/corecapabilities/manage-data/upsert-operations-guide#troubleshooting) |

### `SnapshotCreationFailed`

| Field | Detail |
| - | - |
| **Fires when** | A snapshot task failed in the last hour. |
| **What it means** | No query impact. Restarts take longer. |
| **Check first** | Pinot Upsert → **Snapshot Creation Failure Rate** |
| **What you can do** | Check deep store permissions and the partition config |
| **What StarTree does** | Fixes the task |
| **See also** | [Upsert operations guide: troubleshooting](/corecapabilities/manage-data/upsert-operations-guide#troubleshooting) |

### `RocksDBHighMemoryUsage`

| Field | Detail |
| - | - |
| **Fires when** | RocksDB memtables on a server exceed 10 GB. |
| **What it means** | Upsert metadata is using too much memory. |
| **What you can do** | Nothing. StarTree handles this one. |
| **What StarTree does** | Tunes RocksDB or adds memory |

### `RocksDBHighFlushRate`

| Field | Detail |
| - | - |
| **Fires when** | RocksDB flushes more than 500 times per 2 min on a server. |
| **What it means** | Heavy write churn on upsert metadata. |
| **What you can do** | Nothing. StarTree handles this one. |
| **What StarTree does** | Tunes RocksDB |

### `RocksDBHighCompactionRate`

| Field | Detail |
| - | - |
| **Fires when** | RocksDB compacts more than 10 times per 2 min on a server. |
| **What it means** | Heavy write churn on upsert metadata. |
| **What you can do** | Nothing. StarTree handles this one. |
| **What StarTree does** | Tunes RocksDB |

## Warning alerts

| Alert | Fires when | What it means | What you can do |
| - | - | - | - |
| `UpsertInconsistentReplicas` | Keys weren't replaced while replacing segments, for 2 min. | Replicas can disagree; partial upserts can show defaults. | Set `enforceConsumptionInOrder` and `parallelSegmentConsumptionPolicy: DISALLOW_ALWAYS` |

## Metrics

The series below are available in [Grafana](/corecapabilities/observability/grafana) through the Prometheus data source. For how names, suffixes, and labels are built, see [Reading Pinot metrics](/corecapabilities/observability/alerts/overview#reading-pinot-metrics).

### Metrics worth watching

These are the ones to reach for first; they are not the complete set.

| Metric | What it tells you |
| - | - |
| `pinot_server_upsertPrimaryKeysCount_Value` | Primary keys held in memory. Grows with distinct keys, not rows — the number that outgrows a cluster silently. |
| `pinot_server_documentCount_Value` | Documents per table. Its ratio to primary-key count tells you whether compaction is keeping up. |
| `pinot_server_realtimeUpsertInconsistentRows_Count` | Rows that could not be reconciled between replicas. |
| `pinot_server_partialUpsertKeysNotReplaced_Count` | Partial-upsert keys not replaced during segment replacement. |
| `pinot_server_rocksdb_ticker_total_memtable_bytes` | Off-heap upsert metadata memory. |
| `pinot_server_rocksdb_histogram_flush_time_count`, `pinot_server_rocksdb_histogram_compaction_time_count` | Off-heap upsert write amplification. |

### Metrics behind each alert

The Prometheus series each alert rule evaluates.

| Alert | Metrics in the rule |
| - | - |
| `HighTotalNumberOfPrimaryKeysPerServer` | `pinot_server_upsertPrimaryKeysCount_Value` |
| `HighTotalDocsToPKCountRatio` | `pinot_server_documentCount_Value`, `pinot_server_upsertPrimaryKeysCount_Value`, `pinot_controller_cronSchedulerJobScheduled_Value` |
| `SnapshotCreationFailed` | `pinot_minion_snapshotCreationFailure_Count` |
| `RocksDBHighMemoryUsage` | `pinot_server_rocksdb_ticker_total_memtable_bytes` |
| `RocksDBHighFlushRate` | `pinot_server_rocksdb_histogram_flush_time_count` |
| `RocksDBHighCompactionRate` | `pinot_server_rocksdb_histogram_compaction_time_count` |
| `UpsertInconsistentReplicas` | `pinot_server_realtimeUpsertInconsistentRows_Count`, `pinot_server_partialUpsertKeysNotReplaced_Count` |

## Related

* [Alerts and metrics overview](/corecapabilities/observability/alerts/overview): severity, the symptom router, the alert index, and how metric names are built.
* [Upsert operations guide](/corecapabilities/manage-data/upsert-operations-guide)
* [Diagnosing duplicates and count mismatches](/corecapabilities/manage-data/offheap-upsert#diagnosing-duplicates-and-count-mismatches)
* [Troubleshooting by feature: Upserts and dedup](/corecapabilities/observability/troubleshooting-by-feature#upserts-and-dedup)


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.