How it works
- Server-side loading — each server decodes the Iceberg Puffin deletion files relevant to its segments and applies them when answering a query, so deleted/updated rows are excluded from results.
- Atomic snapshot readiness — before a new Iceberg snapshot becomes visible to queries, the controller polls every server until all of them confirm they’ve loaded the deletion vectors for that snapshot. This avoids a window where some servers see the new snapshot’s deletes and others don’t.
- Broker-side pinning and consistency — the broker pins every query to a single snapshot per table so all servers involved in a query agree on which snapshot (and therefore which deletes) to use. By default it pins the newest active snapshot automatically and injects the
snapshotVersionByTableoption for the servers; a query can also pin an explicit snapshot (see Query-time snapshot pinning). An invalid explicitly-pinned snapshot fails the query rather than silently mixing snapshot state.
Enabling deletion vectors
Set on the table’sExternalTableSyncTask config:
Controller: snapshot readiness
Each server exposes a readiness endpoint, which the controller polls:
{"ready": true|false}. Returns 412 if the table has not initialized a deletion-vector manager (i.e. enableDeletionVectors=false).
Broker: pruning and pinning
Broker-side snapshot pruning activates for OFFLINE External Tables when any of these is true on the table:
enableDeletionVectors=true, enableSnapshotConsistency=true (a flag the onboarding/sync flow sets automatically for snapshotting catalogs — pins queries to the last completed snapshot so they never see a half-ingested one), or segment groups enabled. It is a zero-cost no-op otherwise.
Query-time snapshot pinning
You do not need to pin a snapshot yourself: for a pruning-enabled table the broker automatically resolves each query to the newest active snapshot and injects the server-requiredsnapshotVersionByTable option. Set the option explicitly only to pin a query to a specific snapshot (time travel / reproducing an earlier state):
QUERY_EXECUTION error rather than returning results computed from a mix of snapshot states. A query also fails when the table has no active snapshots at all yet (nothing fully synced).
Storage format
The deletion-vector index itself is stored as a sorted, dictionary-encoded, bloom-filtered Parquet file (dv-index.parquet), which keeps heap usage low even for large delete sets and supports lazy point lookups.
Observability
The run-status endpoint’s
failurePhase field (see Observability) can also report IS_EV_CONVERGENCE or SNAPSHOT_READINESS_POLL when a sync run fails during deletion-vector convergence.
FAQs
Do I need to set snapshotVersionByTable on every query once deletion vectors are enabled?
No — the broker pins each query to the newest active snapshot automatically and injects the option for the servers. Set it explicitly only for time travel to a specific snapshot; an explicit pin that isn’t in the active set fails the query rather than returning potentially inconsistent results.
Does this work for raw S3 Parquet (non-Iceberg) External Tables?
No — deletion vectors are an Iceberg-specific capability (deletionVectorSource=iceberg), tied to Iceberg’s Puffin delete-file format.
