Looking for the API-based approach? See Apache Polaris & Snowflake Open Catalog: Onboarding via API. Using Snowflake Horizon Catalog (Snowflake-managed Iceberg tables, PAT auth) instead? See Snowflake Horizon: Onboarding via Data Portal.
Prerequisites
Before starting, ensure you have:- StarTree 0.16.0 or later with the External Table feature enabled and tiered storage configured for your environment. Contact StarTree support if unsure.
- A running catalog:
- Apache Polaris — a reachable Polaris server and a principal (client ID + client secret) whose principal role can read the catalog.
- Snowflake Open Catalog — an Open Catalog account and a service connection’s principal credentials.
- The catalog name (Polaris/Open Catalog catalog backing the Iceberg tables) and the namespace/table you want to onboard.
- Read access to the underlying S3 object storage — static keys or an IAM role StarTree can assume. See Onboarding via API — Authentication for the permissions required.
Step 1: Open the External Tables
- Log in to Data Portal.
- In the left navigation, go to Tables.
- Click + Connect External Table.
Step 2: Configure the Catalog Connection
Choose Apache Polaris or Snowflake Open Catalog as the catalog type — they appear as top-level tiles alongside AWS Glue, Amazon S3 Tables, and Unity Catalog. Fill in the connection details: Details
Metastore — how StarTree authenticates with the Iceberg REST API.
Storage — credentials for reading the underlying Parquet data files from S3.
Click Validate Connection. Data Portal calls the catalog’s validate endpoint and confirms credentials and connectivity before proceeding.
Step 3: Browse and Select a Table
Once the connection is validated:- Data Portal lists the available namespaces — the field is labeled Namespace (Apache Polaris) or Schema (Snowflake Open Catalog).
- Select one to expand its tables.
- Click the table you want to onboard.
Step 4: Review the Schema
The inferred Pinot schema is displayed for review. For External Tables this step is mostly read-only — columns can’t be added, removed, or renamed, and the inferred data types should be kept (overriding them can break segment generation). One thing is editable inline: the Pinot Field Type column lets you switch a column between Dimension, Metric, and Date-Time where its data type allows (for example, aLONG can be any of the three; a TIMESTAMP can only be Date-Time). Selecting Date-Time opens a dialog to set the value format and granularity.
Multi-value columns are flagged in a dedicated column. Click Next when the schema looks correct.
Step 5: Configure the Table
Set the final table options:
The sync schedule is fixed at every 5 minutes at creation (change it afterwards via the table config, or use Pause Sync / Resume Sync), and null handling is enabled automatically on the generated table config.
Click Create Table to register the schema and table with Pinot. Data Portal automatically triggers the first onboarding run immediately after creation — no manual step required.

