Databricks Collected Data

This page lists what Sifflet imports from Databricks once you connect your workspace.

Imported Assets

Sifflet imports the following objects from the Unity Catalog schemas selected in the source scope:

Databricks objectAsset type in Sifflet
Managed tablesTable
External and foreign tablesExternal table
Views and materialized viewsView
Streaming tablesStreaming table
Volumes (managed and external)Volume

Sifflet skips the internal tables that Lakeflow Declarative Pipelines (formerly Delta Live Tables) create, such as event_log_<pipeline_id> and __materialization_mat_<pipeline_id> tables.

To import Databricks jobs, add the Databricks Workflows integration.

Collected Metadata

Sifflet reads metadata from the information_schema of each catalog in scope.

Asset typeMetadata
Tables, external tables, views, streaming tablesName, type, comment (as description), Unity Catalog tags (key:value), view definition (for views)
ColumnsName, data type, nullability, comment, Unity Catalog tags, and nested fields of STRUCT, ARRAY, and MAP columns
VolumesName, comment (as description), storage location

Lineage

Sifflet builds table-level and column-level lineage from the system.access.column_lineage system table, which records the lineage that Unity Catalog captures for queries, notebooks, jobs, and pipelines. For each table in scope, Sifflet imports its upstream tables, including tables from catalogs and schemas outside the source scope.

When you also add the Databricks Workflows integration, Sifflet links each table to the Databricks job that last wrote it. For tables written by Lakeflow Declarative Pipelines, this requires the optional SELECT permission on system.lakeflow.pipeline_update_timeline (see Permissions Required).

Other integrations extend the lineage downstream and upstream of Databricks, for example BI tools such as Tableau or Power BI, and dbt.

Data Freshness

Sifflet refreshes Databricks metadata and lineage on the source schedule: see Integrations Management for the default frequency and how to change it. Databricks populates system tables with some delay, so lineage from the most recent queries and jobs can appear in Sifflet only at a later refresh.

Limitations

  • Only catalogs governed by Unity Catalog are supported: Sifflet doesn't import assets from the legacy hive_metastore catalog.
  • Lineage is limited to what Unity Catalog captures in system.access.column_lineage. See the Databricks lineage limitations.

Data Exposure

Sifflet reads metadata and runs aggregate queries for monitors. Two features read row-level data from your tables on demand, without storing it: data preview in the Data Catalog, and the sample of failing rows shown for some monitors. Monitor group-by values are cached, encrypted. See Security for what Sifflet stores and accesses.


Did this page help you?