Redshift Collected Data

This page lists what Sifflet imports from Amazon Redshift once you connect your cluster or workgroup.

Imported Assets

Sifflet imports the following objects from the schemas selected in the source scope, on provisioned clusters and Redshift Serverless:

  • Tables
  • Views, including materialized views

Collected Metadata

Asset typeMetadata
Tables and viewsName, type, comment (as description)
ColumnsName, data type, nullability, comment

Lineage

On provisioned clusters, Sifflet builds table-level and field-level lineage by parsing the SQL text of the queries run on the cluster, read from the SVL_STATEMENTTEXT system view. This requires the optional permissions described in Permissions Required: without SYSLOG ACCESS UNRESTRICTED, Sifflet only sees its own queries, and lineage is mostly empty.

On Redshift Serverless, Sifflet doesn't build lineage from query history. Use the dbt integration or declarative lineage instead.

Other integrations extend the lineage upstream and downstream of Redshift, for example dbt and BI tools such as Tableau, Looker, or Amazon QuickSight.

Data Freshness

Sifflet refreshes Redshift metadata and lineage on the source schedule: see Integrations Management for the default frequency and how to change it. Each refresh reads recent queries and adds the lineage they reveal to the lineage Sifflet already knows. Refresh the source at least once a day so that no queries are missed.

Limitations

  • No lineage from query history on Redshift Serverless.
  • External tables (Amazon Redshift Spectrum) aren't imported.

Data Exposure

Sifflet reads metadata and runs aggregate queries for monitors. Two features read row-level data from your tables on demand, without storing it: data preview in the Data Catalog, and the sample of failing rows shown for some monitors. For lineage, Sifflet reads the text of your queries, which can contain literal values. See Security for what Sifflet stores and accesses.


Did this page help you?