Define and deploy your data monitoring strategy

Guidelines for choosing and rolling out a data monitoring strategy with Sifflet.

When you start deploying Sifflet on your data infrastructure, you need to think about which monitoring strategy to start with.

The answer always depends on your specific data organization and use cases, but this page gives you a few guidelines to get started.


1. Connect your sources to Sifflet

Before you start defining monitors on your data, you must first connect Sifflet to your data sources.

For monitoring, the most important sources are your data warehouses. Refer to the documentation page for your data warehouse to connect it to Sifflet (for example, Snowflake, BigQuery, Databricks or Redshift).

2. Define your monitoring strategy

When you're getting started with data observability, there are two possible approaches.

If you're currently experiencing data incidents and want to start reducing their frequency without drowning in noise you can't address, use the strategy Start with the data products.

If you currently don't have a lot of data incidents and want immediate visibility on the entirety of your data stack, use the strategy Monitor everything.

a. Start with the data products

This strategy focuses on reducing the number of data incidents while keeping the alerting noise under control.

The idea is to start with a few monitors on critical assets, and once these monitors no longer report any issue, extend monitoring to another set of assets.

To deploy it:

  1. Identify critical assets. Pick a few data assets (tables, dashboards, etc.) that are important to your business. These are the tables or dashboards that are directly used by teams outside your data engineering teams. The Data Catalog helps you find them.
  2. Group them into data products. In Sifflet, create data products to encompass these assets.
  3. Add a few critical monitors. For each table in these data products, set up one or two critical monitors, typically freshness and volume, with dynamic thresholds. Don't go overboard here: you'll add more monitors once these initial monitors stay green for a while.
  4. Enable data product notifications. Configure data product notifications or notification rules scoped to your data products. Now, you have visibility when you have data issues on these assets.
  5. Fix the underlying issues. Improve your data pipeline until you no longer regularly get alerts on these data products.
    • To help investigate the root cause of issues, use lineage to find the upstream tables of these assets and add monitors on them. Again, freshness and volume monitors are usually a good start.
  6. Expand the scope. Then, identify a few more data products and repeat the process.

b. Define freshness, volume and schema changes monitors on everything

If you want a complete view of the state of your data quality, you can also start by defining common monitors on all your assets.

These common monitors are:

  • Freshness: is my data up-to-date?
  • Volume: does the number of rows in the table match my expectations?
  • Schema change: if the schema of a table changes, this can break its consumers.

You have several options to define all these monitors:

This is the noisiest strategy: with monitors on every table, you'll get failures on assets that nobody depends on. To keep it manageable, don't send notifications for these monitors at first. Instead, review their results in Sifflet, and only enable notifications (for example, with notification rules) on the assets that matter to your business.

Combine both strategies

The two strategies aren't mutually exclusive. You can define common monitors on everything to get visibility over your whole data stack, but only route notifications for your data products and their upstream tables. You get complete coverage, while your team only receives relevant alerts.

3. Additional guidance

Use dynamic thresholds when you have many monitors

When defining freshness and volume monitors at scale, consider dynamic thresholds over static ones.

Your data doesn't behave the same way every day: a table may receive more rows on weekdays than on weekends, grow steadily over time, or spike at the end of each month. A static threshold can't account for this. If it's too tight, it alerts on normal variations and creates noise; if it's too loose, it misses real incidents. And tuning static thresholds by hand for hundreds of tables isn't realistic.

Dynamic thresholds are computed by ML models trained on the history of each monitored metric. They learn the normal behavior of each table, including trends and seasonality, and only alert when the data deviates from it. This makes them a good fit when you deploy monitors at scale and don't know in advance what "normal" looks like for each table.

Keep static thresholds for known business rules, for example when a value must stay within a given range. See Monitor Setup for the available modes.

If a dynamic monitor is too noisy, you can tune it:

  • Adjust its sensitivity to only alert on significant deviations.
  • Use the feedback loop to flag false positives, so the model learns from them.

Mute noisy monitors

If your data team receives too many alerts, they'll start to ignore them. It's generally better to have fewer notifications that are actually actioned by your team, rather than many notifications that are ignored.

If a monitor alerts too frequently, or if your team usually ignores its notifications, mute it until the underlying cause is fixed. In Sifflet, you can pause a monitor's notifications until its next status change: this way, you're notified again as soon as the situation evolves, and you won't forget to turn notifications back on.

You usually don't want to delete monitors that alert frequently: they carry useful information. Instead, adjust their notification configuration, for example by excluding them from a notification rule or by removing their notification channels.

Configure data product incident notifications to alert business teams

Business teams are generally not interested in the failure of individual monitors. However, they are interested to know when a given data product breaks.

Sifflet supports notifying a different team when an incident occurs on a data product:


Did this page help you?