Data Observability software monitors the health and reliability of data pipelines to detect anomalies and quality issues. These tools track data lineage and identify failures in batch or streaming ingestion before they affect downstream analytics. This type of software serves data engineers and analysts who need to verify the accuracy of SQL queries and warehouse tables. Self-hosting these apps ensures that sensitive schema metadata and pipeline logs remain on private infrastructure rather than on external servers.
This page lists 2 open source tools in the Data Observability category. The most popular are Apache Druid and Elementary Data. Most use the Apache-2.0 license, and 2 offer an official Docker image.
A high performance real-time analytics database designed for fast queries and high concurrency data ingestion.
A dbt-native data observability tool for detecting anomalies and monitoring data quality within data warehouse pipelines.
Join our newsletter to get shiny new open source software delivered to your inbox. Unsubscribe anytime.