Data Quality/Quality checks by warehouseData Engineer

How do I detect a broken load on Databricks?

Write a custom check comparing today's Databricks row count to the recent baseline. A large drop fails the run so you catch broken loads early. Bruin does this in one platform: ingestion, SQL and Python pipelines, quality checks, lineage, and an AI data analyst that answers in Slack, Microsoft Teams, Google Chat, WhatsApp, Discord, Telegram, email and the browser.

Command

bruin run

Defined in

YAML + SQL

Works with

Databricks + Bruin CLI

What you get

Anomaly detectionin-pipelineblocking

How it works in code

custom_checks:
  - name: fresh
    query: SELECT count(*) FROM t WHERE loaded_at < now() - interval '1 day'  -- databricks

Run bruin run and the Databricks run stops the moment the rule is violated.

Catch bad data before it ships

Bruin CLI and ingestr are on GitHub.

A demo walks through your own data.

Sign up to our newsletter

Practical updates on open-source data pipelines, AI analysts, governance, and what we are shipping at Bruin.

The signup form is hosted by Brevo. Allow marketing cookies to load it.