Apache Kafka integration

Apache Kafka data, in your warehouse.

A built-in ingestr connector: add credentials, pick tables, schedule it.

name: raw.table
type: ingestr
parameters:
  source_connection: apache_kafka
  source_table: '<table>'
  destination: snowflake
  incremental_strategy: merge

$ bruin run assets/raw/apache_kafka.asset.yml

  1. extract · Apache Kafka dataincremental
  2. load · snowflake raw.datamerged
  3. checks · not_null, unique

Loaded and checked. Downstream models can run.

How it connects

Connected in three steps.

Distributed event streaming platform used for high-performance data pipelines, streaming analytics, data integration, and mission-critical applications.

  1. 01

    Add a Apache Kafka connection with its credentials.

  2. 02

    Pick the tables to load and how: replace, append or merge.

  3. 03

    Bruin runs it on your schedule and checks every load.

Connection parameters

bootstrap_servers
Kafka server or servers to connect to (host:port format)
group_id
Consumer group ID for identifying the client
security_protocol
Protocol for broker communication (e.g., SASL_SSL)
sasl_mechanisms
SASL mechanism for authentication (e.g., PLAIN)
sasl_username
Username for SASL authentication
sasl_password
Password for SASL authentication
batch_size
Number of messages to fetch per batch (default: 3000)
batch_timeout
Maximum wait time for messages in seconds (default: 3)

The platform

Part of the Bruin platform.

Data in, ready for everything downstream: the models, the checks, the lineage and the AI layer.

Frequently asked

Questions about Apache Kafka.

Does Bruin have a built-in Apache Kafka integration?

Yes. Built-in ingestr source.

Where can Apache Kafka data go?

Snowflake, BigQuery, Databricks, Redshift, ClickHouse, Postgres, DuckDB, MotherDuck, Microsoft Fabric and more, plus files on S3 and GCS.

How fresh is the data?

As fresh as your schedule. Incremental loads append, merge or replace a time window, every few minutes if you like.

Do we need Bruin Cloud?

No. The Bruin CLI and ingestr run locally, in CI or in your own orchestrator. Bruin Cloud adds scheduling, lineage, alerts and the AI data analyst on top.

Ready to connect Apache Kafka?

$100 in credits and 50 AI tasks. No credit card.

A demo walks through your own data.

Sign up to our newsletter

Practical updates on open-source data pipelines, AI analysts, governance, and what we are shipping at Bruin.

The signup form is hosted by Brevo. Allow marketing cookies to load it.