Apache Kafka integration
Apache Kafka data, in your warehouse.
A built-in ingestr connector: add credentials, pick tables, schedule it.
name: raw.table
type: ingestr
parameters:
source_connection: apache_kafka
source_table: '<table>'
destination: snowflake
incremental_strategy: merge$ bruin run assets/raw/apache_kafka.asset.yml
- extract · Apache Kafka dataincremental
- load · snowflake raw.datamerged
- checks · not_null, unique
Loaded and checked. Downstream models can run.
How it connects
Connected in three steps.
Distributed event streaming platform used for high-performance data pipelines, streaming analytics, data integration, and mission-critical applications.
- 01
Add a Apache Kafka connection with its credentials.
- 02
Pick the tables to load and how: replace, append or merge.
- 03
Bruin runs it on your schedule and checks every load.
Connection parameters
- bootstrap_servers
- Kafka server or servers to connect to (host:port format)
- group_id
- Consumer group ID for identifying the client
- security_protocol
- Protocol for broker communication (e.g., SASL_SSL)
- sasl_mechanisms
- SASL mechanism for authentication (e.g., PLAIN)
- sasl_username
- Username for SASL authentication
- sasl_password
- Password for SASL authentication
- batch_size
- Number of messages to fetch per batch (default: 3000)
- batch_timeout
- Maximum wait time for messages in seconds (default: 3)
Step-by-step
Pick a destination.
Apache Kafka → Snowflake
Read the guide
Apache Kafka → Google BigQuery
Read the guide
Apache Kafka → Databricks
Read the guide
Apache Kafka → Amazon Redshift
Read the guide
Apache Kafka → PostgreSQL
Read the guide
Apache Kafka → ClickHouse
Read the guide
Apache Kafka → DuckDB
Read the guide
Apache Kafka → MotherDuck
Read the guide
The platform
Part of the Bruin platform.
Data in, ready for everything downstream: the models, the checks, the lineage and the AI layer.
Sources
DatabasesWarehousesApps & APIsFiles & storageStreams & webhooksWeb scrapingMove
Data IngestionStreams & queues
More tools, same category.
Frequently asked
Questions about Apache Kafka.
Does Bruin have a built-in Apache Kafka integration?
Yes. Built-in ingestr source.
Where can Apache Kafka data go?
Snowflake, BigQuery, Databricks, Redshift, ClickHouse, Postgres, DuckDB, MotherDuck, Microsoft Fabric and more, plus files on S3 and GCS.
How fresh is the data?
As fresh as your schedule. Incremental loads append, merge or replace a time window, every few minutes if you like.
Do we need Bruin Cloud?
No. The Bruin CLI and ingestr run locally, in CI or in your own orchestrator. Bruin Cloud adds scheduling, lineage, alerts and the AI data analyst on top.
Ready to connect Apache Kafka?
$100 in credits and 50 AI tasks. No credit card.
A demo walks through your own data.