Data ingestion

Every source, in. Fast.

Any source into your warehouse with ingestr: open source, and the fastest tool in 76 of our 79 benchmarks.

1M rows · Postgres → Postgres · lower is faster

ingestr By Bruin0.00s
Spark 0.00s
Sling 0.00s
dlt 0.00s
Airbyte 0.00s

Benchmarks

Fastest in 76 of 79 routes.

The same copy, on the same machine and data, for every tool.

Mean seconds to copy 10M rows, per route and tool. Lower is faster.
10M rowsingestrdltSlingSparkNext fastest
Postgres → Postgres64.50s (fastest)344s206s65.34sSpark 65.34s
Postgres → ClickHouse45.88s (fastest)157s210s92.24sSpark 92.24s
MySQL → ClickHouse51.69s (fastest)191s272s147sSpark 147s
DuckDB → Postgres45.21s (fastest)258s467s70.56sSpark 70.56s
Postgres → DuckDB73.26s (fastest)158s373s3888sdlt 158s

In your repo

One YAML file. One pipeline.

Each ingestion is a small asset in Git, with checks on the raw table. Run it from the CLI, or on a schedule in Bruin Cloud.

name: raw.orders
type: ingestr
parameters:
  source_connection: postgres
  source_table: 'public.orders'
  destination: bigquery
  incremental_strategy: merge
  incremental_key: updated_at

columns:
  - name: order_id
    primary_key: true
    checks:
      - name: not_null
  - name: status
    checks:
      - name: accepted_values
        value: [paid, refunded, pending]

$ bruin run assets/raw/orders.asset.yml

  1. extract · postgres public.orders84,102 new rows
  2. load · bigquery raw.ordersmerged on order_id
  3. not_null · order_id
  4. accepted_values · status

Loaded and checked. Downstream models can run.

Load strategies

Six ways to load.

Pick one per table. Two are new.

replace

Swap the destination table for the source. The default.

No keys needed

append

Add only rows newer than the last run.

Needs incremental_key

merge

Upsert by primary key: update what changed, insert what is new.

Needs primary_key + incremental_key

delete+insert

Replace a time window. Built for late data and backfills.

Needs incremental_key

truncate+insertNew

Empty the table and load fresh in one operation.

No keys needed

scd2New

Keep full row history with valid_from and valid_to.

Needs primary_key

Sources & destinations

Pick a source. Pick a destination.

Every pair has a step-by-step guide.

Popular
From
To

Pick a source and a destination.

Missing a source? Share test credentials and we build it within 7 days.

Browse all integrations

All available data pipeline guides

Open source

One binary. No per-row fees.

ingestr is source-available under the Functional Source License and runs anywhere a binary runs: your laptop, CI, Airflow or a cron. Bruin Cloud is optional.

Customer results

Numbers from teams on Bruin.

The platform

Part of the Bruin platform.

Data in, ready for everything downstream: the models, the checks, the lineage and the AI layer.

Your stack, your call

Replace the modern data stack.

One layer or every layer. Keep what works, swap what doesn't.

Today

On Bruin

Ingestion

On Bruin: Data Ingestionon open-source ingestr

Transformation & orchestration

On Bruin: SQL & PythonBruin Cloud

Quality, lineage & catalog

On Bruin: Data QualityData Governance

BI & dashboards

  • Power BI

On Bruin: AI DashboardsData Apps

AI on your data

On Bruin: AI Data AnalystScheduled Agents

Frequently asked

Questions about ingestion.

What is ingestr?

ingestr is Bruin's open-source ingestion engine: one CLI that copies data between databases, warehouses and SaaS tools. It is the ingestion layer of the Bruin platform, and it also runs on its own anywhere a binary runs.

Do I need Bruin Cloud to use it?

No. ingestr and the Bruin CLI are open source and run locally, in CI or in your own orchestrator. Bruin Cloud is the optional managed layer: it schedules ingestr jobs alongside SQL and Python pipelines and quality checks, with column-level lineage, audit logs, cost insights and the AI data analyst on top.

Which sources and destinations are supported?

BigQuery, Snowflake, Databricks, Redshift, Postgres, MySQL, SQL Server, ClickHouse, DuckDB, MotherDuck, MongoDB, DynamoDB, Elasticsearch, SQLite, Trino and more on the database side, plus SaaS sources such as Stripe, HubSpot, Salesforce, Google Ads, Facebook Ads, GitHub, Notion, Airtable and Klaviyo. Each supported pair has its own guide, and the full matrix lives in the ingestr README.

Does ingestr do CDC?

It does scheduled incremental replication: append, merge or delete+insert on an incremental key, every few minutes if you like, with no Kafka to run. That covers most replication needs. If you truly need sub-second, log-based streaming, pair Bruin with a dedicated CDC tool; our CDC guide covers when that is worth it.

Is there a per-row fee?

Not for ingestr: it is source-available under the Functional Source License and free to run yourself, so the only bill is your own compute and warehouse. On Bruin Cloud, pricing scales with usage, not seats, and starts free with $100 in credits.

What if our source isn't supported?

Share testing credentials and we implement it within 7 days. You can also write your own source in Python: Bruin runs Python assets next to ingestr ones in the same pipeline.

Which load strategies are supported?

Six: replace (the default), append, merge, delete+insert, truncate+insert and scd2. The last two are new. scd2 keeps a full history per primary key with valid_from and valid_to columns, so you can query any record as it was at any point in time, with no custom merge SQL.

Can I check data quality at ingestion?

Yes. An ingestion asset can declare column checks such as not_null and accepted_values, plus custom checks in SQL, on the raw table. Checks can use templated date ranges to test only the incremental window, and failing checks alert the configured channels before anything downstream runs.

Is the new ingestr a drop-in upgrade?

For the common path, yes: same CLI, same flags, same source and destination URIs. Under the hood it is a rewrite with lower memory use and higher throughput, plus the two new strategies and a double-buffered replace. Two things to check: nested JSON is now kept as one JSON column instead of being flattened, and BigQuery sources need the roles/bigquery.readSessionUser IAM role because ingestr now reads through the Storage Read API.

How were the benchmarks run?

Each tool runs the same source-to-destination copy for a fixed number of rows, on the same machine with identical source data. A warm-up run is discarded and the mean of the timed runs is reported. ingestr was the fastest tool in 76 of the 79 routes; every route is on the benchmarks page.

Can I run ingestr in Docker, CI/CD or Kubernetes?

Yes. It installs as a single binary with one command, so it drops into GitHub Actions, GitLab CI, Argo, Airflow, Dagster, Prefect, Kubernetes jobs or a plain cron.

Is there a web UI?

Yes. ingestr server starts a local web UI to save reusable connections, configure a run, stream its logs and review previous runs. The CLI stays the main surface for production pipelines.

What does ingestr collect about my usage?

Anonymous telemetry only: a hashed machine ID, the ingestr version, OS and architecture, the command that ran and whether it succeeded. No connection strings, no schemas, no data. Set INGESTR_DISABLE_TELEMETRY=true to turn it off.

Every source, in your warehouse.

$100 in credits and 50 AI tasks. No credit card.

A demo walks through your own data.

Sign up to our newsletter

Practical updates on open-source data pipelines, AI analysts, governance, and what we are shipping at Bruin.

The signup form is hosted by Brevo. Allow marketing cookies to load it.