GCP Dataproc Serverless integration

Pipelines that run on GCP Dataproc Serverless.

SQL and Python assets run on GCP Dataproc Serverless, with checks on every table.

  • Query engineRuns SQL and Python assetsDocs →

How it connects

Connected in three steps.

Serverless Spark service on Google Cloud that lets you run Spark workloads without provisioning or managing clusters, with automatic scaling and resource management.

  1. 01

    Add a GCP Dataproc Serverless connection.

  2. 02

    Write SQL or Python assets that run on GCP Dataproc Serverless.

  3. 03

    Bruin builds the graph, runs it and checks every table.

Tables

3 tables, ready to load.

  • spark_tables
  • bigquery_tables
  • hive_metastore_tables

The platform

Part of the Bruin platform.

Data in, ready for everything downstream: the models, the checks, the lineage and the AI layer.

Frequently asked

Questions about GCP Dataproc Serverless.

Does Bruin have a built-in GCP Dataproc Serverless integration?

Yes. Runs SQL and Python assets.

Which GCP Dataproc Serverless tables can Bruin load?

spark_tables, bigquery_tables, hive_metastore_tables.

How fresh is the data?

As fresh as your schedule. Incremental loads append, merge or replace a time window, every few minutes if you like.

Do we need Bruin Cloud?

No. The Bruin CLI and ingestr run locally, in CI or in your own orchestrator. Bruin Cloud adds scheduling, lineage, alerts and the AI data analyst on top.

Ready to connect GCP Dataproc Serverless?

$100 in credits and 50 AI tasks. No credit card.

A demo walks through your own data.

Sign up to our newsletter

Practical updates on open-source data pipelines, AI analysts, governance, and what we are shipping at Bruin.

The signup form is hosted by Brevo. Allow marketing cookies to load it.