Lakeflow Connect is the ingestion part of Databricks Lakeflow. It's a set of connectors that move data from local files, enterprise applications, databases, cloud storage and message buses into the lakehouse.
Databricks' own ingestion overview, last updated on 7 October 2026, organises it by the type of source you read from, and Databricks announced general availability of Lakeflow Connect for Salesforce Platform and Workday Reports on its blog. It sits alongside the other two pieces of Lakeflow, Spark Declarative Pipelines (previously DLT) and Lakeflow Jobs (previously Workflows), which handle transformation and orchestration.
The six kinds of connector #
Databricks groups the connectors by source type.
Database connectors read relational systems such as MySQL, PostgreSQL, Oracle and SQL Server using change data capture. SaaS connectors read enterprise applications including Salesforce, HubSpot, Jira and Workday. File connectors pick up structured and unstructured files from cloud object storage and from file services such as Google Drive and SharePoint. Streaming connectors read continuously from message buses and event streaming sources such as Apache Kafka and RabbitMQ.
Two more sit outside that shape. Query-based connectors read a database by querying the source directly, without change data capture, which matters when you cannot turn CDC on at the source or cannot get the permissions for it. Direct write, the Zerobus Ingest API, goes the other way. Your application writes records straight into streaming tables, so nothing has to poll anything.
Managed service or your own pipeline #
Databricks offers two service models for the same job. The fully managed one gives you out-of-the-box connectors with UIs and APIs, which is how you build an ingestion pipeline without owning the long-term maintenance of it. When you need more control, you write a custom pipeline with Lakeflow pipelines or with Structured Streaming instead, and the documentation treats that as a normal choice rather than an escape hatch.
What holds the two together is the rest of the platform. Databricks says Lakeflow Connect uses Unity Catalog for governance, Lakeflow Jobs for orchestration, and monitoring that spans pipelines, so an ingested table arrives with grants, lineage and run history already attached. It also uses incremental reads and writes, which Databricks says improves ETL performance when combined with incremental transformations downstream.
Where it fits in the Field Guide #
Ingestion covers picking the simplest mechanism that meets a requirement, and Lakeflow Connect is usually that mechanism when a managed connector exists for your source. Real-Time and Change Data Capture covers the database connectors and the latency decision behind them. For what happens after the data lands, see Pipelines and our note on how pipelines declare tables while jobs run tasks.
Questions people ask #
What is Lakeflow Connect in Databricks? #
Lakeflow Connect is the Databricks ingestion layer. It offers connectors that bring data from local files, enterprise applications, databases, cloud storage and message buses into the lakehouse, with governance through Unity Catalog and orchestration through Lakeflow Jobs.
What sources does Lakeflow Connect support? #
Databricks groups its connectors into database connectors using change data capture (MySQL, PostgreSQL, Oracle, SQL Server), SaaS connectors (Salesforce, HubSpot, Jira, Workday), file connectors for cloud object storage and services such as Google Drive and SharePoint, streaming connectors for Apache Kafka and RabbitMQ, query-based connectors that read the source directly, and direct write through the Zerobus Ingest API.
Is Lakeflow Connect the same as Lakeflow? #
No. Lakeflow is the unified data engineering product, and Lakeflow Connect is its ingestion component. The other two are Lakeflow Spark Declarative Pipelines, previously known as DLT, and Lakeflow Jobs, previously known as Workflows.