In Databricks Lakeflow, Lakeflow Pipelines is the declarative layer and Lakeflow Jobs is the orchestration layer, and Databricks' own answer on the question says both are generally available. A pipeline runs as a single task inside a job, so the question is rarely which product to pick and almost always which layer a given problem belongs to.
What each layer actually does #
A job runs tasks in an order you define. You draw the dependency graph yourself: ingest, then transform, then refresh the dashboard, with retries and schedules attached to each step. A pipeline goes the other way. You declare the tables you want and the engine infers the execution graph from the datasets your code references, as the Lakeflow pipelines guide puts it. Because the framework handles orchestration, checkpointing, retries and incremental processing inside the pipeline, the work left to you at each stage is a design decision instead of an implementation.
That inference is the real difference. In a job, if task B depends on task A, you say so. In a pipeline, if your silver table reads the bronze table, the ordering already exists in the code and Databricks reads it out.
CREATE OR REFRESH STREAMING TABLE customers_cdc_clean(
CONSTRAINT valid_id EXPECT (id IS NOT NULL) ON VIOLATION DROP ROW
) AS SELECT * FROM STREAM(customers_cdc);Where the boundary sits #
Databricks draws a second line inside the declarative layer. A single materialized view or streaming table can be defined in SQL as a standalone dataset, and Databricks creates and manages the refresh pipeline behind the scenes. You author and operate a full Lakeflow pipeline as a unit when you need Python authoring, sinks, or multi-stage orchestration. Both run on the same declarative engine and both produce Unity Catalog managed tables.
So the practical sequence starts with a standalone dataset if one table is all you want. Graduate to a pipeline when you have several that depend on each other, and wrap the pipeline in a job when something outside the table graph has to happen in order. Model training, a notification, an export to a downstream system, all of those belong to the job.
The decisions a pipeline still asks of you #
Four choices set the starting configuration, per the Databricks guide. SQL or Python, decided per file so you can mix both. Serverless or classic compute, with serverless the recommended default and classic reserved for specific instance types, custom cluster policies or an init script. Triggered or continuous execution, where triggered only consumes compute while it runs and continuous keeps compute running, which is usually the largest cost factor. And the type of each output, a streaming table for append-heavy incremental data or a materialized view for recomputed aggregates and joins.
That last one drives cost and correctness, because incremental processing scales with the rate of new data rather than with table size. Pick it wrong and you pay to recompute history every refresh.
For the orchestration side, including how run history lands in system tables, see our note on Lakeflow Jobs. The chapters cover each layer in more depth, Pipelines and Orchestration.
Questions people ask #
What is the difference between Databricks pipelines and jobs? #
Lakeflow Pipelines is the declarative layer, where you declare the tables you want and the engine infers the execution order from the datasets your code references. Lakeflow Jobs is the orchestration layer, and it runs tasks in an order you define. A pipeline runs as a single task inside a job.
Are Lakeflow Pipelines and Lakeflow Jobs generally available? #
Yes. Databricks states that both the declarative pipeline layer and the Lakeflow Jobs orchestration layer are generally available.
Should I use triggered or continuous pipeline mode? #
Databricks recommends starting with triggered mode, because it only consumes compute while it runs. Continuous mode keeps compute running to process new data with minimal delay, which is usually the largest cost factor, so it should be reserved for a proven latency requirement.