Databricks Field Guide

Lakeflow Jobs is the orchestrator, and its history is a queryable table

Lakeflow Jobs is the orchestration service built into Databricks, and the documentation, last updated on 11 September 2026, defines it as workflow automation that coordinates and runs multiple tasks as part of a larger workflow. A job holds one or more tasks arranged as a directed acyclic graph, with triggers that decide when it runs, parameters pushed down to every task, notifications on failure or overrun, and Git settings for the task source. It covers ETL, notebooks, machine learning work and integrations with external systems such as dbt.

Jobs, tasks and triggers #

The documentation names three concepts. A job is the resource that coordinates and schedules work, and it can be a single notebook or hundreds of tasks with conditional logic. A task is one unit of work inside it: a notebook task, a pipeline task that runs a Lakeflow pipeline such as a materialised view or streaming table, a Python script task, and many more types. A trigger starts the job, either on a schedule or on an event such as new data arriving in cloud storage.

Control flow is part of the product rather than something you code around. Jobs support branching with if/else and looping with for each, authored visually, and tasks can depend on other tasks or run conditionally.

The part teams underuse #

Every job and pipeline in the account writes to the system.lakeflow schema, documented in the jobs system table reference updated on 8 October 2026. The schema holds six tables: jobs, job_tasks, job_run_timeline, job_task_run_timeline, pipelines and pipeline_update_timeline. All six support streaming, all six have a free retention period of 365 days, and all six are regional, so records from another region need a workspace in that region to read. Records for the jobs and run timeline tables are typically available within an hour.

Three of them, jobs, job_tasks and pipelines, are slowly changing dimension tables. A change emits a new row rather than overwriting, and a deletion is a new row with delete_time populated. Getting the current picture means taking the latest row per entity first and filtering deletions after:

with latest AS (
  SELECT *,
    ROW_NUMBER() OVER (PARTITION BY workspace_id, job_id ORDER BY change_time DESC) as rn
  FROM system.lakeflow.jobs
  QUALIFY rn = 1
)
SELECT * FROM latest WHERE delete_time IS NULL

Filter delete_time inside the same query that reads the table and you get each deleted job's last state instead of omitting it. Access needs metastore admin and account admin, or USE and SELECT on the system schemas.

The rename to watch #

The lakeflow schema was previously called workflow, and Databricks says the content of both schemas is identical. Dashboards and alerts written against the old name still work, though new ones should use lakeflow. Note also that the trigger and trigger_type columns are not populated for rows emitted before early December 2025, which matters if you query far back.

More on designing jobs is in Orchestration, and on what belongs in a pipeline instead in Pipelines.

Questions people ask #

What is the role of Databricks Lakeflow Jobs? #

Lakeflow Jobs is workflow automation for Databricks that provides orchestration for data processing workloads. It schedules and coordinates multiple tasks, such as notebooks, pipelines and Python scripts, as a single workflow represented as a directed acyclic graph.

Where can I see the history of Databricks job runs? #

Job activity is recorded in the system.lakeflow schema, which holds six tables covering jobs, job tasks, job run timelines, pipelines and pipeline updates. Each has a free retention period of 365 days, and records are typically available within one hour.

What was the Databricks lakeflow system schema called before? #

The lakeflow schema was previously known as workflow. Databricks states that the content of both schemas is identical.

Next

Databricks Apps can now run as the user who signed in