Databricks Field Guide

Analysing and serving · Chapter 23

Databricks Apps

Every governed platform eventually meets the same request. Somebody needs a screen. A review queue, an approval form, a small internal tool that reads three tables and writes one. The data is already governed, and the moment you build a screen for it, the governance usually stops at the door.

The short version

Databricks Apps runs a normal web application on Databricks itself, so the thing your users click sits inside the same boundary as the data it reads. You write Streamlit, Dash or Gradio in Python, or React, Angular, Svelte or Express in Node, and you deploy it without provisioning a server, a container registry, a secret store or a login page. The part that matters is not the hosting. It's that the app can act as the person using it, so a user who cannot see a row in a table cannot see it in your interface either, without you writing a single line to enforce that.

What it actually is #

A serverless place to run a web application, integrated with Unity Catalog, SQL warehouses and OAuth. Databricks describes the point as removing "the need for separate infrastructure", and in practice that is four things you no longer stand up: somewhere to run the process, a way to authenticate users, a way to hold credentials, and a network path to the data.

The supported frameworks are deliberately boring, which is the right instinct. Streamlit, Dash and Gradio cover the internal tool that a data team writes in an afternoon. React, Angular, Svelte and Express cover the application that a product team maintains for years.

Billing follows the same shape as the rest of the platform. Apps are billed per hour of compute while running, against provisioned capacity, so an app nobody opens still costs something if it's left running, and a busy one costs what its capacity costs rather than what its traffic costs.

The two identities #

This is the chapter's real subject, and it's the decision people get wrong.

An app can act under its own identity or under the identity of whoever is using it, and the two behave completely differently at the moment somebody asks for data.

Every app gets a dedicated service principal, automatically provisioned, which cannot be changed or reused across applications. When the app acts as that service principal, it holds one set of grants. The documentation is direct about the consequence: "All users who interact with the app share the same permissions defined for the service principal, which prevents the app from enforcing fine-grained policies based on individual user identity."

User authorization, usually called on-behalf-of, does the other thing. The app borrows the identity of the person clicking, and "Databricks enforces all permissions based on the user's existing Unity Catalog policies." The sentence to hold onto is the next one: "Row-level filters and column masks apply automatically when the app accesses data." Whatever you built in Unity Catalog reaches the screen without being rebuilt in the application layer.

App service principal

On behalf of user

Person clicks in the app

Which identity?

One set of grants
for everybody

The clicker's own grants

Query runs once
same rows for all users

Row filters and column masks
applied by Unity Catalog

Query returns
this person's rows

Rendered screen

The same request, resolved under each identity

An app may use both at once, and mature ones usually do. Reading the data a person is entitled to happens on behalf of the user. Writing an audit record, calling an external service or reading shared configuration happens as the service principal, because those actions belong to the application rather than to whoever happened to trigger them.

On-behalf-of apps declare authorization scopes, and a user consents to them on first use. If you declare nothing, Databricks assigns a default that grants no data access at all, which is a sensible failure mode.

The caveat is worth reading twice, because it's the sharpest edge in the feature: "After granting consent, users can't revoke it." A person who clicks through a consent screen has made a decision they cannot personally undo. Workspace admins can restrict which scopes developers are allowed to request, and that admin control is the actual mitigation. Decide the scope list deliberately, before the first user ever sees it, and treat widening it later as a change that needs the same review as widening a grant.

When not to use it #

An application serving anonymous users on the public internet is the wrong shape for this. The whole model assumes an authenticated identity that Unity Catalog recognises, and a public marketing site has no such thing.

An application whose load is spiky and mostly idle will pay for provisioned capacity it isn't using. That is a real cost and it argues for the traffic living somewhere that scales to zero, with only the governed data access reaching back into the platform.

And a team with a deployment platform they like, a working single sign-on, and no governance drift is being asked to solve a problem they've already solved. The honest recommendation there is to leave it alone and spend the effort on the parts of the estate that are still guessing who can see what.