Deploy to Databricks¶
Start with Spark + FileProvider and the platform smoke project. It reads three CSV rows from a Volume and overwrites a new sandbox UC table. These are locally checked setup contracts; no live Databricks run is claimed.
1. Select compute and dependencies¶
Use managed Spark for UC tables, distributed transforms and Delta operations. Native Polars is an option for bounded file workloads that fit one node; it does not provide Spark UC table semantics. See engine selection.
| Compute | Bootstrap |
|---|---|
| Classic jobs/all-purpose | Install matching datacoolie cluster library or notebook %pip |
| Serverless notebook/job | Configure its Environment dependencies and environment version |
| External Python 3.11+ | datacoolie[databricks-external,polars-delta]; SDK unified authentication |
For classic notebook installation:
Use a matching wheel for an unpublished checkout. Managed compute supplies
Spark/Delta; do not install an extra PySpark runtime. Python must satisfy
>=3.11,<4.0. Serverless environment 1 uses Python 3.10 and is unsuitable;
choose a compatible environment (for example environment 2 uses Python 3.11).
Pin and review the chosen environment version.
Serverless dependencies belong in the Environment configuration,
not compute-scoped libraries. Serverless uses Spark Connect and has API/config
restrictions; Python compatibility alone does not qualify every SparkEngine
operation. This guide's local checks do not establish serverless operation parity.
See serverless limitations.
2. Grant the execution identity access¶
Create a sandbox Volume and schema. Grant the job's Run as principal, rather than only the interactive author, the permissions needed for this recipe:
| Operation | UC privileges |
|---|---|
| Read Volume metadata/input | USE CATALOG, USE SCHEMA, READ VOLUME |
| Persist Volume logs/watermarks | Above plus WRITE VOLUME |
| Create the sandbox managed table | USE CATALOG, USE SCHEMA, CREATE TABLE on output schema |
| Write an existing table | MODIFY on table, plus namespace use |
| Verify table output | SELECT on table, plus namespace use |
Volume access does not grant table access. See privileges and permission concepts. The fixture needs no secrets or external database.
3. Build, upload and parameterize¶
Follow the shared build recipe. Edit
metadata/environments/databricks.json: replace the input Volume root and the
output catalog/database. Leave schema_name unset; empty base_path clears
the local output root during overlay merge. Validate/inspect --env databricks.
Upload .builds/current/databricks/metadata/ to
/Volumes/main/default/datacoolie_example/metadata/. Upload the CSV separately
to /Volumes/main/default/datacoolie_example/data/input/orders/orders.csv.
Replace these example namespaces consistently for your workspace.
Import canonical databricks/run_spark.ipynb
(source ·
raw). It selects
runtime="databricks", uses spark and FileProvider. Configure parameter cell
values interactively or the same widget names as Workflow notebook parameters:
| Widget/job parameter | Smoke value |
|---|---|
METADATA_PATH |
/Volumes/main/default/datacoolie_example/metadata/metadata.json |
CONNECTIONS_PATH, SCHEMA_HINTS_PATH |
Empty strings; included in built metadata |
WATERMARK_BASE_PATH |
/Volumes/main/default/datacoolie_example/.runtime/watermarks |
LOG_BASE_PATH |
/Volumes/main/default/datacoolie_example/.runtime/logs |
STAGE |
platform_smoke |
JOB_NUM, JOB_INDEX |
1, 0 |
Set the Workflow Run as identity, notebook task, compatible compute/environment
and dependencies explicitly. Apply shared smoke selection/result guards at the
runner's driver.run boundary for a downstream barrier; generic runners raise
on failed flows. There is no repo-owned Jobs API provisioning payload.
4. Distinguish table addressing from file addressing¶
These are alternative output connection configurations:
{"connection_type": "lakehouse", "format": "delta", "catalog": "main", "database": "default", "configure": {"base_path": ""}}
With destination orders_platform_smoke, this resolves to
main.default.orders_platform_smoke. DataCoolie uses named table addressing
when catalog/database is set; a Volume base_path alongside it is not the
write target. catalog + database + table is UC's three-part name; setting
schema_name adds another component.
For a path-only Delta sandbox, omit catalog/database/schema_name entirely (or clear inherited namespace fields in an overlay):
{"connection_type": "lakehouse", "format": "delta", "configure": {"base_path": "/Volumes/main/default/datacoolie_example/data/output"}}
Its destination is <base_path>/orders_platform_smoke. You may store Delta
files in a Volume, but may not register a UC table on those Volume files; table
and Volume locations cannot overlap. See Volume paths.
dbfs:/Volumes/... is an accepted alias; canonical paths are /Volumes/....
Raw s3://, abfss://, gs:// are native-only platform paths. Workspace Files
are outside the platform's portable contract. Avoid new DBFS root/mount usage;
see DBFS guidance.
5. Verify the named table¶
Require selection orders_platform_smoke and result counts (1, 1, 0, 0) for
total, succeeded, failed, pending. Independently read the chosen namespace:
output = spark.table("main.default.orders_platform_smoke").select(
"order_id", "customer_id", "amount"
).orderBy("order_id")
assert [tuple(row) for row in output.collect()] == [(1, 100, 20), (2, 100, 43), (3, 101, 7)]
assert output.dtypes == [("order_id", "bigint"), ("customer_id", "bigint"), ("amount", "bigint")]
Inspect Volume execution/system logs and the Workflow task failure status.
For the path-only branch use spark.read.format("delta").load(...) instead.
Optional providers, secrets and Polars¶
Database metadata is a separate DatabaseProvider setup,
including bootstrap/migrations, credentials, networking and the SQLAlchemy
backend driver (for PostgreSQL, psycopg2-binary). It is not needed for this run.
Secret resolution uses native dbutils.secrets or external SDK secrets:
{"configure": {"password": "sample-db-password"}, "secrets_ref": {"datacoolie-scope": ["password"]}}
Grant the executing principal secret-scope read access separately. Keep actual values out of notebooks/logs.
For native file-only Polars use the native Polars scenario notebook; install the
required engine/format profile and prepare file/path metadata rather than the
smoke project's named UC target.
For external SDK use databricks/run_polars_sdk.py
(source ·
raw); run --help.
SDK Files API access to Volume control files does not mount /Volumes locally,
submit remote Spark or configure Polars business storage credentials. Business
connections must address storage directly accessible to that process.
Troubleshooting and next steps¶
| Symptom | Check |
|---|---|
| Volume reads work but table denied | Output namespace/table grants to Run as identity |
| Output appears as a table instead of Volume files | Inherited catalog/database in effective metadata |
| Serverless dependency/API error | Environment version/dependencies and Spark Connect limitations |
| Zero selected flows | Built metadata root, stage and shard widgets |
External Polars cannot read /Volumes |
SDK access is not a local filesystem mount |
Use operations for replay, sharding, maintenance and functions. The Databricks simulator and WWI walkthrough are larger scenarios requiring separate preparation.