Skip to content

Deploy to Microsoft Fabric

Start with Spark + FileProvider and the shared platform smoke project. It reads three CSV rows and overwrites a new sandbox Delta path. This guide is a locally checked setup contract; it does not report a live Fabric run.

1. Choose the host and prepare access

Host Installation Boundary
Fabric Spark notebook Base datacoolie; reuse supplied spark NotebookUtils for control files; Spark connector for business data
Fabric native Python datacoolie[polars-delta], Python 3.11+ Attached lakehouse mount for this fixture
External Python 3.11+ datacoolie[fabric-external,polars-delta] Azure SDK control access; separate engine storage credentials

Use Spark for distributed work, MERGE/SCD and Gold tables requiring Spark Delta features. Use Polars for bounded single-node file work after measuring memory/runtime. See engine selection.

Create a sandbox lakehouse, attach it as the notebook's default lakehouse, and choose a runtime compatible with Python >=3.11,<4.0. Install a matching release:

%pip install "datacoolie==0.2.0"

For an unpublished checkout, install its matching wheel. Fabric supplies Spark/Delta; avoid installing another PySpark runtime into the host.

The execution identity needs metadata/input read access and output/log/watermark read/write access. Workspace Contributor is a convenient sandbox role with broader access than a restricted production policy. See OneLake security. Interactive runs use the current user; schedules use their creator/last updater. For pipelines, inspect the Notebook activity's configured connection identity; Workspace Identity authentication is also available and needs access to the notebook and data workspaces. Interactive success does not prove scheduler access. Notebook identities, pipeline Notebook activity.

2. Build and upload

Follow the shared download/build recipe. Replace workspace/lakehouse identifiers in both roots in metadata/environments/fabric.json; validate, inspect --env fabric, then build. Use a qualified prefix:

abfss://your-workspace@onelake.dfs.fabric.microsoft.com/your-lakehouse.Lakehouse/Files/datacoolie-example

Upload .builds/current/fabric/metadata/ to <root>/metadata/. Upload the CSV separately to <root>/data/input/orders/orders.csv. Output is <root>/data/output/orders_platform_smoke; reserve it for this overwrite test. Changing notebook parameters does not rewrite connection paths in metadata.

3. Import and run

Import canonical fabric/run_spark.ipynb (source · raw). It explicitly selects runtime="fabric" and reuses the supplied Spark session.

Parameter cell / pipeline parameter Smoke value
METADATA_PATH <root>/metadata/metadata.json (exact built file)
CONNECTIONS_PATH, SCHEMA_HINTS_PATH None; included in built metadata
WATERMARK_BASE_PATH <root>/.runtime/watermarks
LOG_BASE_PATH <root>/.runtime/logs
STAGE platform_smoke
JOB_NUM, JOB_INDEX 1, 0

Edit the parameter cell or pass these exact names through the pipeline activity. Apply the shared smoke selection/result guards at its driver.run boundary when gating downstream work. The generic runner raises on failed dataflows.

4. Verify output and logs

Require selection orders_platform_smoke and counts total=1, succeeded=1, failed=0, pending=0. Read the actual Delta destination:

output = spark.read.format("delta").load(
    "abfss://your-workspace@onelake.dfs.fabric.microsoft.com/"
    "your-lakehouse.Lakehouse/Files/datacoolie-example/data/output/orders_platform_smoke"
).select("order_id", "customer_id", "amount").orderBy("order_id")
assert [tuple(row) for row in output.collect()] == [(1, 100, 20), (2, 100, 43), (3, 101, 7)]
assert output.dtypes == [("order_id", "bigint"), ("customer_id", "bigint"), ("amount", "bigint")]

Inspect execution/system logs under LOG_BASE_PATH for job ID and diagnostics. A successful notebook cell alone does not prove selection or fresh output. This fixture writes under Files/; it does not register a lakehouse Tables/ table.

Native Python and external SDK

NotebookUtils relative paths differ by kernel. The platform passes paths through to its backend.

Path Spark notebook Native Python External SDK
Files/... Default lakehouse relative Python working-directory relative; avoid here No default lakehouse
/lakehouse/default/Files/... Attached mount where available Use for this fixture Unavailable on laptop
Qualified abfss://... Host connectors Needs suitable engine connector/auth Azure SDK for control files

For native Python choose a supported Python 3.11+ kernel, install datacoolie[polars-delta], and use fabric/run_polars.ipynb (source · raw). Rebuild the Fabric overlay with both business roots under /lakehouse/default/Files/datacoolie-example/data/{input,output}. Use that mount prefix for metadata/log/watermark parameters too. Put sizing before execution:

%%configure -f
{"vCores": 4}

This is Fabric cell syntax. See Python kernels and sizing.

For external execution use fabric/run_polars_azure_sdk.py (source · raw); run --help for its required qualified cloud paths and CLI parameters. It selects runtime="external" and uses DefaultAzureCredential. Azure SDK credentials cover platform metadata/log/secrets; they do not automatically configure Polars CSV/Delta storage auth. Provide engine storage options or directly accessible business paths before running.

Optional secrets and Gold tables

No secrets are required by the fixture. For a secret-backed connection:

{"configure": {"password": "sql-password"}, "secrets_ref": {"https://myvault.vault.azure.net/": ["password"]}}

The executing identity needs Key Vault secret Get separately from lakehouse access. Native mode uses NotebookUtils; external mode uses Azure SecretClient. See credentials.

For later Direct Lake Gold workloads, explicitly choose spark.conf.set("spark.sql.parquet.vorder.default", "true") before writing when read benefits justify write costs. Review V-Order and Direct Lake storage. The smoke output alone does not create a semantic model.

Troubleshooting and next steps

Symptom Check
Python cannot find Files/... Attached mount and rebuilt business roots
Pipeline denied after interactive success Pipeline connection identity and cross-workspace access
Empty selection Built metadata root, stage, active flag and shard parameters
SDK metadata works, engine fails Business connector credentials and path support

Continue with operations for replay, maintenance, sharding and functions. The larger Fabric simulator and WWI walkthrough require their own input/dependency preparation.