Skip to content

Platforms and deployment

Platform pages explain the host-specific decisions around engines, storage paths, secrets, notebooks, jobs, and deployment packaging. They complement the portable runtime and operations guide, which owns the runner lifecycle.

Platform Start here when Guide
Microsoft Fabric You run in OneLake notebooks and need to choose Spark or Polars by layer Microsoft Fabric
Databricks You use managed Spark, Unity Catalog, or UC Volumes Databricks
AWS Glue You run managed Spark ETL with S3, Glue Catalog, or Secrets Manager AWS Glue

Start with the platform smoke download/build/local rehearsal, then follow one leaf guide to upload input and built metadata, choose the execution identity and pass host parameters. The fixture's overwrite output must be a fresh sandbox. Generic job success is insufficient: check required selection, result counts, exact rows/types and any downstream catalog entry.

Coverage and verification boundary

Case Route and boundary
Native Spark + file metadata First-run recipe in all three guides; host adapters tested locally
Native Python/Polars Fabric mount and Databricks file-only branches; single-node and connector limits
External SDK + Polars Canonical Azure, Databricks SDK and AWS S3 runners; control auth differs from engine auth
Named table vs path files Databricks UC alternatives; Glue Iceberg catalog vs S3 Delta; Fabric Files smoke
Secrets/database metadata Optional per-host secrets; provider setup owns bootstrap
Dependencies/runtime Host Spark reused; classic/serverless and bundled/custom Glue recipes separated
Scheduler and failure Parameter tables, required-selection guards, output/log checks; generic runners allow valid empty shards
Functions/replay/maintenance/sharding Operations and canonical runner catalog

Local evidence includes metadata validation/build, downloaded-project execution and exact output checks, simulated host adapters, and generated source/raw/ZIP pages. It does not qualify live cloud access or every serverless/Spark Connect operation. A branch not executed on its cloud host is not thereby unsupported. Explicit exclusions include Glue Python Shell's incompatible Python version, UC table registration on Volume files, and portable Workspace Files support. EMR provisioning and Jobs API payloads have no repo-owned recipe.

Keep runtime configuration explicit. For larger multi-stage scenarios, use the linked simulator/WWI handoffs after preparing their additional assets and dependencies.