Skip to content

Deploy to AWS Glue

Start with Spark + FileProvider and the shared platform smoke project. The baseline writes three rows to a new S3 Delta sandbox without Athena/catalog registration. These are locally checked setup contracts, not a live Glue receipt.

1. Choose the runtime and package

Runtime Scope
Glue 5.0 Spark ETL Chosen example baseline: Python 3.11, Spark 3.5.4, bundled Delta 3.3.0
Python 3.11+ container/VM/local Polars + S3 through AWSPlatform; separate engine installation
EMR / EMR Serverless Spark Runtime/packaging must be adapted; no repo-backed deployment recipe here
Glue Python Shell (Python 3.9) Incompatible with DataCoolie's Python >=3.11,<4.0 requirement

Glue 5.0 is a selected baseline, not the latest/default runtime. Do not infer support for another Glue version from Python compatibility alone. Review Glue releases and job types.

Attach the matching datacoolie[aws] package to the job. For an available published release, the job argument value is:

--additional-python-modules = datacoolie[aws]==0.2.0

Use the matching wheel plus resolved dependencies for an unpublished checkout. Freeze compatible Python wheels, including boto3/botocore, rather than resolving unpinned packages on every production startup. Glue supplies Spark; do not add a profile that installs another PySpark runtime. pyiceberg is for PyIceberg operations, not a substitute for Spark Iceberg JARs. Glue Python libraries.

2. Prepare IAM and S3

Use a Glue execution role trusted by Glue, with CloudWatch logging and access to the script, wheels, temporary directories and fixture storage. The role needs bucket listing and object reads for metadata/input and Delta transaction logs; output/log/watermark roots also need writes and deletes as their operations require. Narrow policies to the chosen prefixes. With SSE-KMS, reads require Decrypt and writes GenerateDataKey (multipart operations can require both), plus a permitting key policy. See Glue role setup, minimum job access and S3 KMS permissions.

The baseline needs no Secrets Manager, Athena or Glue table-registration grants. Spark S3 access and AWSPlatform's boto3 access must both work for the execution role; successful control-file access alone does not prove engine access.

3. Build and upload the Delta fixture

Follow the shared download/build recipe. Edit both S3 roots in metadata/environments/aws.json, validate and inspect --env aws, then build. Example prefix: s3://your-bucket/datacoolie-example.

Upload .builds/current/aws/metadata/ to <root>/metadata/. Upload input CSV separately to <root>/data/input/orders/orders.csv. Output will be <root>/data/output/orders_platform_smoke; reserve it for this overwrite test. Upload canonical aws/run_glue_spark.py as the Glue job script (source · raw). The script reuses GlueContext's session and uses DataCoolie state, not Glue bookmarks.

4. Set bundled Delta session and job arguments

Configure these before the Spark session starts. In Glue job parameters, each row is a key and its complete value; repeated --conf segments belong inside one --conf value:

Key Value
--datalake-formats delta
--conf spark.sql.extensions=io.delta.sql.DeltaSparkSessionExtension --conf spark.sql.catalog.spark_catalog=org.apache.spark.sql.delta.catalog.DeltaCatalog --conf spark.delta.logStore.class=org.apache.spark.sql.delta.storage.S3SingleDriverLogStore
--REGION Your bucket/job region, e.g. us-east-1
--METADATA_PATH <root>/metadata/metadata.json
--WATERMARK_BASE_PATH <root>/.runtime/watermarks
--LOG_BASE_PATH <root>/.runtime/logs
--STAGE platform_smoke
--JOB_NUM, --JOB_INDEX 1, 0 respectively

REGION, metadata, watermark and log arguments are required by the script. Omit optional --CONNECTIONS_PATH/--SCHEMA_HINTS_PATH; built metadata embeds both. These uppercase script argument names differ from Glue's own lowercase arguments. Apply shared smoke selection/result guards at driver.run if this job gates downstream work; generic runner failure raises a job error. The complete bundled Delta recipe comes from AWS Delta setup.

Custom Delta version alternative

Use this branch only when deliberately replacing the bundled runtime: omit delta from --datalake-formats, provide matching Delta JARs through --extra-jars, and set --user-jars-first=true on Glue 5.0+. Provide the matching Python API through --extra-py-files using the artifact layout documented by AWS (their Delta JAR includes the Python library), and retain the Delta extension, catalog and S3 log-store settings. Match Spark/Scala/JAR/Python versions as one set. Python pip install delta-spark alone does not provision this custom Glue JVM runtime. Do not combine bundled and custom Delta JARs.

5. Verify Delta output and logs

Require selected name orders_platform_smoke and counts total=1, succeeded=1, failed=0, pending=0. Independently read output using a compatible Spark session, or place this verification after the run in the job:

output = spark.read.format("delta").load(
    "s3://your-bucket/datacoolie-example/data/output/orders_platform_smoke"
).select("order_id", "customer_id", "amount").orderBy("order_id")
assert [tuple(row) for row in output.collect()] == [(1, 100, 20), (2, 100, 43), (3, 101, 7)]
assert output.dtypes == [("order_id", "bigint"), ("customer_id", "bigint"), ("amount", "bigint")]

Inspect DataCoolie execution/system logs under LOG_BASE_PATH and Glue's CloudWatch error logs. Job success alone proves neither intended selection nor Athena queryability.

Bundled Iceberg alternative

Use environment aws-iceberg, replace the input bucket and output Glue database, validate/inspect/build, and upload its metadata instead. Create the sandbox Glue database and grant the role the required Glue catalog access; account for Lake Formation permissions if that database is governed by it. Keep output format="iceberg", catalog="glue_catalog", database="datacoolie_example", no schema_name, and empty base_path. This resolves a named Iceberg table.

Replace the Delta session settings with --datalake-formats=iceberg and this single --conf value (substitute the warehouse bucket/prefix):

spark.sql.extensions=org.apache.iceberg.spark.extensions.IcebergSparkSessionExtensions --conf spark.sql.catalog.glue_catalog=org.apache.iceberg.spark.SparkCatalog --conf spark.sql.catalog.glue_catalog.warehouse=s3://your-bucket/datacoolie-example/warehouse --conf spark.sql.catalog.glue_catalog.catalog-impl=org.apache.iceberg.aws.glue.GlueCatalog --conf spark.sql.catalog.glue_catalog.io-impl=org.apache.iceberg.aws.s3.S3FileIO

The remaining script parameters are unchanged. Verify with spark.table("glue_catalog.datacoolie_example.orders_platform_smoke") and the same three business rows/types. The warehouse is session configuration, not a Volume-style connection path. For custom Iceberg, omit bundled iceberg, supply matching JARs and use --user-jars-first=true on Glue 5.0+; qualify that version set separately. See AWS Iceberg setup.

Optional catalog handoff and secrets

For AWS path-addressed Delta, athena_output_location opts into native Delta catalog registration via Athena; register_symlink_table requests the legacy symlink route and implies manifest generation. Registration is conditional: unchanged-schema writes can skip it, and registration exceptions are logged as warnings. Successful data writes/job status therefore do not guarantee a new or refreshed catalog entry. If downstream requires Athena, independently check the expected Glue database/table, location and an Athena query result.

Grant optional Athena StartQueryExecution/GetQueryExecution in the chosen workgroup, result-bucket access and Glue database/table permissions needed by that DDL; these are additional to the smoke baseline. Workgroup policies, Glue resource access.

Secret references map configured field values to JSON keys in the secret:

{"configure": {"username": "db_user", "password": "db_pass"}, "secrets_ref": {"arn:aws:secretsmanager:us-east-1:123456789012:secret:datacoolie/rds-example": ["username", "password"]}}

The secret payload must contain db_user and db_pass. Use your actual region, ARN or exact secret name; grant GetSecretValue and, for a custom encryption key, KMS Decrypt. Secret retrieval permissions.

Polars, troubleshooting and next steps

Use aws/run_polars_s3.py in controlled Python 3.11+ (source · raw); install datacoolie[aws,polars-delta] and run --help for S3 parameters. This is not a Glue Python Shell job. AWSPlatform uses boto3 credentials for control files; this script does not copy those credentials into Polars storage options. Configure compatible ambient engine credentials or explicit storage options for CSV/Delta access, and verify format-specific auth before adding Iceberg.

Symptom Check
Delta class missing Bundled flag or complete custom JAR recipe before session creation
Iceberg catalog missing Extension, catalog, warehouse and Glue permissions
Metadata works, output denied Spark S3 role access, output reads/writes/deletes and KMS
Job succeeds, Athena cannot see table Conditional/warning-only registration and separate catalog/query check
Zero selected flows Uploaded built environment, stage and shard arguments

Continue with operations. The larger AWS simulator and WWI walkthrough need their own input/dependency preparation.