Skip to content

Troubleshooting

Start with the terminal dataflow status and message, then correlate its dataflow_run_id and session log_session_id with system diagnostics. Check the selected metadata, resolved paths and engine before changing state or rerunning writes. See logging for persistence prerequisites.

Run reports no failures but expected data is missing

Check: inspect total, succeeded, skipped and pending, compare the selected dataflow IDs with the required stage inventory, and read skip reasons. Check activation on both connections and the flow, stage/connection filters, shard assignment and source eligibility.

Action: correct selection/configuration, rerun the intended scope and validate output freshness and completeness. A legitimately empty shard differs from missing required work across the job; wait for all upstream shards before releasing dependent stages. See run checks.

Logs or the job summary are missing

Check: confirm that an ExecutionLogger exists and is activated, has a platform and output path, and received the RunConfig. Check credentials, upload warnings, dropped-record counters and whether Driver close completed.

Action: configure a writable log/state root or inject configured loggers, use the Driver context manager, and verify files after close. Business success does not guarantee every log upload succeeded. See flush and failure behavior.

"pl.sql_expr unknown function …"

Polars uses a SQL subset. Common offenders:

Doesn't work in Polars Use
current_timestamp() Literal cast, or framework-added __updated_at.
date_format(col, 'yyyy-MM-dd') CAST(col AS DATE).
year(col) EXTRACT(YEAR FROM col).

See Partition expression portability.

Row count mismatch for multi-line JSON

Check: distinguish a JSON array/document from JSON Lines. JSONL requires one complete JSON value per physical line; escaped \n inside a string is valid, but a pretty-printed record spanning lines is not JSONL. Literal unescaped newlines inside JSON strings are invalid JSON.

Action: use format: "json" for a JSON document or array, and format: "jsonl" for one-record-per-line input. Fix malformed input and compare IDs/row counts after parsing rather than assuming file-line count is record count.

Excel is_active column loads everything as inactive

Check: inspect the flow and both connection activation values in the loaded metadata. A blank value uses the default active behavior; explicit FALSE means inactive and may be deliberate.

Action: set TRUE only for flows and connections that should run. Keep intentional inactive flags. See activation and selection.

"dead lock detected" on Delta optimize

Check: inspect overlapping job invocations and resolved destination paths or catalog identities. Deduplication covers one maintenance invocation; it does not coordinate external jobs. Also check concurrent ingestion/writes and the backend conflict message.

Action: serialize conflicting operations for the same physical target and retry according to the backend's conflict policy. See deduplication scope.

Spark or a source service is unavailable

Check: use the failing operation's actual runtime coordinates: Java/Spark versions, connector packages, endpoint, catalog and credentials. Resolve network/DNS/authentication failures before investigating row selection.

Action: configure the host using the platform guides and choose a supported dependency profile from installation. Repository simulation setup belongs to testing; production runners should use their deployed services and project configuration.

Iceberg writes do not appear in the expected catalog

Check: confirm format: "iceberg", the catalog implementation and endpoint, warehouse, namespace, table identifier and storage credentials used by both writer and reader. An AWSPlatform with an S3-compatible endpoint does not by itself imply a Glue catalog. Check engine/catalog initialization and permissions for the configured REST, Glue or other supported catalog.

Action: align the writer/reader catalog and target, then query that target and inspect its snapshots. Delta symlink options generate_manifest and register_symlink_table do not register Iceberg tables.

WatermarkManager throws on first run

Check: distinguish missing stored state from a provider/storage error or invalid source configuration. A missing state is normally an initial read; it does not guarantee every source can issue an unbounded request. For example, an API using incremental range splitting may require watermark_range_start when no lower state exists.

Action: configure the source's documented first-read behavior and a valid provider-owned state root. Custom readers must handle absent state according to their contract or reject an unsupported initial read clearly. Use an explicit bounded replay for a history load when supported; do not fabricate watermark state to hide a storage failure. See API source configuration.

Maintenance succeeded but compaction or cleanup had no effect

Check: read warning logs and destination_operation_details, then inspect the selected table's history, snapshots and files. A requested backend action can be unavailable, skipped or have no eligible work. Polars Iceberg warns about unsupported compaction/orphan removal and can warn on snapshot-expiry failure without failing the aggregate.

Action: use the capability matrix to select an engine/catalog that implements the required action. Review the retention policy and target identity before any cleanup retry.

Replay wrote rows but the watermark did not advance

Replay writes the destination before saving the reader-produced watermark candidate. A process interruption or provider failure in that gap can leave committed rows with the previous watermark. The next invocation still reads every requested [start, end) chunk; the saved value is never a replay checkpoint. Rerun the complete range with a keyed or otherwise idempotent destination strategy when repeated delivery would create duplicates. An empty or all-null observation does not advance state, while a real zero is a valid watermark value.