Skip to content

Configuration examples

Configuration belongs at the boundary that owns it. A runner constructs the engine, platform, provider and Driver session; metadata preparation resolves queries and secrets before a reader is created.

The standalone provider fixture below is a local contract check. It requires Python 3.11+ and the pinned provider profiles datacoolie[polars,metadata-db,source-api]==0.2.0: Polars executes the optional dataflow, metadata-db supplies SQLAlchemy for SQLite metadata, and source-api supplies the HTTP client used by the loopback API provider. See provider installation and the versioned installation handoff before running it. If the handoff uses a preview/source wheel, apply [polars,metadata-db,source-api] to that matching wheel; the installation guide's generic cli,polars-delta profile supplies Polars but does not include the SQLAlchemy and HTTPX dependencies required by this fixture. The other files on this page are snippets or authoring references; they do not seed a production metadata service.

The three common path modes

Mode Driver input When to use
Artifact artifact_base_path=".../current/dev" A CLI-built environment contains metadata/; SQL may be artifact-relative.
Explicit file metadata metadata_base_path=".../metadata" or a FileProvider A standalone metadata folder is owned by the caller.
Non-file metadata DatabaseProvider or APIProvider plus optional SQL roots Metadata is remote/provider-owned; file-only watermark roots are not inferred for it.

The artifact and explicit roots implementation is configuration/provider_construction.py (source · raw). The standalone FileProvider variant is configuration/standalone_file_provider.py (source · raw). The runtime configuration guide is the normative explanation of fallback and conflict rules.

A minimal explicit provider

from datacoolie.engines.polars_engine import PolarsEngine
from datacoolie.metadata.file_provider import FileProvider
from datacoolie.orchestration.driver import DataCoolieDriver
from datacoolie.platforms.local_platform import LocalPlatform

platform = LocalPlatform()
engine = PolarsEngine(platform=platform)
metadata = FileProvider(
    metadata_base_path="./metadata",
    platform=platform,
    sql_base_path=["./sql"],
)

metadata.initialize()
connections = metadata.get_connections(active_only=False)
dataflows = metadata.get_dataflows(active_only=False, attach_schema_hints=False)
try:
    with DataCoolieDriver(
        engine=engine,
        metadata_provider=metadata,
        state_base_path="./.runtime",
    ) as driver:
        result = driver.run(stage="bronze2silver")
finally:
    # The provider was injected, so the caller owns its lifecycle.
    metadata.close()

Constructing FileProvider without a platform is also supported when no I/O occurs during construction; the Driver startup must bind a platform before metadata is initialized. An injected provider remains caller-owned.

Local Database and API contract fixture

configuration/provider_fixtures.py (source · raw) supplies a small orders dataflow from an in-memory SQLite database or a temporary loopback HTTP service. It is a local contract check for provider startup and Driver handoff, not an application seeder or service template.

Run it from the DataCoolie package checkout directory that contains docs/. The default command only hydrates metadata. The opt-in execution creates a new work directory, writes a three-row Parquet input, and runs one bronze2silver stage through both providers:

python -m pip install "datacoolie[polars,metadata-db,source-api]==0.2.0"
python docs/examples/files/configuration/provider_fixtures.py --provider both
python docs/examples/files/configuration/provider_fixtures.py --provider both --run-dataflow --work-dir .scratch/provider-orders-both

Choose a work directory that does not exist yet; this command will not reuse an existing path. The SQLite and API examples use the same local input and output format, so the metadata backend is the meaningful change. If the checkout is outside the framework repository, copy or download this standalone fixture and install the same package profiles before running it; it is not an extracted project archive.

The command reports connections=2 dataflows=1 for each provider and exits non-zero if the local stage does not succeed. It validates provider metadata and the Driver handoff. With --run-dataflow, it reports executed=1 connections=2 dataflows=1 for both providers and writes each provider's input, output and state below the new work directory. It does not prove permissions or connectivity to a production source or destination. To repeat the check from a clean state, choose another new work directory rather than removing the first one.

Run attributes and logging

DataCoolieRunConfig.run_attributes is a strict JSON object for external correlation values such as a Data Factory pipeline run ID or a Glue job ID. It is copied at construction and is recorded with the job runtime; it is not a replacement for job_id, sharding fields or framework status.

from datacoolie.core.models.run_config import DataCoolieRunConfig
from datacoolie.logging import LogConfig

config = DataCoolieRunConfig(
    run_attributes={"factory_pipeline_run_id": "pipeline-2026-09-15-001"},
)
log_config = LogConfig(
    output_path="./.runtime/logs",
    persistence_mode="snapshot",
)

Logging modes

The complete logging example is configuration/logging_modes.py (source · raw). Snapshot mode replaces the stable projection; batch mode emits JSON Lines records in .json files.

Multiple SQL roots

The SQL-roots sample is configuration/sql_roots.py (source · raw). It shows how a relative query reference is resolved through more than one explicit root. The framework does not reserve a fixed sql/ folder; provider roots own metadata-associated SQL files, while Driver roots remain the session fallback.

Persisted records keep the original metadata query reference while runtime records may include the actual SQL sent to the reader in source_action["query"]. Console formatting is a presentation choice; it does not change the structured log payload.

Configuration source files

The catalog remains the discovery inventory. When the configuration sample is known, use configuration/provider_construction.py (source · raw). Project-owned configuration belongs to the project's project-files section in the catalog; its complete checkout is the project's single download action (a .zip archive).

Use CLI project workflow for authoring and building datacoolie.yml; do not copy Driver session settings into that project contract.