Dataflow examples¶
Each focused dataflow demonstrates one primary behavior. The framework accepts
inline SQL or a file reference; preparation resolves the reference before the
reader is built, while metadata logging keeps the original source.query
value.
The extracted project recipes require Python 3.11+. Their pinned 0.2.0
profiles and the matching source-wheel handoff are documented in the
installation guide.
When using that preview/source-wheel handoff, apply the profile named by each
recipe to the matching wheel for a minimal install. Artifact additionally
needs polars-sql (and its SQLGlot dependency); the guide's generic
cli,polars-delta profile already supplies Polars for Function and Transform.
Artifact project fixture¶
The fixture below is intentionally small:
artifact/
├── datacoolie.yml
├── metadata/
│ ├── connections.json
│ ├── schema_hints.json
│ └── dataflows/orders_query.json
└── queries/orders.sql
Open Artifact SQL project (project-files · download). The dataflow is metadata/dataflows/orders_query.json (source · raw), and the companion query is queries/orders.sql (source · raw).
SQL file resolution¶
The metadata uses artifact:/queries/orders.sql, so the SQL file is resolved
relative to the artifact root. A relative reference such as
queries/orders.sql can instead be resolved through one or more explicit SQL
roots. The framework does not reserve a fixed sql/ folder.
The project runner registers orders and order_categories as Polars
relations before calling the Driver. The SQL file joins the two qualified
relations. This is the important boundary for qualified SQL: table
registration is project/engine code, while query-file resolution is framework
preparation.
Artifact project recipe¶
For a complete extracted run, use Python 3.11+ with the qualified-SQL profile,
then change to the extracted artifact/ root (the directory
containing metadata/, queries/ and runners/):
python -m pip install "datacoolie[polars-sql]==0.2.0"
python runners/dev/run.py --state-base-path .runtime
The project runner reports one succeeded dataflow and writes
data/output/orders/orders.parquet plus runtime records under .runtime/.
Read the business result independently:
import polars as pl
rows = (
pl.read_parquet("data/output/orders/orders.parquet")
.select("order_id", "amount", "category_group")
.sort("order_id")
)
assert rows.rows() == [
(1, 19.99, "physical"),
(2, 29.0, "digital"),
(3, 5.5, "physical"),
]
The destination uses overwrite, so rerunning the command keeps three business
rows. To adapt the SQL, change queries/orders.sql and register every new
relation in runners/dev/run.py; changing the reference CSV in data/input/
alone does not change this in-memory relation example. To reset without a
destructive cleanup command, extract a fresh ZIP copy into another directory;
generated output and .runtime/ belong to each extracted project.
Function project¶
Function project recipe¶
Open Function project (project-files ·
download). Its metadata is
metadata/dataflows/orders_function.json
(source ·
raw) and its
function package is functions/sources.py
(source ·
raw).
The python_function path is metadata, while packaging/import setup remains
project runner code. Extract the ZIP and change to its function/ root:
python -m pip install "datacoolie[polars]==0.2.0"
python runners/dev/run.py --state-base-path .runtime
functions/sources.py returns three rows without external input. The runner
keeps the functions prefix explicit and writes
data/output/orders/orders.parquet:
import polars as pl
rows = (
pl.read_parquet("data/output/orders/orders.parquet")
.select("order_id", "amount", "category")
.sort("order_id")
)
assert rows.rows() == [
(1, 19.99, "hardware"),
(2, 29.0, "software"),
(3, 5.5, "hardware"),
]
Rerunning keeps the three-row overwrite result. To adapt the function, edit
functions/sources.py and the python_function value in the dataflow; keep
functions/__init__.py and install any function dependencies in the execution
environment. For the automatic packaging check, add the cli extra and run
python -m pip install "datacoolie[cli,polars]==0.2.0", then
dc --project-dir . build --format json from the extracted root. A fresh
extraction resets generated output and .runtime/.
The archive also includes functions/range_source.py, a separate
custom-reader extension fixture.
The CLI packages it with the function root, but this project's metadata selects
functions.sources.load_orders; the normal Function recipe does not execute
that custom reader.
One focused transform project¶
The Transform project (project-files ·
download) keeps one dataflow per feature: it
trims and lowercases category, projects the three business columns, and
overwrites a Parquet destination. Its runner creates a
tiny synthetic CSV only when the input is absent, so both the project archive and a
CLI-built environment run without a checked-in runtime directory.
Transform project recipe¶
Extract the ZIP and change to its transform/ root. Install the Polars
profile, then run the project-owned runner:
python -m pip install "datacoolie[polars]==0.2.0"
python runners/dev/run.py --state-base-path .runtime
The runner creates the three-row CSV when needed. The output keeps exactly the
three business columns and normalizes category:
import polars as pl
rows = pl.read_parquet("data/output/orders/orders.parquet").sort("order_id")
assert rows.columns[:3] == ["order_id", "category", "amount"]
assert rows.select("category").to_series().to_list() == [
"hardware",
"software",
"hardware",
]
Rerunning the overwrite flow keeps three rows. To adapt it, edit the source
fixture or metadata/dataflows/orders_clean.json and keep select_columns
aligned with the columns produced by the value rules. Extract a fresh copy to
reset the generated data/output/ and .runtime/ directories.
Focused metadata references¶
The catalog owns the complete metadata inventory. The focused source files are inline_sql.json (source · raw), sql_file.json (source · raw), format_connections.json (source · raw), transform_patterns.json (source · raw) and load_strategies.json (source · raw).
These compact files are authoring references rather than one combined runnable pipeline. Keep each production dataflow focused on one behavior, then validate the complete project with the CLI before a runner executes it.
Validate and build without execution¶
dc --project-dir docs/examples/files/projects/artifact validate --only config --only metadata --format json
dc --project-dir docs/examples/files/projects/artifact build --dry-run --format json
These commands validate the project and show the build digest. They do not run a Driver, read business data or write runtime state.
Authoring patterns¶
| Pattern | Use when | Reference |
|---|---|---|
| Inline SQL | The query is short and belongs with the dataflow metadata | Source patterns |
| SQL file | The query is medium/complex, reviewed as code or shared by a project | Query preparation guide |
| Qualified Polars SQL | A runner registers tables before calling the Driver | Polars SQL guide |
| Python function source | Business logic needs a packaged callable | Python functions contract |
Function packages are prepared before Driver startup. The dataflow metadata declares the import path; the runner only supplies the fixed allowed prefix and does not discover or package source files at runtime.