Skip to content

Metadata guide for new users

If you are a new Data Engineer or Data Analyst who just installed DataCoolie and is not sure how to configure your first pipeline, start here.

DataCoolie is driven entirely by metadata — a JSON (or YAML, or Excel) document that tells the framework what to read, how to transform it, where to write it, and when to re-run incrementally. You do not write Python for each pipeline; you fill in a structured document.

This guide walks through that document from zero, but it also covers the cases that usually appear right after the first successful run: incremental loads, query-based sources, API pagination, function sources, partitioning, merge strategies, secrets, and metadata-provider differences.

Use the Metadata reference when you need the complete authored field contract. The topic pages below explain how fields work together; each case keeps its JSON configuration beside the explanation. The runnable projects under Examples are separate projects with runners, SQL and fixtures for cases that need an executable setup.

Choose your path

New to DataCoolie?

Start with Build your first metadata file. It creates a small JSON document, validates it, and runs the first dataflow.

Authoring or reviewing production metadata?

Follow the authoring workflow: begin with reusable connections, compose dataflows, then configure the source, transform, and destination phases.

The first path is optimized for a quick successful run. The second is the conceptual route for understanding how a larger metadata document is organized.

Configure features and solve combined cases

Configure metadata teaches the supported fields and conditions of each component, including OAuth authentication, pagination, look-back, partitioning and SCD2. Start with the page that owns the component. Advanced contains worked situations that combine those already explained components: complete window replacement, incremental paginated APIs, late files, protected keyed outputs, and incremental SCD2. It is a task route, not a mandatory step.

The generated Metadata reference provides exact field types, values and defaults when a guide case leaves you needing the precise authored shape.

Metadata document

Start with one authored document containing connections and dataflows. The optional root schema_hints groups reusable column types across matching sources. The Metadata document reference lists the exact root fields; the pages for Connections, Dataflows, and Datatypes and schema hints explain how each section is configured.

This is a root-level fragment with no configured flows yet:

{
  "$schema": "https://datacoolie.github.io/datacoolie/schema/latest/metadata.schema.json",
  "connections": [],
  "dataflows": [],
  "schema_hints": [],
  "extensions": {"owner": "analytics"}
}

Choose a schema marker

The optional $schema helps editors and validators find the authored contract. Use the published latest alias for current authoring, or pin a compatible versioned URL when an artifact must remain reproducible. The installed CLI validates with its bundled compatible schema; it does not fetch the URL. The Driver does not download it at runtime.

Keep names readable

Connections always need a nonblank name. Keep that name unique in the active name-lookup scope; distinct explicit connection IDs may share a display name when every operation uses the ID. Dataflows normally use a unique name, but an explicit dataflow_id can be used without a name for ID-based integrations. For ordinary authoring, refer to a connection using source.connection_name, destination.connection_name, and root schema_hints[].connection_name. DataCoolie derives connection and dataflow IDs when they are absent. Use an explicit stable ID only when another system owns that identity; when a workspace_id is present, name scope is that workspace. See Connection identity and Dataflow identity.

Add project-owned extensions

The root extensions object can hold project annotations, for example owner, data product, and ticket identifiers:

{
  "extensions": {
    "owner": "analytics",
    "data_product": "orders",
    "change_ticket": "DATA-1234"
  }
}

DataCoolie core does not interpret those keys. An external runner or project extension may read them; they do not configure connections, change loads, or bypass validation. Put operational settings in their documented fields.

Prepare an environment variant

An environment overlay is applied to a common metadata snapshot during project preparation, before validation. It is not a connection or dataflow field and the Driver does not read overlay files. Follow the Environment overlays procedure, then validate the effective document. The overlay cannot change identity fields or delete existing definitions.

Authoring workflow

Recommended workflow

Start with the reusable connections, compose a dataflow, then follow its runtime phases from source through transform to destination. This is a navigation path, not a requirement to read every advanced page before your first run. Validate before executing and repeat the check after changes.

Order Concern What you decide
1 Metadata document Root structure, schema marker and extensions
2 Connections Reusable endpoints, formats, defaults, workspace scope, secrets and authentication
3 Dataflows Pipeline identity, scheduling envelope, and phase boundaries
4 Source Table, SQL/query, file, API, function, filter, pagination, watermark and look-back behavior
5 Transform Cast, normalize, deduplicate, hash, mask, project, and compute columns
6 Destination Target, load_type, keys, partitions, and write behavior
7 Datatypes and schema hints Inline or shared type hints, source dialects, decimals, and timestamps
8 Validation checklist Provider, path, secret, cross-field, and pre-run checks

The order follows the authored document and then the dataflow phases. Destination strategy can constrain transform choices, so revisit it before finalizing a complex transform. For combinations, choose window replacement, paginated incremental API, late files, stable protected keys, or incremental SCD2.

Coverage map

Area Configure it here Built-in cases
Metadata shape Metadata document, dataflows connections[], dataflows[], shared schema_hints[], orchestration fields
Metadata backends Provider chooser JSON, YAML, Excel, database provider, API provider
Connections and secrets Connections File, lakehouse, database, API, function; configure, secrets_ref
Source types Source File, Delta, Iceberg, database table/query/SQL file, REST API, Python function
Destination and load Destination patterns File outputs, Delta, Iceberg; append, overwrite, full_load, merge, SCD2, window replacement
Transform and partition Transform patterns, destination partitioning Normalize, hash, deduplicate, compute, mask, project, partition, system columns
Datatypes and hints Datatypes and schema hints Inline and shared hints, source types, timestamp and decimal semantics
Incremental behavior Source Watermarks, look-back, closing day, API ranges, replay boundaries
Identity and preparation Metadata document, Connections, Dataflows, Project overlays $schema, derived/explicit IDs, project-owned extensions, environment overlays
Validation and safety Validation checklist Secrets, cross-field checks, smoke tests, common errors

For every area, the Metadata reference is the field lookup: type, schema enum, default, and authored structure. The linked guide page explains which fields must be combined for a working case.

Important edge cases

This guide covers the real behavior of the current framework, including:

  • connection_type can be derived automatically from format
  • Excel is a supported source format, not a writable destination
  • flat-file destinations support append, overwrite, and full_load, but not merge or SCD2
  • connection_type: "streaming" exists in the model but has no supported formats yet

Where metadata lives

DataCoolie supports three metadata backends. Choose one:

Backend Good for Operational note Guide
JSON / YAML / Excel file Local dev, small teams, single-machine runs JSON should stay canonical; YAML/Excel are alternative views or generated siblings Configure file metadata
Relational database Shared team configuration, multi-workspace governance Rows are workspace-scoped via workspace_id Configure database metadata
REST API Enterprise ops, Git-backed or approval-gated config Endpoints are workspace-scoped under /workspaces/{workspace_id}/... Configure API metadata

Recommendation for beginners

Start with a JSON file. The file backend requires no database and no service — just create a .json file and point FileProvider at it. You can migrate to the database or API backend later while keeping the same logical metadata model. Verify provider round-trips before rollout: database/API storage has its own schema and payload contract, and API source.filter_expression is still evaluated locally after the DataFrame is created.

What metadata tells the framework

metadata.json
├── connections[]            ← WHERE to read from and write to
│   ├── name / format / configure
│   ├── catalog / database / base_path
│   └── secrets_ref / is_active / workspace_id
└── dataflows[]              ← HOW to move data
  ├── name / stage / description
  ├── group_number / execution_order / processing_mode / is_active
  ├── source               ← which connection + table/query/function to read
  ├── destination          ← which connection + table + load_type to write
  └── transform            ← normalization, schema, hashing, masking, projection (optional)

Start with connections and dataflows. The rest is optional and you can add it incrementally.

If you are unsure where a field belongs, use this rule:

  • Connection.configure = reusable endpoint defaults
  • source.configure / destination.configure = per-dataflow overrides
  • transform.configure = transformer behavior flags

Quick example (30 seconds)

{
  "connections": [
    {
      "name": "csv_input",
      "connection_type": "file",
      "format": "csv",
      "configure": { "base_path": "data/input" }
    },
    {
      "name": "bronze",
      "format": "delta",
      "configure": { "base_path": "data/output/bronze" }
    }
  ],
  "dataflows": [
    {
      "name": "orders_to_bronze",
      "stage": "ingest",
      "source":      { "connection_name": "csv_input", "table": "orders" },
      "destination": { "connection_name": "bronze", "schema_name": "sales", "table": "orders", "load_type": "append" }
    }
  ]
}

This reads data/input/orders (a folder of CSV files) and appends to a Delta table at data/output/bronze/sales/orders. In this example DataCoolie derives connection_type: "lakehouse" from format: "delta".

→ Choose a path: Build your first metadata file · Connections · Dataflows