Configure API metadata¶
Prerequisites · pip install "datacoolie[polars,source-api]" for the local
loopback and Parquet example. For your own service, use an HTTPS metadata API
that implements the HTTP contract below. The source-api extra installs
httpx for the DataCoolie client.
End state · APIProvider reads a complete workspace scope, exposes the
same metadata models as the other providers, and reads/writes watermarks over
HTTP.
Try a local orders dataflow¶
From a DataCoolie repository checkout, run the maintained
loopback example. It
starts a temporary HTTP metadata service on 127.0.0.1, returns two
connections and one dataflow, and runs the same small local Parquet flow used
by the Database provider example:
python docs/examples/files/configuration/provider_fixtures.py --provider api --run-dataflow --work-dir ../.scratch/provider-api-orders
Choose a --work-dir path that does not exist yet. Expect
provider=api executed=1 connections=2 dataflows=1 and an output file at
<work-dir>/api/output/orders/orders.parquet. This local service is a contract
fixture: the dataflow reads and writes local files, while APIProvider gets
its metadata over HTTP. The example does not set up a production service.
Expected endpoints¶
All paths are scoped under /workspaces/{workspace_id}/.
| Method | Path | Purpose |
|---|---|---|
GET |
/workspaces/{workspace_id}/connections |
Paginated connection listing. Supports name and active_only. |
GET |
/workspaces/{workspace_id}/connections/{connection_id} |
One connection; 404 means it was not found. |
GET |
/workspaces/{workspace_id}/dataflows?stage=X&active_only=true |
Paginated dataflows, optionally filtered by stage and activity. |
GET |
/workspaces/{workspace_id}/dataflows/{dataflow_id} |
One dataflow with expanded source and destination connections. |
GET |
/workspaces/{workspace_id}/schema-hints |
Paginated hints; supports connection/table/schema filters. |
GET |
/workspaces/{workspace_id}/watermarks/{dataflow_id} |
Watermark envelope for one dataflow. |
PUT |
/workspaces/{workspace_id}/watermarks/{dataflow_id} |
Replace the current watermark from a JSON body. |
The client sends the API key in the X-API-Key header. Collection endpoints
should return an object with a list under data and a pagination object. The
client accepts an omitted data as an empty list and an omitted pagination
as one page, but a service should send the explicit envelope so other clients
can inspect the response consistently:
Response fields inside data follow the metadata models defined in the
focused modules under datacoolie.core.models — see
Reference · Metadata document.
Dataflow response contract¶
Each dataflow response must include expanded source.connection and
destination.connection objects. This response describes the same local
orders flow used by the loopback example above:
{
"dataflow_id": "orders-dataflow",
"workspace_id": "example-workspace",
"name": "orders",
"stage": "bronze2silver",
"source": {
"connection": {
"connection_id": "source-connection",
"workspace_id": "example-workspace",
"name": "source",
"connection_type": "file",
"format": "parquet",
"configure": {"base_path": "./input"}
},
"table": "orders"
},
"destination": {
"connection": {
"connection_id": "destination-connection",
"workspace_id": "example-workspace",
"name": "destination",
"connection_type": "file",
"format": "parquet",
"configure": {"base_path": "./output"}
},
"table": "orders",
"load_type": "overwrite"
},
"transform": {}
}
The nested connection objects must include at least connection_id, name,
connection_type, and format; the client validates those fields while it
builds the runtime model. A dataflow response must include both expanded
connection objects so startup can validate references without guessing another
API call.
For a source with a row predicate, put filter_expression directly in the
source object. The API provider preserves it for the source reader; it does
not turn it into HTTP request parameters. Use source.configure.params or
source.configure.body when a data source requires endpoint push-down.
Send destination.load_type explicitly. The API mapper currently falls back to
overwrite when the field is missing, while authored metadata models use
append as their default. Pinning the value avoids an accidental load strategy
when a service omits a field.
Watermark requests use a separate envelope. A GET response for a dataflow
that has no stored watermark is:
A PUT request sends the serialized JSON watermark string together with the
optional run identifiers:
{
"current_value": "{\"updated_at\": \"2026-09-28T00:00:00+00:00\"}",
"job_id": "optional-external-or-framework-job-id",
"dataflow_run_id": "optional-dataflow-run-id"
}
The Python method get_watermark() returns the raw serialized value or
None; that Python return is not the HTTP response shape. A valid 404 for a
watermark also means that no watermark has been created. Authentication,
transport, server, and malformed-response failures remain errors.
Loading¶
import os
from datacoolie.metadata.api_provider import APIProvider
provider = APIProvider(
base_url=os.environ["DATACOOLIE_METADATA_API_URL"],
api_key=os.environ["DATACOOLIE_METADATA_API_KEY"],
workspace_id="your-workspace-id",
enable_cache=True,
timeout=10.0,
sql_base_path="./sql",
)
Construction is lazy. Validate the complete workspace before handing the provider to a Driver:
provider.initialize()
connections = provider.get_connections(active_only=False)
dataflows = provider.get_dataflows(
stage="bronze2silver",
active_only=False,
attach_schema_hints=False,
)
print(f"connections={len(connections)} dataflows={len(dataflows)}")
Use the provider configuration example
for the next step: construct an engine and platform, pass this provider to
DataCoolieDriver, inspect the ExecutionResult, then close the injected
provider in the caller's finally block. Driver closes only providers it
created itself.
sql_base_path is configuration carried by the metadata provider, even though
the API provider never reads local SQL files. Driver preparation resolves a
relative source.query through its execution platform. If the provider omits
the value, the Driver-level SQL root is the session fallback; two different
explicit values are rejected during startup.
Caching¶
Set enable_cache=True (the default) to avoid re-fetching connections,
dataflows, and schema hints on every call. Call provider.clear_cache() after a
known metadata change to invalidate the snapshot; the next read initializes the
provider again. Watermark reads/writes go directly to the API; a successful
watermark update does not clear the metadata cache. APIProvider.close() closes
the HTTP client created by the provider and is safe to call once the Driver has
finished using it.
Common failures¶
| Symptom | Meaning | Fix |
|---|---|---|
API paginated response must be an object |
A collection endpoint returned a top-level array | Return an object with data as a list and optional pagination object. |
Dataflow response is missing a source or destination connection |
A dataflow omitted an expanded connection | Include both nested connection objects and their identity/type fields. |
Invalid watermark response |
The watermark body is not an object with current_value |
Return {"current_value": null} for an uninitialized watermark or a serialized JSON value. |
HTTP 401/403/transport error |
Authentication or service connectivity failed | Check the X-API-Key, HTTPS URL, timeout and service logs; do not treat it as an empty metadata scope. |
Never place an API key in committed metadata, examples or logs. Use an environment variable or the host's secret mechanism and use HTTPS outside a loopback fixture.
Larger integration testbed¶
The repository's
pg_api_metadata_server.py
is a Flask service backed by the usecase-sim PostgreSQL metadata tables. Use
it to validate a larger set of repository scenarios after the local example.
It requires its own Flask, database, schema, and seeded rows; the
datacoolie[source-api] client extra does not install or configure that
service. An application service still owns authentication, authorization,
migrations, and deployment.
Related¶
- Metadata guide for new users — understand the metadata shape (connections, dataflows, sources, destinations) before configuring a backend
- Concepts · Metadata providers · API provider
reference/api/metadata