Runtime configuration¶
Runtime configuration is distinct from authored metadata. A runner supplies these values when creating a Driver session; the Driver does not read a project manifest at runtime. Types and defaults below come from the installed source models. Semantic constraints describe current runtime behavior.
Run configuration¶
DataCoolieRunConfig accepts declared fields as keyword arguments through its model constructor. Its dataclass declaration disables the generated initializer, so an empty generated signature would be misleading. These fields are not keys in authored metadata JSON.
| Field | Python type | Default | Meaning / constraints |
|---|---|---|---|
job_id |
str |
generated per instance (generate_unique_id) |
Caller correlation identity; must be non-empty. Logging also supplies session and operation identities. |
job_num |
int |
1 |
Total shard count; at least 1. |
job_index |
int |
0 |
Zero-based shard index: 0 <= job_index < job_num. |
max_workers |
int |
8 |
Maximum concurrent dataflow workers; at least 1. |
stop_on_error |
bool |
False |
Stop admitting new dataflows after a failure; already admitted work may finish. |
retry_count |
int |
0 |
Additional attempts after a failed execution; non-negative. |
retry_delay |
float |
5.0 |
Seconds between retries; non-negative. |
dry_run |
bool |
False |
Use the Driver dry-run validation path. This still prepares runtime dependencies; it is not an offline schema-only check. |
retention_hours |
int |
168 |
Retention passed to destination maintenance; non-negative. Backend support determines the effect. |
allowed_function_prefixes |
List[str] |
new [] per instance |
Import-path prefixes for Python function sources. An empty list applies no prefix restriction. A non-empty list requires the function path to start with one of these strings; use trusted module prefixes such as my_project.sources. to constrain imports. |
run_attributes |
Optional[Dict[str, Any]] |
None |
Optional caller correlation object, copied on construction. Keys must be strings; nested values must be JSON-compatible with finite numbers and no cycles. Raw JSON strings and arbitrary objects are rejected. |
from datacoolie.core import DataCoolieRunConfig
config = DataCoolieRunConfig(
job_id="nightly-orders", max_workers=2, retry_count=1,
run_attributes={"scheduler": "local", "attempt": 1},
)
Supply this object as config= to DataCoolieDriver. The create_driver factory instead accepts run fields such as job_id, max_workers and retry_count directly; it builds the configuration object internally. See runtime setup for a full runner.
DataCoolieRunConfig
dataclass
¶
Validated execution parameters for a DataCoolie run.
Replay configuration¶
| Field | Python type | Default | Meaning / constraints |
|---|---|---|---|
start |
Any |
required | Required inclusive lower bound, not None. |
end |
Any |
required | Required exclusive upper bound, not None. |
chunk_interval |
Optional[Dict[str, int]] |
None |
None selects one bounded read. Time chunks use years/months/weeks/days/hours/minutes; integer chunks use {"step": N}. Constructor checks the mapping shape; execution validates the interval and range. |
save_watermark |
bool |
False |
Persist source-observed watermarks after successful chunks. Does not create a replay checkpoint or skip chunks on a later run. |
chunk_column |
Optional[str] |
None |
Non-empty override, or the first source watermark column. The reader must support bounded reads for the selected column. |
Replay uses [start, end). Construction does not prove reader capabilities or validate every range/interval combination; execution performs those checks. See replay and backfill.
ReplayConfig
dataclass
¶
ReplayConfig(
start: Any,
end: Any,
chunk_interval: Optional[Dict[str, int]] = None,
save_watermark: bool = False,
chunk_column: Optional[str] = None,
)
Configuration for replaying a bounded time range in chunks.
Used by :meth:DataCoolieDriver.run_replay to reprocess historical
data without corrupting the production watermark.
The range uses the left-closed, right-open [start, end) convention:
start is inclusive, end is exclusive. This aligns chunks
to whole calendar units (days, weeks, months, etc.) and is the
industry-standard interval convention used by Python’s range(),
Spark partition pruning, and PostgreSQL range types.
Example::
# Replay all of Q1 2025 in monthly chunks:
ReplayConfig(
start="2025-01-01", # inclusive
end="2025-04-01", # exclusive (first day NOT included)
chunk_interval={"months": 1},
)
# Produces chunks: [Jan 1, Feb 1), [Feb 1, Mar 1), [Mar 1, Apr 1)
The chunk column is auto-resolved from
dataflow.source.watermark_columns[0] at runtime. Override with
chunk_column for a source-supported independent bounded-read column
or when the first watermark column is not the one to chunk on.
Type detection is automatic:
strparseable to date/datetime → time-based chunkingdatetime/dateobjects → time-based chunkingint→ integer-based chunking
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
start
|
Any
|
Inclusive lower bound of the replay range. |
required |
end
|
Any
|
Exclusive upper bound of the replay range. |
required |
chunk_interval
|
Optional[Dict[str, int]]
|
Chunking interval. Time-based keys ( |
None
|
save_watermark
|
bool
|
When |
False
|
chunk_column
|
Optional[str]
|
Override the auto-resolved chunk column. Use this for
an independent bounded-read column when the selected source
reader supports it, or when the first watermark column is not the
desired chunking dimension. API readers require a matching
|
None
|
Logging configuration¶
| Field | Python type | Default | Meaning / constraints |
|---|---|---|---|
log_level |
str |
'INFO' |
Console threshold: DEBUG, INFO, WARNING, ERROR, CRITICAL. Names normalize to uppercase. |
file_level |
str |
'DEBUG' |
Persisted Python-record threshold; same supported levels as log_level. |
storage_mode |
str |
'memory' |
Local capture: memory or file; names normalize to lowercase. |
output_path |
Optional[str] |
None |
Standalone logger component root, or None for no remote writer. Must be a non-empty path when supplied. Driver log_base_path instead creates category roots. |
partition_by_date |
bool |
True |
Boolean controlling UTC date partitioning. |
partition_pattern |
str |
'__run_date={year}-{month}-{day}' |
Non-empty template with simple placeholders in ordered prefix: {year}, {month}, {day}, {hour}. Each path segment needs a placeholder; literals cannot contain digits, %, or braces. |
persistence_mode |
str |
'snapshot' |
snapshot or batch; names normalize to lowercase. See the logging guide for write layout and backend requirements. |
flush_interval_seconds |
float |
300.0 |
Finite non-negative seconds; 0 disables time-triggered flushing. Booleans are rejected. |
flush_batch_bytes |
int |
4194304 |
Positive integer byte threshold; booleans are rejected. |
buffer_memory_bytes |
int |
67108864 |
Positive integer memory buffer budget; booleans are rejected. |
spool_max_bytes |
int |
536870912 |
Positive integer spool budget, at least buffer_memory_bytes; booleans are rejected. |
spool_directory |
Optional[str] |
None |
Optional non-empty local spool path; normalized when supplied. |
close_timeout_seconds |
float |
10.0 |
Positive finite close timeout; booleans are rejected. |
console_color |
str |
'auto' |
auto, always, or never; names normalize to lowercase. Auto respects NO_COLOR and TERM=dumb before terminal detection. |
Logging capture is bounded. Buffer or spool exhaustion can lose records and is reported through logging statistics; logging persistence is not a transactional guarantee for business execution. For activation, flush, close and storage behavior, see logging and the logging API.
LogConfig
dataclass
¶
LogConfig(
log_level: str = INFO.value,
file_level: str = DEBUG.value,
storage_mode: str = MEMORY.value,
output_path: Optional[str] = None,
partition_by_date: bool = True,
partition_pattern: str = DEFAULT_PARTITION_PATTERN,
persistence_mode: str = SNAPSHOT.value,
flush_interval_seconds: float = 300.0,
flush_batch_bytes: int = 4 * 1024 * 1024,
buffer_memory_bytes: int = 64 * 1024 * 1024,
spool_max_bytes: int = 512 * 1024 * 1024,
spool_directory: Optional[str] = None,
close_timeout_seconds: float = 10.0,
console_color: str = AUTO.value,
)
Declared configuration for framework loggers.