Skip to content

Runtime configuration

Runtime configuration is distinct from authored metadata. A runner supplies these values when creating a Driver session; the Driver does not read a project manifest at runtime. Types and defaults below come from the installed source models. Semantic constraints describe current runtime behavior.

Run configuration

DataCoolieRunConfig accepts declared fields as keyword arguments through its model constructor. Its dataclass declaration disables the generated initializer, so an empty generated signature would be misleading. These fields are not keys in authored metadata JSON.

Field Python type Default Meaning / constraints
job_id str generated per instance (generate_unique_id) Caller correlation identity; must be non-empty. Logging also supplies session and operation identities.
job_num int 1 Total shard count; at least 1.
job_index int 0 Zero-based shard index: 0 <= job_index < job_num.
max_workers int 8 Maximum concurrent dataflow workers; at least 1.
stop_on_error bool False Stop admitting new dataflows after a failure; already admitted work may finish.
retry_count int 0 Additional attempts after a failed execution; non-negative.
retry_delay float 5.0 Seconds between retries; non-negative.
dry_run bool False Use the Driver dry-run validation path. This still prepares runtime dependencies; it is not an offline schema-only check.
retention_hours int 168 Retention passed to destination maintenance; non-negative. Backend support determines the effect.
allowed_function_prefixes List[str] new [] per instance Import-path prefixes for Python function sources. An empty list applies no prefix restriction. A non-empty list requires the function path to start with one of these strings; use trusted module prefixes such as my_project.sources. to constrain imports.
run_attributes Optional[Dict[str, Any]] None Optional caller correlation object, copied on construction. Keys must be strings; nested values must be JSON-compatible with finite numbers and no cycles. Raw JSON strings and arbitrary objects are rejected.
from datacoolie.core import DataCoolieRunConfig

config = DataCoolieRunConfig(
    job_id="nightly-orders", max_workers=2, retry_count=1,
    run_attributes={"scheduler": "local", "attempt": 1},
)

Supply this object as config= to DataCoolieDriver. The create_driver factory instead accepts run fields such as job_id, max_workers and retry_count directly; it builds the configuration object internally. See runtime setup for a full runner.

DataCoolieRunConfig dataclass

DataCoolieRunConfig

Validated execution parameters for a DataCoolie run.

Replay configuration

Field Python type Default Meaning / constraints
start Any required Required inclusive lower bound, not None.
end Any required Required exclusive upper bound, not None.
chunk_interval Optional[Dict[str, int]] None None selects one bounded read. Time chunks use years/months/weeks/days/hours/minutes; integer chunks use {"step": N}. Constructor checks the mapping shape; execution validates the interval and range.
save_watermark bool False Persist source-observed watermarks after successful chunks. Does not create a replay checkpoint or skip chunks on a later run.
chunk_column Optional[str] None Non-empty override, or the first source watermark column. The reader must support bounded reads for the selected column.

Replay uses [start, end). Construction does not prove reader capabilities or validate every range/interval combination; execution performs those checks. See replay and backfill.

ReplayConfig dataclass

ReplayConfig(
    start: Any,
    end: Any,
    chunk_interval: Optional[Dict[str, int]] = None,
    save_watermark: bool = False,
    chunk_column: Optional[str] = None,
)

Configuration for replaying a bounded time range in chunks.

Used by :meth:DataCoolieDriver.run_replay to reprocess historical data without corrupting the production watermark.

The range uses the left-closed, right-open [start, end) convention: start is inclusive, end is exclusive. This aligns chunks to whole calendar units (days, weeks, months, etc.) and is the industry-standard interval convention used by Python’s range(), Spark partition pruning, and PostgreSQL range types.

Example::

# Replay all of Q1 2025 in monthly chunks:
ReplayConfig(
    start="2025-01-01",  # inclusive
    end="2025-04-01",    # exclusive (first day NOT included)
    chunk_interval={"months": 1},
)
# Produces chunks: [Jan 1, Feb 1), [Feb 1, Mar 1), [Mar 1, Apr 1)

The chunk column is auto-resolved from dataflow.source.watermark_columns[0] at runtime. Override with chunk_column for a source-supported independent bounded-read column or when the first watermark column is not the one to chunk on.

Type detection is automatic:

  • str parseable to date/datetime → time-based chunking
  • datetime / date objects → time-based chunking
  • int → integer-based chunking

Parameters:

Name Type Description Default
start Any

Inclusive lower bound of the replay range.

required
end Any

Exclusive upper bound of the replay range.

required
chunk_interval Optional[Dict[str, int]]

Chunking interval. Time-based keys (months, days, hours, minutes, weeks, years) use relativedelta; step key is for integer watermarks. None disables chunking (single-shot replay).

None
save_watermark bool

When True, persist the source-observed watermark after each successful chunk. Replay is always re-runnable: this flag does not create a checkpoint or skip chunks on a later run. When False, the stored watermark is never touched.

False
chunk_column Optional[str]

Override the auto-resolved chunk column. Use this for an independent bounded-read column when the selected source reader supports it, or when the first watermark column is not the desired chunking dimension. API readers require a matching range_param_mapping binding.

None

Logging configuration

Field Python type Default Meaning / constraints
log_level str 'INFO' Console threshold: DEBUG, INFO, WARNING, ERROR, CRITICAL. Names normalize to uppercase.
file_level str 'DEBUG' Persisted Python-record threshold; same supported levels as log_level.
storage_mode str 'memory' Local capture: memory or file; names normalize to lowercase.
output_path Optional[str] None Standalone logger component root, or None for no remote writer. Must be a non-empty path when supplied. Driver log_base_path instead creates category roots.
partition_by_date bool True Boolean controlling UTC date partitioning.
partition_pattern str '__run_date={year}-{month}-{day}' Non-empty template with simple placeholders in ordered prefix: {year}, {month}, {day}, {hour}. Each path segment needs a placeholder; literals cannot contain digits, %, or braces.
persistence_mode str 'snapshot' snapshot or batch; names normalize to lowercase. See the logging guide for write layout and backend requirements.
flush_interval_seconds float 300.0 Finite non-negative seconds; 0 disables time-triggered flushing. Booleans are rejected.
flush_batch_bytes int 4194304 Positive integer byte threshold; booleans are rejected.
buffer_memory_bytes int 67108864 Positive integer memory buffer budget; booleans are rejected.
spool_max_bytes int 536870912 Positive integer spool budget, at least buffer_memory_bytes; booleans are rejected.
spool_directory Optional[str] None Optional non-empty local spool path; normalized when supplied.
close_timeout_seconds float 10.0 Positive finite close timeout; booleans are rejected.
console_color str 'auto' auto, always, or never; names normalize to lowercase. Auto respects NO_COLOR and TERM=dumb before terminal detection.

Logging capture is bounded. Buffer or spool exhaustion can lose records and is reported through logging statistics; logging persistence is not a transactional guarantee for business execution. For activation, flush, close and storage behavior, see logging and the logging API.

LogConfig dataclass

LogConfig(
    log_level: str = INFO.value,
    file_level: str = DEBUG.value,
    storage_mode: str = MEMORY.value,
    output_path: Optional[str] = None,
    partition_by_date: bool = True,
    partition_pattern: str = DEFAULT_PARTITION_PATTERN,
    persistence_mode: str = SNAPSHOT.value,
    flush_interval_seconds: float = 300.0,
    flush_batch_bytes: int = 4 * 1024 * 1024,
    buffer_memory_bytes: int = 64 * 1024 * 1024,
    spool_max_bytes: int = 512 * 1024 * 1024,
    spool_directory: Optional[str] = None,
    close_timeout_seconds: float = 10.0,
    console_color: str = AUTO.value,
)

Declared configuration for framework loggers.