Jobs

The scheduled jobs behind cairndb snapshot and cairndb gc. See Operations.

Snapshot builder - creates SQLite snapshots from log replay.

Designed to run as a scheduled serverless job. Building is idempotent and race-safe: replay is deterministic and the upload is put-if-absent, so two jobs racing produce one object with identical content either way.

class cairndb.jobs.snapshot.SnapshotBuilder(storage: BlobStorage, registry: HandlerRegistry, init_schema: Callable[[str], Awaitable[None]] | None = None, schema_version: str = DEFAULT_SCHEMA_VERSION)[source]

Builds SQLite snapshots by replaying the commit log from scratch.

Always replays from the beginning of the log to guarantee determinism — two builds over the same commits produce equivalent databases.

Workflow: 1. Create a temporary SQLite database (metadata table + app schema) 2. Replay all commits (up to end_at if specified) 3. Upload the database with put-if-absent under its commit number 4. Clean up temporary files

Initialize the snapshot builder.

Parameters:
  • storage – Blob storage backend (read commits, write snapshots)

  • registry – Event handler registry with all handlers registered

  • init_schema – Optional async callback creating app tables at a given db path before replay

  • schema_version – Projection schema version — selects the snapshots/v{version}/ prefix

async build(end_at: int | None = None) → int[source]

Build a snapshot by replaying the log from the beginning.

Parameters:

end_at – Optional last commit number to include (point-in-time snapshot). None means everything currently in the log.

Returns:

The commit number the snapshot represents.

Raises:

Garbage collection: retention for snapshots and the commit log.

Deletion is the only mutation in the system besides put-if-absent, so GC is deliberately conservative:

  • Snapshots: keep the newest keep_snapshots for the given schema version.

  • Log: only when prune_log is set, delete commits at or below the OLDEST KEPT snapshot — every kept snapshot can still bootstrap, and full history from that point remains replayable.

Caveat: log pruning considers a single schema version. If multiple projection schema versions are active, run prune_log only for the version whose oldest kept snapshot is lowest.

async cairndb.jobs.gc.collect_garbage(storage: BlobStorage, schema_version: str, keep_snapshots: int = 3, prune_log: bool = False) → GCResult[source]

Apply retention policy.

Parameters:
  • storage – Blob storage backend

  • schema_version – Projection schema version whose snapshots to prune

  • keep_snapshots – Number of most recent snapshots to keep (>= 1)

  • prune_log – Also delete commits fully covered by the oldest kept snapshot. Leave False to retain full event history (cheap in archive-tier storage, and history is usually the point).

Returns:

GCResult with deletion counts.

class cairndb.jobs.gc.GCResult(snapshots_deleted: int, commits_deleted: int, oldest_kept_snapshot: int | None)[source]

What a garbage-collection run removed.