Enumerate constraints first: input format (jpg, png, multi-page tiff), model I/O (CPU vs GPU), expected cardinality (100k images), target SLA. Standard skeleton:
- A globbing iterator that yields
(path, stable_id) to preserve deterministic ordering even after sort.
- A thread or process pool sized to
min(cpu_count, model_concurrency).
- A producer-consumer queue decoupling disk I/O from inference.
- An output writer that flushes per shard to avoid data loss on crash.
- An optional retry queue keyed by
stable_id, not path (filenames can collide).
paths = sorted(glob("*.jpg"))
pool = ThreadPoolExecutor(max_workers=GPU_CONCURRENCY)
for path, result in zip(paths, pool.map(run_pipeline, paths)):
write(stable_id(path), result)
Gotchas that separate strong from weak: blocking the event loop on numpy decode; hashing by path when timestamps change; holding the entire batch in memory; not handling intermittent CUDA OOM.