Why I joined Blacksmith to work on storage again
16 hours ago
- CI storage problems are unique due to ephemeral VMs, high churn, and concurrency—Blacksmith creates and destroys ~2.5 million volumes per day, writing ~2 PB with 95% discarded.
- The workload includes read fan-out, concurrent writes to shared caches, and a need for correctness without corruption, requiring reliability during a job but tiered durability.
- Off-the-shelf solutions (block devices, tool caches, object stores) fail because they assume persistent state or don't handle CI's churn, concurrency, and granularity.
- Key open problems: inbound delivery of root filesystems, efficient environment assembly exploiting overlap, classifying write durability online, reliable artifact survival, and preventing shared blast radius.
- Pooling diverse CI traffic smooths peaks, making the economics of shared infrastructure a major opportunity.
- The work is full-stack, production-grounded, and green-field but constrained, offering a rare chance to rethink storage foundations.