Self-Healing GPU Fleets: Auto-Repairing SUNK Safely at Scale
About this session
At GPU scale, healthy hardware sits idle. Slurm nodes get drained for transient faults - failed health checks, prolog/epilog errors, insufficient timeouts, connectivity issues, etc. - that clear on their own, yet stay drained until a human resumes them. Multiply that by a large fleet and it's a standing tax on GPU-hours and on-call attention. This session shows how we turned that manual toil into a self-healing fleet. SUNK Auto Repair is a policy-gated auto-repair controller that classifies every drain against declarative safety rules and resumes only what passes - behind verification gates, rate limits, and a dry-run mode for trustworthy rollout. We will also talk about future enhancements on the roadmap that include repairing not just nodes but also other components in the SUNK/Slurm ecosystem that frequently fail. You'll leave with a blueprint for closing the loop from observability to safe, automated action.
Share this session


Get in the room. San Francisco, September 29.
Fully Connected 2026 is where the engineers, leaders, and operators running AI in production come together for three days of depth, access, and real conversation.