Expanse Debuts GPU Cluster Optimization
TL;DR. Expanse (YC P26) launched a platform to boost GPU cluster utilization and predict job needs, addressing widespread compute waste. - The system analyzes job source code and hardware telemetry to provide accurate resource recommendations. - Its deep learning models improve usage rates and detect potential job failures before they occur. - Expanse reports 59% compute waste on one national HPC cluster, representing millions in lost resources.
- Expanse seeks to improve GPU cluster utilization from 30-40% to higher rates.
- The platform uses deep learning models to predict resource needs, detect failures, and suggest optimizations for HPC/GPU workloads.
- Expanse integrates with schedulers like Kubernetes and SLURM, analyzing job scripts and hardware data.
- One study cited 59% compute waste, equating to an estimated $8.5M monthly, on a national-scale HPC cluster.
Sources
- Launch HN: Expanse (YC P26) – Unlock Wasted GPU Capacity — news.ycombinator.com