Datacenter Nvidia GPU Powers Local LLM Inference for £200
TL;DR. A user repurposed an Nvidia Tesla V100 datacenter GPU, originally for DGX servers, to significantly boost local LLM inference capabilities on a gaming PC. - The customized setup provides 32GB of VRAM across two GPUs, running a 27-billion-parameter model at 32 tokens per second. - The V100 GPU cost £150 on eBay and was connected using a £50 SXM2-to-PCIe adapter. - This configuration delivers superior memory bandwidth for LLM inference compared to newer consumer cards at a fraction of the cost.
- A datacenter NVIDIA Tesla V100 SXM2 GPU was adapted for a gaming PC.
- The setup provides a total of 32GB VRAM for local LLM inference, running a 27B parameter model.
- The V100 offers 900 GB/s HBM2 memory bandwidth, outperforming many newer consumer GPUs.
- The total cost for the V100 and adapter was approximately £200.
- This solution offers significant cost savings compared to high-end consumer GPUs for LLM tasks.
Sources
- I Put a Datacenter GPU in My Gaming PC for £200 — blog.tymscar.com