A computing fanatic has repurposed a really noisy and largely out of date enterprise GPU (with a number of VRAM) for native LLM inference functions. They are actually having fun with a system that has doubled its complete VRAM quota to 32GB for only a $266 (£200) outlay. That’s a superb consequence, particularly within the midst of a RAMpocalypse.
Oscar Molnar explains that an inexpensive Tesla V100 SXM2 with 16GB HBM2 was sourced, as was an SXM2-to-PCIe adapter, and a PWM mod for the loud-as-a-lawnmower cooler, to finish this VRAM growth for the hefty native LLMs challenge. Indeed, these GPUs do look low-cost proper now, as I can see them listed on eBay US for under $140 every, in case you don’t thoughts shopping for from China.
As talked about above, you possibly can’t simply get certainly one of these Tesla V100 SXM2 playing cards with considerable VRAM and plug it into your PC. Molnar says they spent about $66 on an SXM2-to-PCIe adapter, also on eBay.
Latest Videos FromTom’s Hardware
You might think that was enough. However, the PC and local LLMs enthusiast baulked at the noise of “the fan from hell,” which came as standard with the Tesla V100 SXM2. That shrieking cooler was measured outputting 82dB of noise. Molnar described it as “somewhere between a garbage disposal and a lawnmower.” This may be the most complicated tweak yet, but basically the existing fan wires just needed rerouting and plugging into the motherboard PWM fan header. You could also simply purchase a “2.54mm male to PH2.0 female jumper cable” for the task. Apparently, the fan only needs to run at 10% to keep the Tesla V100 under 50C at full load.
27 billion parameter LLM runs at 32 tokens per second
With the hardware all now fitted and finessed, Molnar had a 32GB VRAM system at their disposal – that’s a PC with RTX 4080: 16GB VRAM, Ada architecture and Tesla V100: 16GB VRAM, Volta architecture. They note you can get Tesla V100s with 32GB of VRAM, but they are double the price.
Getting the system to make use of this 32GB of total VRAM for LLMs wasn’t tricky, says the DIYer. They used NixOS with a legacy Nvidia driver that overlapped support for both Volta and Ada architectures. Testing a local LLM, they got a 27 billion parameter model running at 32 tokens per second, which they say is “fast enough for interactive use” and faster than most cloud API alternatives.
Follow Tom’s Hardware on Google News, or add us as a preferred source, to get our newest information, evaluation, & opinions in your feeds.