A Modder’s RTX 4080 Was Sufficient To Play AAA Video games, However Not For Running LLMs, So He Built-in NVIDIA’s Tesla V100 At A Throwaway Value To Run 27B AI Fashions

This web page was created programmatically, to learn the article in its unique location you’ll be able to go to the hyperlink bellow:
https://wccftech.com/nvidia-tesla-v100-modded-into-gaming-pc-to-run-aaa-titles-and-27b-ai-models/
and if you wish to take away this text from our web site please contact us


Running higher-quality AI fashions means your present GPU should be geared up with a ton of VRAM to make the expertise pleasing. For one RTX 4080 proprietor, enjoying probably the most visually taxing and graphically demanding video games may be a stroll within the park, however for operating LLMs, it’s a Herculean job. Wanting to perform each feats in a single gaming PC, a modder efficiently ran NVIDIA’s Tesla V100 in his system, however encountered a couple of challenges alongside the best way.

Adding the Tesla V100 grants the modder 32GB of usable VRAM, enough for operating considerably improved AI fashions like Qwen 3.6 at 32 tokens/second

Before you ask, it’s unattainable to connect the Tesla V100 to a desktop motherboard like a “plug and play” job, requiring Tymscar to buy an SXM2-to-PCIe adapter. Sourcing the GPUs with 16GB of HBM2 reminiscence and the accent value him £200, which interprets into roughly $266. On eBay, you’ll be able to seize these for under $100 apiece. With the Tesla V100’s 5,120 CUDA cores and 4,096-bit bus width that delivers 900GB/s of bandwidth, the GPU nonetheless has some computing juice remaining.

Successfully sourcing an SXM2-to-PCIe adapter wasn’t probably the most tough of challenges, however it isn’t easy both, particularly if you discover out later that the Tesla V100 doesn’t have a PCIe slot, show inputs, or PCIe energy connectors. As you’ll be able to see within the picture beneath, attaching the GPU to the vapor chamber-style cooler may be completely effective for many who don’t thoughts the extreme noise, however at 82dB, it’ll make anybody uncomfortable.

With slightly tweaking right here and there to scale back the fan noise by utilizing a 9V battery and a PWM jumper, the V100’s modded cooler was now working at 10 % of the unique most RPM. With this drawback out of the best way, the modder had efficiently discovered a option to get 32GB of usable VRAM into his system. Now, no recreation will ever require a whopping 32GB of video reminiscence, except you determine to run a more recent title at 16K decision, so the perfect use case could be to fireside up Qwen3.6 27B.

The modder was operating Qwen3.6-27B-MTP quantized at Q5_K_M, which is available in at 19GB, and with a context dimension of 128K tokens, there was enough VRAM to run the LLM at 32 tokens per second. Prompt processing was between 133 and 160 tokens per second, making it first rate efficiency should you occur to stumble throughout previous-generation AI GPUs for home-computing functions. Best of all, for lower than $300, you’ll be able to have your very personal small to medium-sized AI fashions operating at dwelling, freed from value and with none web connection.

News Source: Tymscar

Omar Sohail Photo

About the creator: Omar Sohail is a reporter and analyst for Wccftech’s cell part, specializing within the know-how and enterprise of the cell trade. His experience lies within the intricate {hardware} provide chain, masking developments in semiconductor manufacturing, chip lithography, and digital camera sensor know-how.

Follow Wccftech on Google to get extra of our information protection in your feeds.


This web page was created programmatically, to learn the article in its unique location you’ll be able to go to the hyperlink bellow:
https://wccftech.com/nvidia-tesla-v100-modded-into-gaming-pc-to-run-aaa-titles-and-27b-ai-models/
and if you wish to take away this text from our web site please contact us