NVIDIA Made This GPU Useless, So Modders Unlocked It
47sThe dramatic story of turning e-waste into a powerful AI card taps into the tech community's love for modding and sticking it to big corporations.
▶ Play Clip"The title promises a deep dive into modding a mining card, but the video spends significant time on sponsor content and rambling, delivering only a surface-level overview."
The video revisits NVIDIA's CMP170HX mining GPU, originally crippled to be useless for AI, and explores how modders have unlocked its hidden potential. It covers the card's limitations, the modding process, and its practical value for AI workloads.
The CMP170HX had no graphics hardware, crippled PCIe bus, memory locked to 8GB, and floating point performance locked down.
Modders unlocked floating point performance, hidden memory, and PCI Express interface, turning it from e-waste to e-grit.
Defective cards may have memory or other manufacturing defects, often binned and sold in bulk.
Cooling solutions include shrouds (quiet but space-consuming) or water cooling.
Modded cards can run standard large models (20-40GB), support large context, and are fast enough for real-time interaction and agentic tasks.
CMP170HX's Original Limitations
Explains why the card was considered worthless, setting up the modding story.
Modding Unlocks
Details the specific unlocks modders achieved, showing the card's potential.
01:11Defect Lottery
Highlights the risk of buying these cards due to potential manufacturing defects.
13:49Practical Value
Summarizes the card's real-world utility for AI tasks, despite limitations.
16:46[00:00] Five years ago, we first covered NVIDIA's CMP170HX, a mining GPU so utterly worthless that its very existence offended me. See, unlike a real GPU, these crypto cards didn't have any graphics hardware built into them whatsoever.
[00:19] All they could do was crunch crypto numbers, meaning that as soon as they were unprofitable, they were garbage. But why? Compute CPUs can be used for lots of things. And in fact, the A100 cards that
[00:35] these mining cards were derived from in the first place are still in high demand for AI applications. Well, because NVIDIA went out of their way to make them useless. They crippled
[00:47] the PCI Express bus, locked the memory down from 64GB of HDMI to just 8GB, meaning it meaning it could barely run any decent sized models, and to add insult to injury,
[00:59] they locked down the floating point performance too. Which raises the question, Why am I so happy about that? Because, as Lucas from the lab is about to show us,
[01:11] enterprising modders have found ways to unlock the floating point performance, the hidden memory, and even the PCI Express interface, turning it from e-waste to e-grit!
[01:24] race racing like they're fast oops sorry I stepped on you it was I can't fit all of you in my field of view so we're going to show you guys
[01:36] how to do it and talk about whether one of these cards could be right for you after the segue to our quest behold wisdom from the unmoistened one
[01:54] Thank you. Step one is to stand on this box so that our camera operator can frame for both of us. Step two is to establish a baseline of what kind of models we can actually run on this card
[02:11] in its stock configuration and what its performance looks like. We will be using LalaBench since it's the simplest. It's not a comprehensive benchmark, but it gives us a number, which is all we really need to do a comparison like for like.
[02:24] Literally the exact same car running the exact same architecture, but more stonker. You might have noticed we've got a lot of GPUs on this bench. The 3090Ti's are here for sort of a...
[02:36] They're not quite equivalent in price, are they? These are a little more? They were equivalent in price before, and now the CNPs have shot up in price. So when we were looking before, they were about $1100 each, $1100 to $1500,
[02:49] and now the CNPs are going for like $1500 to $1900. 3090Gi's are here as sort of a faster compute but less RAM point of comparison. We've got a 1650 that's just here for video output and then we have two of these mining cards to show
[03:04] you how well things can go and also how not as well things can go with the unlock. We have a lot of CVT running here on, I would say, Granite's 4.1 model, 8 billion, so we can answer some questions like... Sure.
[03:17] ...a question of the day. He's the best tech YouTuber. Wow, that is out of date. And also really stupid. Smosh? Lots of love to Ian and Anthony, but, like, tech YouTuber?
[03:29] Yeah, we had to do that basically because that's all we can fit on the 8 gigabytes. So if we put that onto the 3090 first, we can see it loads much faster. Oh, that did load much faster.
[03:42] And it rips away. Right. And that just comes down to the way faster compute because both of them are running really small models. Now, what would happen if we tried to run a larger model on our 8 gigabyte mining card?
[03:56] Please, Client 3.8, do not go overage now. I think this is like 28 gigabytes-ish. To be clear, we do have over 28 gigs of system memory for it to fall back on.
[04:08] It's probably going to fail very soon, yeah. Also, the context we did for the small model, it could only be 16k instead of 64 on it. But here you can see it's... Yeah, so red text is good, right?
[04:20] Yeah, it's the warning one. And this should, this will fit the... Actually, this might not fit on the 3090. But it'll fit on two, which we can do as well. Do we have an NVLink bridge? Do we even need it? We do, but we don't really need it.
[04:32] We actually didn't see too much of a gain of it on the NVLink bridge. That is another thing The 390 is sometimes inconsistent and they crash BRB With our baseline out of the way it time to perform the unlock How would you rate this on a scale of
[04:51] 1 to 10 for difficulty? Pretty easy. 6. Okay. 6, you got a million little things. I think the re-lock afterwards
[05:04] is where it got more difficult, but I'll say that's the best. But yeah, it's not too bad. And it's like steps on in the repo. from this guy Amogh Mooney Kostay, I'm going to butcher his name, but we'll have it linked
[05:16] of course somewhere. And we'll actually have a link down below to the full LTC Labs companion article that'll go with this video, so you're going to want to be sure to check that out. Yeah, but we just have, you know, the clone github repo, you have to disable secure boot,
[05:30] and then nuke all the existing NVIDIA drivers, just purge those, blacklist the new row drivers, so those don't kind of try to jump in. Yeah, oh, dude, I ran into that a little while ago with Fedora. It just, like, kept installing drivers that I didn't want
[05:45] over and over and over again when I was trying to get some Volta cards working on it. It was a whole thing. Yeah, but there's a kind of config file you can write to them. It's a little stuff app. Cool. Rebuild the RAM file system and all that. And then specify which drivers you want
[05:59] because there's specific drivers that you just built for that you can go for and this unlock will act on. Now, tell me this, another thing I'm going up against is that on Linux, or at least on Fedora, it's a little bit more challenging to run different NVIDIA drivers for your different
[06:14] NVIDIA cards like you can on Windows. Do you have to run the same NVIDIA driver across all your cards? I actually don't know. I kind of just stuck with the single driver to make sure it works in this limited time. And with the unlock.
[06:26] Chat with us in the comments. In the meantime, though, are we pretty much ready to press execute and show the unlocked cards? Yeah. So basically, instead of unlocking the card, you're just running a modified version of
[06:40] the driver software that unlocks any card that you put in the system now, right? Yeah, that's correct. There's some papers in my earlier research on unlocking the cards specifically, and they're actually really interesting papers to read, but this one acts on the system and the driver.
[06:52] So any card you plug in will be unlocked, and if you take that same card and plug it to a different system, it'll be locked again until you unlock that system. I love this kind of black magic, Tomfoolery. It's very cool to see, and kind of the unlocking of all this potential and hardware.
[07:07] I mean, some of the papers I read were like, oh, you know, investigating this is like, you know, things we should watch out for, so people can't unlock it in the future, but, you know. Oh. Close, guys. Yeah. There you go.
[07:19] Great. Nvidia CMT-178HX. We didn't even talk about the power limit going up, too, then. Yes, well, actually, it's the same power limit, but because the computers fell lockdown before, all the floating point operations, it never even used anywhere near the 250 watt.
[07:34] But now we can get up near that 250 watt. Definitely, yeah. When it's running. It depends on what's running, of course. Alright. Why don't we start with that little model and have a look at what our better floating
[07:46] port performance gets us. I think it's quite slow. But that's mostly because of the PCIe lane kind of thing. Because it unlocked it to PCIe Gen 2. But it's still a x4. So our bandwidth is seriously limited.
[07:59] Still limited, yeah. That can be unlocked, but that takes us past Ramah, that we'll get into a bit more later. Certainly. And my shot is kind of like what we were seeing with the 3090. Wow, that is world faster. Still completely crap analysis.
[08:15] We'll give it a larger prompt too. Sure. Because you don't really see the prompt costing speed unless you give it a large prompt like that. As you see, the 3090 Ti, a single one of them, was getting about 7,000 I think it was, or 4,000, I don't know exactly.
[08:29] Just so you know, it's also up there. Right, yeah, it's slower, right? But this is a completely different world of speed compared to what it was doing before. And now let's see what happens when we load that Qen model that everyone is mayonnaise-ing themselves over.
[08:44] There we go, hopefully it runs. It'll always be slower to load up, that's kind of the main lock from the PCIe thing. Right, because even though we've got that 64 gigs in this, HBM2E, which is still very fast, like right on par with the GDDR7 that you'd get on something like a 5090.
[09:02] If the data's got to be loaded through the PCIe bus, if that bus is slow, then you've got to wait around. And that wasn't that bad, though. That's not too bad. It's a larger model. But that's the same thing, and probably get a way better answer.
[09:15] So we'll see how this runs, because Quint likes to think a lot. And I didn't give it any reasoning effort command, So it may just stink for a really long time before it's put in there.
[09:27] Okay Yeah I get it It not ordered I knew this was going to happen It not ordered I knew it was going to happen It a single best Maybe that is a smart decision You know, watching multiple reviews. This is very usable though. Very good. Enter D&C. Cut
[09:47] that, cut that, cut that. Alright, let's see what happens if we give it a much longer countdown. This will be slow, it will be some time first token, but still quite quick.
[09:59] That's a thousand words of warm wisdom. Now you mentioned before that because we've got twice the memory of the 3090 Ti, we could actually run even larger models on our mining card. But question for you, given that it doesn't have the fastest compute,
[10:15] would an agentic workload make more sense for a card like this? and would we instead want to just use that extra VRAM for more context and then let it turn away at things in the background? It's really good to have all the extra money, obviously, for the larger models
[10:31] and also if there's multiple concurrency going on, so they don't have large context. Or if you're doing some kind of task that needs multiple different models, if you have your main agent you're talking to and it's also doing embedding or image generation,
[10:45] So that's kind of the advantage you get from this huge amount of memory. So basically yes, and also yes, and also yes. It's just way more flexibility. I'd love to buy one off of Logistics. Now let's try running Client on a 3090 Ti.
[10:59] Or should I say 3090 Ti's because we kind of have to pool the memory in order to run that without it crashing? Sure, you can try running on just one right now. Okay. I mean, did we crash last time?
[11:11] Yeah, it will work. Because it's all like 20, 29 gigabytes. and then we add 200 GMA just on layer split mode this will work since we got 48 gigabytes. Right. We should get a very similar quality of answer and we should actually
[11:24] get measurably more performance in this case. It looks really similar. Lower, yeah. So prime got 154 but it said it's multiple ounce. Generation O's you got 30.
[11:36] Not as much of a difference as you might think but let's go ahead and throw a bigger prompt in there. It was a bit of a pause, but that was like, I think, 2,000 words or something. It'll think for a while. It's about 2,000 tokens per second.
[11:48] Right, okay. So that's where the better compute of the 39 API comes into play. Of course, just like we can run two 39 APIs, we can also run two mining cards. This is where, once again, the PCIe bands with limitations could rear their ugly heads.
[12:04] Yeah, this depends totally as well on the tensor versus layer split mode, like parallelism. and it's kind of the same thing where it's like there's so many factors that could, you know, limit it, bottleneck it, or make it speed up.
[12:16] That's his way of saying that these are not certified labs test results. It depends. These are quick and dirty illicit numbers, and anyone actually deploying these would do a bunch more tuning and almost certainly get better results.
[12:29] As we've been doing, like, in the Discord, they had got there for the CMP and lock. They've done a lot of different really cool things there, so you should check that out as well. So it's kind of freaking, you know, half the processing speed of it, the same token generation.
[12:41] Still really usable. Yeah, this is great for like if you're using some local coding or whatever agentic thing. I think over like 10 to 15 is usually pretty good for kind of the specs, but yeah. But, it's not all sunshine and rainbows either.
[12:55] As Lucas already alluded to, in the time since this unlock was discovered, to the time that we're filming, these have gone up very significantly in price, and there's no guarantee whatsoever that you're going to actually get a working one.
[13:08] In fact, even though we were able to process a prompt just now on these two cards, one of them is bad. Let me show you what that means here. So we'll mem test this, and we'll try to target one of the HXs.
[13:23] This test will just go through all of the memory and check it for any errors, write and read it back, see if anything's changed. It should show up within the first five minutes of the test, even faster on my previous test.
[13:35] But this, I would say, seems like the good one. See, there's lots of reasons that a card might be demoted from A100 to crap mining card. Maybe it doesn't meet the power threshold.
[13:49] Maybe it doesn't meet the quality standards. Or maybe the HBM memory, which, if you guys recall, is right on the package, and once it's installed, cannot be removed,
[14:01] was damaged or otherwise failed during manufacturing. And when that happened what NVIDIA was doing was they were basically using off most of it and selling it as a crap mining cart So whether you get one that has a memory defect or some other kind of defect pure luck of the draw And like Lucas mentioned maybe not even pure luck because a lot of people are going to be buying these in bulk
[14:27] binning them, and selling the crap ones to take advantage of the current market situation. Oh, we found a bad one. So if it's all good, it'll kind of keep showing through the list and show iterations of the code. Yeah. Turn it five minutes and you should let it go for longer for this one.
[14:41] Bad factors. And what does that mean? That basically means that at some point, we're going to ask it to do some math, and it's going to burp out. Whenever you're running a crash, I saw some stuff on the Discord
[14:54] where people were trying to run it, and they were seeing inconsistency. We didn't see any inconsistency, but we didn't run it for a long time like that. Who we borrowed this card from said that as it gets really hot, sometimes that fixes their kind of inconsistency.
[15:07] Like, when it's above, like, 85 degrees, kind of like, the limit there. But who knows? I wouldn't rely on that for production. Right. If you do get a good card, there's another step that you can go to to unlock even further performance.
[15:19] Using some loop schematics on the 170th Street repo, our resident slaughtering wizard, Dan, managed to put the 24 little tiny capacitors back on this thing, above the PCIe slot,
[15:32] unlocking 4x16 functionality. So let's go ahead and throw two modded cards into our system and see how they fly. While we're performing this operation, it seems as good a time as any to talk about cooling these monsters.
[15:48] This is the same shroud that we used last time around. We'll have it linked in the video description. What's good about this approach is it works pretty well and is reasonably quiet, but what sucks about it is it takes up a lot of space, forcing the use of risers, which just adds another variable whenever you run into psyche
[16:03] behavior. So if you're more interested in water cooling, you'll find more information about that on 178 Street. I guess you're waiting for me to give you this, so you have the, can have the fan. I'm so sorry. I don't think we're going to get into the nuance of how different concurrency methods might benefit more or less from the additional bandwidth.
[16:19] I just want to see something everyone will benefit from, which is faster model load-in. So, want to throw Quinn in there? Loading model. And this was, if you remember, quite slow the first time around on a single card.
[16:33] And that was quite, not nearly as slow. So, what's your takeaway? Is there a charcoal for me? No, no, we're going to answer it in a second. I mean, sure, yeah, take your crack at it. Buy one, I guess. If you can get a working one.
[16:46] Yeah, yeah. No, I think it's great. It kind of fits all the standard large kind of models you'd fit, because they're kind of now contained between like the 20 to 40 gigabyte range. And it leaves a lot of room for context to make it half a million or a quarter of a million context.
[17:00] Or running additional models. It's fast enough to interact with in real time, but definitely fast enough to work on agentic tasks in the background. There are some inconveniences, but even at the current pricing, I think it's a pretty good deal.
[17:13] And there will be more optimizations and people finding workarounds for things, maybe some more unlocks as well. So, pretty interesting. Future's bright. because it contains this segment. To our sponsor.
[17:29] Water? No. Water? No. No, I haven't seen it. I know, with cats and dogs out there.
[17:41] I hope it's wearing a dirty, cranky, filthy. Damperton Falls has never been this damp. Like an infection. The rain wormed its way into every crevice.
[17:54] There wasn't a dry sulk in the city. Except the one inside that. Stain. Expect the unexpected with the Bessie Weekend Chelsea.
[18:07] Lightweight, breathable, comfortable, and waterproof. Bet you didn't see that coming. Simple, stylish, and versatile. It's one pair for whatever might happen. Its existence is a paradox.
[18:20] Does that mean it can't be real? Or that you can't get 15% off your first order with a one-year warranty and free shipping on North American orders over $110?
[18:36] Shh. Just embrace Vessi. Now to return most of these to their rightful owner. If you guys enjoyed this video, maybe check out the one where we used a different cut-down mining card for gaming by routing
[18:53] its GPU power through our onboard graphics. Which, if that sounds complicated, it's because it was a little bit.
⚡ Saved you 0h 19m reading this? Transcribe any YouTube video for free — no signup needed.