[00:00] Nvidia showed this thing off 9 months before anyone could actually buy it, then raised the price 17% after it shipped. This is the box that started the supercomputer on your desk category, and it might be the worst deal in it. [00:16] I dug into all three of these machines for this series, and today you get the one number that decides which one to buy. I'm also correcting two numbers I got wrong on air. Rewind to CES in January 2025. [00:31] Jensen Huang holds up a box roughly the size of a Mac mini and calls it Project Digits, a personal AI supercomputer aimed at $3,000. At GPC that March, it gets its real name, VGX Spark. [00:44] The thing finally ships on October 15, 2025 at $3,999. Huang personally hand-delivers some of the very first units, one to Elon Musk at SpaceX, and one to Sam Altman at OpenAI. [00:58] Then it goes quiet for months. Then February 27, 2026 happened. NVIDIA moved the Founders Edition to $4,699, a 17.5% jump on a box that didn't change in any way. [01:14] Same chip, same memory, and $700 more. The stated reason was memory supply, and that part is genuinely true. As of this week, the price hasn't come back down. Cured the part, the friendly reviews skipped. [01:27] NVIDIA announced this category first and shipped it close to last. By the time Sparks were landing on desks in October, you could already buy an AMD Strix Halo box with the same kind of big unified memory pool. [01:40] Plenty of them were already running in people's homes. Being early with a keynote and late with a product is how you hand your own category to somebody else. One macro thing frames all three episodes in this series. Every local AI box got more expensive in 2026 for exactly one reason, and that reason is memory. [01:58] TrendForce reported conventional DRAM contract prices up 93 to 98% quarter over quarter in Q1 2026, with PC DRAM projected to climb over 100%. [02:10] Apple called it an extraordinary surge in demand in June, and Tim Cook described it to the Wall Street Journal as a 100-year flood. Wong said in Seoul that the whole supply chain is short and stays short for years, and Micron's CEO expected tight past 2027. [02:26] Decode speed is bound by one thing, and that thing is memory bandwidth. Decode is the rate tokens come out once the model starts talking, which is the speed you actually feel in a chat window. [02:38] Divide the bandwidth by how much weight the model pulls for every token, and that lands close to your tokens per second. The Spark reads memory at 273 gigabytes per second, and that's 273, not the 300 you'll still see quoted off the hotchip slide. [02:54] The AMD box reads at 256 on paper. Those two numbers sit 7% apart, so on chat speed, they're basically the same machine. Pre-fill is a completely different animal. [03:06] Pre-sale is the model reading your prompt before it answers, and that work is compute-bound, so it scales with tensor cores instead of bandwidth. This is where a Blackwell GPU with 6,144 CUDA cores walks away from an AMD APU. [03:23] On the same test, the Spark chews through prompt about 5 times faster while answering only 13% quicker. If you bought this box for chatting, you'd pay 3 grand extra for 13%. [03:35] If you shove huge prompts and whole documents into it all day, that 5x is the entire reason it exists. Which brings me straight to my own receipt. In the Mac Mini episode, I said the M4 Pro moves 550 gigabytes per second, and that's wrong. [03:53] That number belongs to the M4 Max, not the M4 Pro. The M4 Pro Mini runs at 273, which is the exact same bandwidth as that $4,700 NVIDIA box. [04:07] The decode ceiling on those two machines is identical, and the price absolutely is not. Here's the second one. In the AMD episode, I used 122 gigabytes per second as the Strix Halo's real-world bandwidth, [04:21] and 122 is the floor of the range rather than the typical result. Measured numbers land anywhere from about 122 up to 215 against 256 on paper, [04:34] so I made that box look slower than it usually is. The prices from both episodes have aged badly, too. That $1,500 GMK Tech Evo X2 now feeds closer to $2,000 and sometimes $2,700. [04:48] Apple raised the M4 Pro mini base to on June 25th after killing the 64GB option in May so 48GB that you stealing there now Apple also paid Google roughly a billion dollars a year for a custom Gemini model just for Siri The BLMAR German broke in January [05:07] Even Apple's wall garden depends on outside AI when it actually matters. Every one of these boxes sells the same promise, which is that you stop renting compute from somebody else's data center. I buy that argument completely. [05:20] But look at what's actually draining your card every month. because for most people, it isn't keeping you time at all. It's Claude, plus chat GPTs, plus a video generator, plus a voice tool, plus a stock image plan. [05:33] That stack quietly costs more per year than the AMD box does, and it's the exact problem I built my own platform to kill. This part is my product, so treat it as a first-party pitch and judge it that way. [05:46] This is AL Master, and I'm opening the multi-model chat first. Claude, ChadGBT, Grok, and Gemini all fit in one window. I can show the same prompt at three of them side by side without paying for four separate subscriptions. [06:00] The per-token cost runs cheaper than going direct, which matters a lot once you're running real volume. Now watch me switch tabs into generation, because image, audio, and video all live in the same workspace. [06:13] It's also cheaper than stacking individual tools everywhere else. The feature I use most is consistent characters. I lock a character once and it looks like the same person across every generation. That's the wall most people hit with raw image tools. [06:27] You can publish that character on the platform and let it earn money on its own. Anything you generate can also be shared into the feed, so your content gets real reach instead of dying in a downloads folder. [06:39] Then there's the AI content engine, which runs a whole channel for you. AI agents write the scripts, descriptions, and thumbnails, so you're operating a channel instead of touching every single asset yourself. [06:51] There is a community of over 13,000 people inside. I'm scrolling the result page right now so you can see real cases and screenshots instead of just taking my word for it. There are also testimonials from people actually using it in the demo videos. [07:05] The Academy sits alongside all of it, with more than 200 lessons across roughly 30 hours. Here's how you get access. Go to the link in the description and hit buy. Select the annual plan, fill in your details, and you'll get a confirmation email. [07:19] Then log in, and you're inside the same workspace I just showed you. The whole setup takes under three minutes. The link is waiting in the description. There's a seven-day money-back guarantee, [07:32] which is me telling you to test it properly and take your money back if it doesn't fit. I pulled the same test across all four machines from Public Benchmarks. I picked GGT OSS 120B because it's the biggest thing most of these boxes can actually hold. [07:47] On llama.ccp, the DGX Spark does 1,723 tokens per second on pre-fill and 38.55 on decode. For a llama, it reached 1,169 pre-fill and 41.14 decode. [08:03] You'll see a figure near 1,821 floating around for pre-fill, and I'm not using it because it traces back to a single source. Next up is the $1,500 class AMD box. [08:15] Strix Halo turns in 339.87 pre-fill and 34.13 decode. Read those two rows next to each other, and the whole video snaps into focus. pre-fill is five times faster on the NVIDIA box, and decode is 13% faster for roughly $3,000 more. [08:34] The Mac Studio M3 Ultra does 863 pre-fill and 70.79 decode. Apple loses the prompt reading race badly. Then when's the part you feel in a chat? Because it reads memory at 819 gigabytes [08:48] per second. And the Mac Mini M4 Pro, it doesn't post a number at all, because 48 gigabytes can hold the model. Price per gigabyte of usable memory tells the same story from a different angle. [09:01] Strix Halo comes in around $11.70 a gig, the Spark around $36.70, the M4 Pro Mini around $41.60, and the M3 Ultra around $55.20. The Spark also gives you about 119 gigabytes usable out of its [09:17] 128, and there's no ECC on that memory, which is worth knowing before you leave a week-long fine-tuned running unattended. Two more specs matter here. The GB10 Grace Blackwell chip [09:30] carries a 20-core ARM CPU, 10x925 cores, and 10a725 cores, with a 6,144-core Blackwell GPU. [09:42] And that famous one-petaflop headline is sparse FP4 marketing math. Dense BF16, the precision you'd actually train in measures around 60 teraflops and Amy Hannon's [09:54] own micro-benchmarks landed in the same place. One teraflop is a real number for a very specific thing that is not the thing you doing Around January 8th 2026 Business Insider published an internal email chain from Inside [10:09] NVIDIA about how badly the Spark Food section was going, and Huang personally stepped in to defend it. He called it the Ultimate Developer's Platform, which is a very specific defense. He didn't defend the box on speed for the money, he defended who it's built for. [10:25] The lightning rod in every one of those threads was that same 273 gigabytes per second, because reviewers had spent a year expecting a desktop DGX to feel like a data center part. [10:38] Then the independent measurements started coming in, and they were messy in both directions. John Carmack put one on a meter. On October 27, 2025, he found it pulling roughly 100 watts against the 240-watt rating, [10:52] while delivering about half the quoted BF16 compute. Serve to Home measured closer to 200 watts under real load, so the box is genuinely efficient and genuinely slower [11:04] than the spec sheet implies at the same time. Both of those things can be true, and the marketing only mentioned one of them. The clustering story got rough too. Two units link over a single 200-git QSFP cable, [11:18] which pulls 256GB and officially supports 405B inference in FP4. NVIDIA officially supports scaling up to 4 nodes for roughly 700B class inference, [11:31] though almost everyone testing this in public only ever changed to. The catch is that the half-meter QSFP112 cable was back-ordered for months, priced between $159 and $229. [11:44] That left people with a $9,000 cluster waiting on a cable, and the scaling disappoints. Neo Zool No GPTOS 120B goes from 58.82 tokens per second to 75.96, which is nowhere near double for double the hardware. [12:03] Two rebuttals deserve airtime because they're the strongest defense this box has. The first is concurrency. currency. Push it to 256 simultaneous requests, and it pushes about 862 tokens per second [12:16] in aggregate. If you're serving a small team instead of chatting solo, that changes the math completely. The second comes from Sebastian Roshka, who owns an M4 Pro Mini, and still defends the spark on sustained fine-tuning. His point is that NPS and macOS is still unstable, [12:33] and fine-tuning often fails to converge, while CUDA and PyTorch just work. NVIDIA also shared the CES 2026 software update, claiming up to two and a half times in select workloads [12:45] through Tensor RT-LRM, MDSP4, and Eagle 3 speculative decoding. Select is doing real work in that sentence. And here's my favorite result in this entire series. [12:59] The team at XO Labs did something obvious in hindsight and slightly beautiful in practice. Remember the split? Pre-fill is compute-bound, and the Spark wins it. Decode is bandwidth-bound, and Apple wins that. [13:12] So they wired a DGX Spark to the Mac Studio MC Ultra and gave each machine the hack it's actually good at. NVIDIA reads the prompt. Apple generates the answer. On Lama, 3.18b was an 8192 token prompt. [13:29] that hybrid ran 2.8 times faster end-to-end than either box on its own. Regency dropped from 6.42 seconds to 2.32. These are two companies that spend all year taking shots at each other, [13:44] and their hardware turns out to be complementary rather than competing. If you already own a Mac and you're eyeing the Spark, that combination is a much more interesting reason to buy one than any benchmark row I just read you. [13:56] The rumor layer around this box is a swap. So I'm rating each item on screen as confirmed, credible, or unverified. Number one is confirmed. RTX Spark, codename N1X, was announced at Computex and GPC Taipei on June 1st, 2026, [14:13] and it ships this fall from Asus, Dell, HP, Lenovo, MSI, and has a Microsoft Surface. It carries up to 128GB of memory, around a petaflop of FP4, [14:26] and the same 20-core Grace ARM CPU with 6,144 CUDA cores. It runs 120 B-models with context up to a million tokens, and Adobe is rebuilding Photoshop and Premiere for it. [14:40] That is the machine most of you are actually waiting for. Number two is credible, not confirmed. RTX Spark pricing is being reported around $1799 for the N1 [14:52] and around $2899 for the N1X. sourced from Morgan Stanley's supply chain checks via Max Weinbach PC World separately cites a to down No part of that is an NVIDIA MSRP so treat it as a well guess Number three is confirmed but constantly misread [15:13] DGX station for Windows is real. Built on GB300, Grace Blackwell Ultra with up to 748 gigabytes, and Huang showed that same 748 gigabyte workstation on stage for Q4 2026. [15:27] It is a separate high-end product and not a Spark replacement. Number four is Roadmap, which I call credible as a plan and worthless as a purchase decision. Vera Rubin Spark in 2027 to 2028 and Rosa Feynman Spark in 2029 to 2030. [15:44] Both come from Computex's 2026 slides. Slides are not shipping products, and I've watched that exact gap in nine months once already in this very story. While we're here, the Vera Rubin Data Center platform is confirmed and in full production [15:59] with 50 petaflops NDSP4, 288 gigabytes of HBM4, and 22 terabytes per second. It's a rack, not a desk. Don't let anyone blend those two things together. [16:11] Number five is the one I need to say out loud. There is no BGX Spark 2. It hasn't been announced, and it isn't on a roadmap slide. Every post you've seen about it is somebody's guess with a render attached. [16:24] If you're holding off on a purchase because a sequel is coming in three months, you're waiting on something that does not exist yet. The next real thing in this lineup is RTX Spark this fall. And that one has a date. [16:36] Step one. And this is the buy case. If your work lives inside CUDA or you fine tune regularly, get a GD10 box. The same goes if you're serving a small team instead of chatting alone. [16:49] But skip the Founders Edition because the OEM twins do the same job for less. MSI Edge Expert, Asus Ascend GX10, and Dell Pro Max start around $29.99 to $30.99 with a full terabyte of storage. [17:05] That's roughly $1,600 under the NVIDIA branded unit. For performance differences, you will struggle to notice. Paying extra for the Gold Box is the single most avoidable mistake in this category. Step two is the weight case. [17:19] and it covers most of you. If you're a pro-sumer who wants a fast local machine that also happens to be a normal computer, wait. RTX Spark lands this fall with real OEM support and reported pricing well under the Founders Edition. [17:33] Waiting a few weeks for that beats spending $4,700 today. On the AMD side, Gorgon Halo is confirmed for Q3 2026 through Asus, HP, and Lenovo. It's built on Ryzen AI Max Plus Pro 495. [17:49] with 192 gigabytes of LPDDR5 X8533 and up to 160 of it addressable as DRAM. And if someone tells you Medusa Halo with LPDDR6 at 460 to 691 gigabytes per second is coming, [18:07] that's weaker source for 2027 or 2028. So fire it under speculation. Step three is the skip case. If you just want a private model in a chat window with no cloud involved, [18:21] buy a strict Halo box and keep the difference. You're giving up 5x pre-sale and getting decode within 13% at about $11.70 per gigabyte instead of $36.70. AMD's own Ryzen AI Halo dev kit at $39.99, [18:37] a micro-center exclusive with pickup around July 10th. It's a hard sell when third-party boxes do it for half. Here's one honest note on Apple. The M4 Pro Mini can't run the 120B model at all, now that 48 gigs is the ceiling. [18:53] So it's a great small model machine and not a big model machine. One last thread ties the trilogy together. Arnie Hanoon, the co-creator of MLX, left Apple on February 27th, 2026 and joined Anthropik. [19:07] And he's the same person whose micro-benchmarks tend to spark at roughly 60 teraflops dense BS-16s. The people measuring this hardware honestly keep showing up in all three episodes, which tells you where to look when marketing and reality disagree. [19:22] My verdict after digging into all three of them is that the Spark is a genuinely good machine sold at a bad price to the wrong audience. As a CUDA development box, it's the real thing, and I defended there all day. [19:36] As the personal AI supercomputer from that keynote, it was overtaken by cheaper hardware before it even shipped. The one number is what I want you to keep, because bandwidth divided by model size still predicts your decode speed on machines that don't exist yet. [19:52] And if you'd rather not buy any of this, the platform I showed you earlier does the same job from a browser, no extra hardware required. Thanks for watching, and I'll see you in the next one.