How I Killed My ChatGPT Subscription
45sThe opening hook promises a money-saving hack, immediately grabbing attention from viewers tired of subscription costs.
โถ Play Clip"The title promises a subscription killer, but the video delivers a nuanced, conditional review that admits the Mac Mini isn't the best for all use cases."
This video reviews the Mac Mini M4 Pro for local AI inference, comparing it to the AMD Strix Halo box and discrete GPU rigs. It covers real-world performance, power efficiency, pricing, and the impact of recent Apple silicon roadmap changes, including a rumored skip to M7. The verdict is conditional: it's the cleanest, most energy-efficient inference box, but only if you have accurate expectations about its memory and bandwidth.
The video's premise is replacing a $20/month ChatGPT subscription with a Mac Mini M4 Pro for local AI inference, offering privacy and ownership of the model.
Apple's AI hardware story has shifted: a price hike on June 25th, a killed memory configuration, and Bloomberg leaks about the M6/M7 roadmap change the buying calculus.
The Mac Mini M4 Pro with 48GB of unified memory costs around $2,039. Anything below 48GB is too compromised for local AI.
The often-cited 550 GB/s memory bandwidth is for the M4 Max, not the M4 Pro. The M4 Pro in the Mac Mini delivers 273 GB/s. This is a critical correction for buyers.
Under inference load, the Mac Mini M4 Pro draws 30-65W, compared to 45-140W for the AMD Strix Halo box and 700-800W for an RTX tower. This makes it ideal for 24/7 always-on inference.
On June 25, 2026, Apple raised Mac Mini M4 Pro prices by $200 and killed the 64GB memory option. Tim Cook called it a '100-year flood' in DRAM supply, and Apple's stock dropped 6%.
The Mac Mini M4 Pro costs $42 per GB, far more than the AMD Strix Halo box at $12/GB, the Framework desktop at $18/GB, and the DGX Spark at $37/GB. Apple loses on pure memory value.
On 48GB, 7-8B models run at 20-30 tokens/s, 14B at 10-20, 30-32B dense at 12-18, and 70B models at a slow 3-5 tokens/s. The Qwen 330B A3B MoE runs at 40-75 tokens/s, and GPT-OSS 20B at 34 tokens/s. GPT-OSS 120B does not fit in 48GB.
Deb Brooks, Apple's silicon product marketing lead, describes the Mac Mini as ideal for 24/7 isolated inference, and frames AI as a 'whole chip problem, not a GPU one.'
Bloomberg's Mark Gurman reports Apple is skipping M6 Pro/Max/Ultra, fast-tracking to M7 (target H1 2027) with up to 1.5TB unified memory. The M5 Mini is rumored for late 2026.
With the Llama 0.19 MLX backend, Qwen 3.5 35B A3B decode jumped 93% on M5 Max (58 to 112 tokens/s). M5 Ultra is internally tested at 768GB unified memory.
A dozen AI researchers, including MLX co-creator Aani Hanan, have left Apple. Apple pays Google ~$1B/year for Gemini to power Siri, and fine-tuning via MPS is still unstable.
Buy the M4 Pro Mac Mini now for 24/7 inference with 7B-32B models. Wait 90 days for the M5 Mini if you care about throughput. Skip it if 70B+ models are your primary need; consider the M4 Max or AMD Strix Halo instead.
The Mac Mini M4 Pro is the cleanest, most energy-efficient inference box you can buy, but only if you go in with accurate expectations about its 48GB memory and 273 GB/s bandwidth. The Apple tax is real, but so is the software advantage.
What is the actual memory bandwidth of the Mac Mini M4 Pro?
273 GB/s (not the often-cited 550 GB/s, which is for the M4 Max).
01:51
What is the price per gigabyte of the Mac Mini M4 Pro compared to the AMD Strix Halo box?
Mac Mini M4 Pro is $42/GB, while the AMD Strix Halo box is $12/GB.
06:46
What happened to the 64GB memory option for the Mac Mini M4 Pro?
Apple killed the 64GB memory option entirely on June 25, 2026.
05:39
What is the power draw of the Mac Mini M4 Pro under inference load?
30-65 watts.
02:47
Which model does NOT fit in 48GB of unified memory?
GPT-OSS 120B (requires at least 60GB).
09:07
What is the rumored memory target for the M7 Ultra?
1.5 terabytes of unified memory.
11:33
What is the main limitation of fine-tuning on Apple silicon via MPS?
It is still unstable.
14:36
Who left Apple to join Anthropic, and what did they co-create?
Aani Hanan, who co-created MLX.
13:41
What is the approximate cost of the Mac Mini M4 Pro with 48GB?
Around $2,039.
01:23
Bandwidth Correction
Corrects a common misconception that could mislead buyers about the Mac Mini's actual performance.
01:51Power Efficiency Advantage
Highlights the Mac Mini's key strength: 30-65W draw makes it ideal for 24/7 always-on inference.
02:47Price per GB Reality
Provides a clear, honest comparison showing Apple loses on pure memory value.
06:46Roadmap Leak
Reveals Apple's strategic shift to skip M6 and fast-track M7, which changes the buying decision.
10:46Talent Exodus
Points out the departure of key AI researchers, undermining Apple's software advantage narrative.
13:28[00:02] box and killed a chat GPT subscription with it. Today I'm swapping it for this, the Mac Mini M4 Pro. Same promise, different silicon, and honestly, a completely different verdict. If you watched the AMD Strix Halo video, you
[00:15] already know the pitch. You stop paying $20 a month per AI subscription, you run the model locally, and you own the inference. Nothing you type ever leaves your machine. That video landed well. The most common follow-up question in
[00:29] the comments was exactly what I expected. Okay, but what about the Mac Mini? That's actually a fair question to ask. A lot of users already own a Mac. Some of you run agentic workflows that need something always on, isolated from
[00:43] your daily driver, drawing 30 watts instead of 700. And Apple has been talking a very big game about Apple silicon being purpose-built for AI. So, let's actually check the receipts. The timing matters here, too. This isn't
[00:56] just a spec video. The ground has been moving under Apple's AI hardware story all year. There was a significant price hike on June 25th. Apple quietly killed a memory configuration, and there are leaks out of Bloomberg that, if
[01:11] accurate, change the entire calculus of when to buy. We're going to cover all of it. Let's start with what you're actually buying. The Mac Mini M4 Pro actually buying. The Mac Mini M4 Pro with 48 GB of unified memory runs you
[01:23] somewhere around $2,039 right now. That's worth double-checking on Apple's site before you buy, since pricing has been moving fast this year. That's the config worth talking about for local AI. Anything below 48 GB, and
[01:38] you're making too many compromises on which models actually fit. Now, here's the number I need you to write down. The figure you'll see everywhere online is figure you'll see everywhere online is 550 GB per second of memory bandwidth.
[01:51] That number is wrong for this machine. That 550 That 550 GB per second figure. 546 to be precise. That belongs to the M4 Max. The M4 Pro, which is what lives
[02:04] Max. The M4 Pro, which is what lives inside the Mac mini, delivers 273 GB/s. everyone quotes. If someone is benchmarking a Mac mini and citing 550 GB/s, they either have the wrong chip or the wrong box. Compare that to the AMD
[02:20] Strix Halo box from our last video. The GMK Tech Evo X2 gives you 128 GB of unified memory. On paper, its bandwidth looks close to this machine's.
[02:32] 256 versus 273. But real world, Strix Halo measured closer to 122 GB/s in our testing. This machine actually beats that in practice, and it comes with a completely different
[02:47] software story, which we'll get to. The power draw is where Apple does something genuinely impressive. Under a real inference load, not just sitting idle, the Mac mini M4 Pro pulls between 30 and
[03:00] 65 W. Your AMD Strix Halo box draws somewhere between 45 and 140 W. That's actually efficient, too. But an RTX tower doing the same workload runs tower doing the same workload runs closer to 700 to 800 W. Mac mini still
[03:15] edges out Strix Halo here, but the real contrast is against discrete GPU rigs, not against AMD's box. I researched the Mac mini M4 Pro comparison for this exact video inside AI Master. The AI producer agent structured the research
[03:30] brief. Then the script writer agent drafted the sections you're watching right now. I built AI Master for my own production pipeline first. My team runs that's actually packed inside the platform. You get every top LLM in one
[03:46] window, and you're paying less per token than going direct. We built our own studio right into it. So, image, audio, and video generation all run from the and video generation all run from the same login. Nano Banana Pro, Seed uns 2,
[03:59] and VO3.1 all run without a separate subscription. You can build consistent characters that carry across your content, and you can monetize those characters directly on the platform. There's a built-in sharing system, so
[04:12] your work reaches the community. Over 12,000 creators are already inside. This is where users find collaborators, share what is working, and help each other ship faster. The counter on the page updates live. Then, there's the academy
[04:27] behind all of it. Over 200 lessons and close to 30 hours take you from complete beginner to someone actually shipping content. That's the exact path I used to build this channel, and students keep using it to launch or scale their own.
[04:41] Here is what real users are saying about the platform, and the case studies show exactly what people shipped with this pipeline. Real projects and real results, not marketing copy. There is an annual plan with a real discount, and
[04:55] the promo code from this video stacks on top of that. Every purchase has a 7-day money-back guarantee. You try it, and if it's not right for you, you get your money back. That's zero risk. Here's how you get access. Go to the link in the
[05:10] description and hit buy. Select the annual plan and enter the promo code and you'll get a confirmation email. Then, log in, and you're inside the same workspace I just showed you. The whole setup takes under 3 minutes. The link is
[05:25] waiting in the description. Okay, back to the box itself. Because what the price hike actually did to this machine is where things get interesting. Let's talk about what happened on June 25th, 2026, because this is important context
[05:39] for whether you should buy right now. Apple raised the price of every Mac mini Apple raised the price of every Mac mini M4 Pro configuration by $200. The base M4 Pro configuration by $200. The base config went from $1,399
[05:52] config went from $1,399 to $1,599 want for serious local AI work now sits at roughly 2,039.
[06:04] And on the same day, Apple killed the 64 GB memory option entirely. It's gone. If you were waiting for that config, the ship has sailed. Tim Cook addressed the hike by calling it a 100-year flood in DRAM supply. Micron has been even more
[06:20] direct. They say supply stays tight beyond 2027. Apple's stock dropped 6% on the day of the announcement, their worst single-day drop since April 2025. The market clearly didn't love that announcement. Neither did Mac Studio
[06:34] buyers, who also lost their 256 and 512 GB storage tier options in the same move. So, where does the Mac mini land on the dollar per gigabyte metric?
[06:46] Here's the honest table on price per gigabyte. The GMK Tech Evo X2, the AMD Strix Halo box, comes in at roughly $12 per gigabyte. The Framework desktop sits per gigabyte. The Framework desktop sits around $18. The DGX Spark lands at about
[07:01] around $18. The DGX Spark lands at about $37. And the Mac mini M4 Pro, $42 per gigabyte. The Mac Studio M3 Ultra pushes that even further to $55. Apple loses this comparison badly on pure memory value. If getting the most gigabytes for
[07:16] your dollar is your primary goal, you should be looking at the AMD boxes we covered last time. I want to be honest about that. Before we get to where Apple actually earns its premium. All right, let's get to the model fit reality,
[07:28] because this is is where most reviews either get it wrong or avoid it either get it wrong or avoid it entirely. On 48 GB of unified memory entirely. On 48 GB of unified memory with the M4 Pro's 273 GB per second
[07:40] bandwidth, here's what you're working with. 7 to 8 billion parameter models run at 20 to 30 tokens per second. That's fast and responsive, genuinely useful for everyday tasks. 14 billion
[07:53] parameter models drop to 10 to 20 tokens per second, still comfortable. 30 to 32 billion dense models land around 12 to 18 tokens per second. That starts to feel like inference rather than conversation, but it's workable. Here's
[08:08] where this actually gets interesting. The Qwen 330B A3B is a mixture of experts model and it runs at 40 to 75 tokens per second on this machine. That's the fastest feeling model you'll run on a Mac mini and it punches well
[08:23] above its weight on coding and reasoning tasks. GPT-OSS 20B comes in around 34 tokens per second. Both of those feel genuinely fast and real use. Then, we
[08:36] hit the real wall. 70 billion parameter models on 48 GB run at 3 to 5 tokens per second. That's not unusable, but it's slow enough that you'll feel it on every single response. It's the difference between a conversation and watching
[08:51] paint dry. If 70B is your primary use case, you need to be looking at the M4 case, you need to be looking at the M4 Max with 128 GB where 70B runs at a much more usable 12 to 13 tokens per second. And GPT-OSS 120B
[09:07] And GPT-OSS 120B simply does not fit in 48 GB. You need at least 60 GB of unified memory to load it. This is where I have to give AMD the win directly. In our last video, the Strix Halo box ran GPT-OSS 120B at 34
[09:23] tokens per second. The Mac mini can't run that model at all. If 120B class inference is your primary need, the AMD box from our last video is the right call, not this one. Here's the use case Apple's own team keeps coming back to,
[09:37] and honestly, it's the most compelling argument for the Mac mini specifically. Deb Brooks leads Apple's silicon product marketing. He gave an interview to the deep view in July 2026, right before WWDC. He described what users actually
[09:52] want, a system that's isolated from their primary machine and capable of their primary machine and capable of running 24 hours a day, 7 days a week. A Mac mini, in his words, is an amazing system for exactly that. He also said
[10:05] something that stuck with me. AI, in his words, is a whole chip problem, not a GPU one. That framing matters because it's Apple's answer to the benchmark wars. The Mac mini actually earns that framing at 30 to 65 watts. You leave it
[10:20] running an agent loop overnight. It handles tool calls and writes back to the database. It processes a queue of tasks while you sleep. When you wake up, it's done, and it costs you almost nothing in electricity. Try that with an
[10:33] nothing in electricity. Try that with an RTX tower drawing 700 to 800 watts, and the math starts to hurt fast. This is the section that actually changes the buying decision, so pay attention here. Mark Gurman at Bloomberg has reported
[10:46] something that might be the most significant Apple silicon news in years. significant Apple silicon news in years. Apple is skipping M6 Pro, M6 Max, and M6 Ultra entirely. That would be a first in the Apple silicon era. The only M6 chip
[11:01] shipping this year is the base M6, which lands in late 2026 at around 200 GB/s. If you were expecting an M6 Mac mini with serious memory bandwidth, there isn't one coming. Instead, Apple is fast-tracking directly to M7. Gurman
[11:18] reports the target is first half of 2027 for M7, with M7 Ultra in 2028. The internal framing at Apple is apparently approaching Nvidia Blackwell, and the approaching Nvidia Blackwell, and the memory target for M7 Ultra is 1.5
[11:33] terabytes of unified memory. That's not a typo. That's the gap Apple is trying to close with a single generation jump. Now, for the near term, the M5 Mini is Now, for the near term, the M5 Mini is rumored for October or November 2026.
[11:47] Still unconfirmed, but close. That's according to Gurman. The M5 chip already exists in MacBook Pros. The M5 Max delivers up to 128 delivers up to 128 GB of unified memory 614
[12:01] GB per second of bandwidth. More importantly, a Llama 0.19, which dropped in March 2026, added an MLX backend. On M5 Max, the Qwen 3.5 35
[12:15] MLX backend. On M5 Max, the Qwen 3.5 35 BA3B went from 58 to 112 tokens per second decode. That's a 93% jump. Prefill went from 1,154 to 1,810 tokens per second. Those are real
[12:30] numbers, even with a caveat that the comparison isn't perfectly apples to apples across quantization formats. And at the top of the Mac Studio line, Apple has internally tested the M5 Ultra configuration at 768
[12:45] GB of unified memory. That number is significant. It puts Apple-branded hardware in the same conversation a small cluster builds for a single box sitting on a desk pulling maybe 200 W. So, the road map picture looks like
[13:01] this. Buy an M4 Pro Mac Mini today, and you're buying a machine that's two generations behind the performance curve by the time M7 lands. The M5 Mini is close, even if the exact date isn't locked yet. M7 is where Apple silicon
[13:16] gets genuinely competitive with Blackwell tier hardware. The question isn't whether a better machine is coming. It's whether you can afford to wait, and whether of machine you'd buy today actually serves your needs in the
[13:28] meantime. All right, here's the part where I push back on the Apple silicon optimism because there are a few things happening underneath the surface that don't get enough coverage. In February 2026, Aani Hanan left Apple for
[13:41] Anthropic. He co-created MLX, the framework that makes Apple silicon so compelling for local inference, and he wasn't the only one who left. Around a dozen AI researchers have departed over the same period including Apple's head
[13:55] of foundation models. The people who built the software stack that makes the built the software stack that makes the Mac mini a real AI inference machine are leaving. There's also a detail that undercuts Apple's AI narrative in a
[14:07] specific way. Apple is paying Google approximately $1 billion a year for a custom version of Gemini to power Siri. They reportedly passed on Claude because Anthropic wanted around $1.5 billion. So, the company marketing Apple silicon
[14:22] as the future of on-device AI is licensing a cloud model from Google to run its flagship AI feature. That tension is worth holding in your head when you're evaluating the platform long-term. On the software side,
[14:36] researcher Sebastian Raschka has done a lot of work with Apple's MPS backend. He's been clear that fine-tuning on Apple silicon via MPS is still unstable. If you're planning to fine-tune models locally, the Mac mini isn't your box
[14:49] right now. It's an inference-only machine for the foreseeable future. None of this makes the Mac mini a bad machine. But, that's the honest picture right now. The software advantage Apple has built is real, and it's also fragile
[15:02] in ways the company's marketing doesn't acknowledge. So, here's exactly where I land on this. I've got three categories and three clear calls to make. Buy the and three clear calls to make. Buy the M4 Pro Mac mini now if you need a 24/7
[15:15] always-on inference box today. That means you're running 7B to 32B models or Maui models in the 30B range. Your main use case is agentic workflows or private inference. The software stack is mature. A llama and LM Studio just work out of
[15:31] the box and 30 to 65 watts of draw is hard to argue with for machine that never sleeps. Just go in knowing you're buying a 273 GB/s machine, not a 550
[15:44] GB/s one. Wait 90 days if you're not in a rush and you primarily care about throughput. The M5 Mac mini is rumored for this fall, still unconfirmed but close. The A llama 0.19 MLX numbers on M5 Max are legitimately
[16:00] exciting. If that efficiency translates to the M5 mini, the performance per dollar picture changes meaningfully. Skip the Mac mini entirely if 70B or larger models are your primary use case and you need speed, not just
[16:15] and you need speed, not just compatibility. The M4 Max with 128 GB is a better fit for 70B. The AMD Strix Halo box from our last video is the right box from our last video is the right call if you need 120B class models and
[16:29] you're willing to deal with a Rock ham setup. And if you can wait until 2027, setup. And if you can wait until 2027, the M7 with 1.5 TB unified memory is where Apple silicon stops making apologies and starts making the
[16:42] Blackwell argument on its own terms. Last time I recommended the AMD box for raw model capacity at the best dollar per GB on the market. This time, the verdict is more conditional. The Mac mini earns its place as the cleanest,
[16:55] most energy efficient inference box you can buy. That's as long as you go in with accurate expectations about what 48 GB and 273 GB per second actually means.
[17:07] The Apple tax here is genuinely real, so is Apple's software advantage, to be fair. You're paying a premium for both of those things. If this video helped you figure out your local AI setup, hit subscribe. [music] We do this kind of
[17:21] honest hardware breakdown every time a new box enters the conversation. [music] And if you want to learn and practice AI tools the way we do at AI Master, the link is in the description. See you in the next one.
โก Saved you 0h 17m reading this? Transcribe any YouTube video for free โ no signup needed.