At Computex 2026, Jensen Huang put a new chip on a slide and called it the start of the PC's next chapter. RTX Spark pairs a 20-core Grace CPU with a Blackwell GPU and up to 128 GB of unified memory, soldered into laptops and compact desktops from Dell, HP, Microsoft, Lenovo, MSI, and ASUS. The headline number: 1 petaflop of AI compute, enough to run a 120-billion-parameter LLM locally with a million tokens of context. The catch: it ships in fall 2026. Until then, the local-AI desk you want already exists, and it doesn't have an NVIDIA logo on it. It's a Mac Studio or a Strix Halo mini PC.
Apple Mac Studio (M4 Max, 64 GB unified)
M4 Max 16C/40C GPU · 64 GB unified · 1 TB SSD · 546 GB/s
Apple's unified memory architecture treats GPU and CPU as one memory pool. At 546 GB/s on M4 Max, and 800 GB/s on the M3 Ultra config, that pool holds models too big for any consumer NVIDIA card's VRAM. A 64 GB Mac Studio runs Llama 3.3 70B at Q4 quantization at conversational speed; the 128 GB and 192 GB configs run it at higher precision.
- Highest memory bandwidth per dollar in the local-AI category
- Silent under load. Sits on a desk and disappears
- Ollama, llama.cpp, MLX, and LM Studio all have first-class M-series support
- The default recommendation in the r/LocalLLaMA community for two years running
- macOS only. No native CUDA, no Windows AI tooling
- Storage upgrades cost more than the rest of the machine combined
- For image and video diffusion workloads, NVIDIA's CUDA ecosystem still wins on throughput
The Mac Studio is the machine local-AI hobbyists tell new buyers to start with. You can buy one today, plug it in, install Ollama, and have a 70B model talking back to you in twenty minutes. The 64 GB config handles models up to about 50 GB on disk; if you want Llama 3.3 70B at higher precision or a multi-model setup, jump to 128 GB. The M3 Ultra version goes to 192 GB but costs more than two of these. Apple's M5 chip is rumored for late 2026, so this is the right time to buy an M4 Max and the wrong time to buy the Ultra.
GMKtec EVO-X2 AI Mini PC (128 GB)
Ryzen AI Max+ 395 · 128 GB LPDDR5X-8000 · 2 TB SSD · 54-140 W
If RTX Spark is the future of unified-memory AI PCs, the GMKtec EVO-X2 is what that future looks like with an AMD logo on it and a shipping date of last quarter. The Ryzen AI Max+ 395 pairs 16 Zen 5 cores with a 40-CU Radeon GPU and 128 GB of LPDDR5X-8000 unified memory soldered to the board. Same architectural bet NVIDIA just announced. The EVO-X2 packages it in a 35 dB box you can fit on a bookshelf.
- 128 GB unified memory at a price RTX Spark laptops probably won't match
- Quad 8K display output, USB4, WiFi 7
- 35 dB at idle in Quiet Mode; sustained 50+ TOPS XDNA 2 NPU for agentic loads
- Ships now, no fall-2026 wait
- Strix Halo memory bandwidth sits close to RTX Spark's reported figure, well below Mac Studio Ultra's 800 GB/s
- ROCm on Windows is less mature than CUDA for AI tooling
- Performance Mode at 140 W spins the fans audibly louder than the 35 dB Quiet Mode rating
The EVO-X2 is the answer to 'what if I want RTX Spark but I want it now, and I don't need a laptop.' You get the unified-memory architecture, you get the 128 GB ceiling, you skip the wait. The Ryzen AI Max+ 395 is the same chip Framework puts in its Desktop ($1,999, ships now from frame.work direct, no Amazon SKU), so this category is real, not vapor. ROCm is improving; for local LLM inference via llama.cpp or LM Studio's ROCm path, the EVO-X2 runs 70B models. If you want CUDA specifically, this isn't the machine.
Beelink GTR9 Pro (128 GB, dual 10 GbE)
Ryzen AI Max+ 395 · 128 GB LPDDR5X · 2 TB SSD · dual 10 GbE · USB4 ×2
The Strix Halo platform's strongest argument for itself isn't raw AI throughput. It's that you can have 128 GB of unified memory in a mini PC under $2,000. The Beelink GTR9 Pro takes the same Ryzen AI Max+ 395 chip as the GMKtec and adds dual 10 GbE Ethernet. Every local AI lab eventually wants fast NAS access for models and datasets, so the 10 GbE is more than a spec sheet flourish.
- Cheapest 128 GB unified-memory machine on Amazon (~$1,899 as of June 2026)
- Dual 10 GbE for NAS/SAN AI workflows
- Beelink markets DeepSeek 70B support out of the box
- Two USB4 ports (40 Gbps) for external storage or eGPU experiments
- White cube aesthetic. Looks like a smart speaker on a desk
- Beelink's first revision of a new chip historically ships with BIOS quirks that need a firmware update or two to settle
- Same ROCm-vs-CUDA tradeoff as the GMKtec
The GTR9 Pro is the pick if you've already decided on Strix Halo and your top criterion is price-per-GB-of-unified-memory. It's $300 below the GMKtec EVO-X2 for the same 128 GB / 2 TB config. The dual 10 GbE is the differentiator: if you run a local AI lab with a Synology, a TrueNAS box, or a second machine for batch processing, 10 GbE between them changes what's possible. Single-user inference on the GTR9 Pro is identical to the EVO-X2 because they're the same silicon.
Apple Mac Mini (M4 Pro, 64 GB unified, 10 GbE)
M4 Pro 14C/20C GPU · 64 GB unified · 2 TB SSD · 10 GbE · 273 GB/s
The cheapest way into the 64-GB-unified-memory club. Same M4 Pro chip Apple ships in its higher-end MacBook Pro, in a 5-inch silver puck that costs about half what a Mac Studio does for the same memory tier. It won't run 120B models. It runs Llama 3.3 70B at Q4, Mistral Small 3, and most quantized 30B reasoning models faster than first-time buyers expect.
- Sub-$2,000 entry into 64 GB unified memory territory
- Same software ecosystem as Mac Studio (Ollama, MLX, LM Studio)
- 5-inch footprint sits behind a monitor, drawing under 30 W at chat workloads
- 10 GbE built in, matching the Beelink GTR9 Pro on networking
- 273 GB/s memory bandwidth is half the M4 Max and a third of the M3 Ultra
- 64 GB is the chip ceiling; no upgrade path to 128 GB without moving to Studio
- Apple's storage tax is real (2 TB is a $600 upgrade over base 512 GB)
If your goal is to learn local AI without Mac Studio money, the M4 Pro Mac Mini is the right entry point. The 64 GB unified config runs every small-to-medium model the community releases (Llama 3.3 70B at Q4, Phi-4, Mistral Small 3, Qwen 3 32B) and does so silently. The 273 GB/s bandwidth ceiling slows token generation on bigger models, so it suits one-at-a-time chat better than batch agentic workflows. If you outgrow it in twelve months, the upgrade path to a Mac Studio is part of the plan.
The numbers.
| Mac Studio M4 Max | GMKtec EVO-X2 | Beelink GTR9 Pro | Mac Mini M4 Pro | |
|---|---|---|---|---|
| Architecture | Apple unified | AMD unified | AMD unified | Apple unified |
| Memory for AI | 64 GB | 128 GB | 128 GB | 64 GB |
| Memory bandwidth | 546 GB/s | ~256 GB/s | ~256 GB/s | 273 GB/s |
| Native CUDA | No | No | No | No |
| OS | macOS | Windows / Linux | Windows / Linux | macOS |
| Power at AI load | ~80 W | 85-140 W | 85-140 W | ~30 W |
| Form factor | Desktop | Mini PC | Mini PC | Mini desktop |
| Street price | ~$2,499 | ~$2,199 | ~$1,899 | ~$2,000 |
Other strong options.
NVIDIA DGX Spark (GB10 Grace Blackwell)
NVIDIA's reference local-AI workstation, already shipping at ~$3,000. GB10 Grace Blackwell Superchip with 128 GB of unified memory, sold by NVIDIA, ASUS, Dell, and HP direct (not on Amazon). Same Spark family name as the new RTX Spark consumer chip, different target buyer: data scientists who want NVIDIA's stack on a desk now, rather than Windows AI agent users waiting for fall. Editorial mention rather than affiliate pick.
Framework Desktop (Ryzen AI Max+ 395)
Framework's first desktop, sold direct from frame.work at $1,999. Same Strix Halo silicon as the GMKtec EVO-X2 and Beelink GTR9 Pro, in Framework's signature open-hardware chassis with full repairability scores. Worth knowing about for the open-platform crowd that values long-term parts availability over price. Currently no Amazon SKU.
The buying guide.
Pick I, the LLM enthusiast
You mostly want to run 70B-class LLMs locally and you don't need Windows. Buy the Mac Studio M4 Max. It's been the default recommendation in the r/LocalLLaMA community for two years and stays the right call until M5 Ultra arrives in late 2026. The 64 GB config we list runs Llama 3.3 70B at Q4 conversationally; step up to 128 GB if you want higher precision or simultaneous models loaded.
Pick II, the Windows-native AI early adopter
You want unified-memory AI on Windows or Linux right now, not in fall 2026 when RTX Spark ships. Buy the GMKtec EVO-X2 or Beelink GTR9 Pro. Identical silicon, identical inference performance. The Beelink is $300 cheaper and has dual 10 GbE LAN; the GMKtec has more polished firmware and a sturdier chassis. Either way you're betting on AMD's ROCm software stack maturing through 2026, which it is.
Pick III, the curious beginner
Your budget is under $2,000 and you want into the local AI hobby without a Mac Studio commitment. Buy the Mac Mini M4 Pro with 64 GB unified. The 273 GB/s bandwidth ceiling caps growth, but at this price the eventual upgrade to a Mac Studio is part of the plan. Same software ecosystem on the way up: no relearning when you move to the bigger machine.
For diffusion workloads, build your own
None of the four picks above suit image and video diffusion (Stable Diffusion 3, FLUX, Wan 2.2, Hunyuan Video). CUDA on an RTX 5090 still wins on those workloads per dollar, but there's no Amazon prebuild we'd recommend on its merits, the brand options at this tier are uneven. Build your own tower around an RTX 5090 instead. The GPU Buying Guide below covers current cards and pairing; the CPU and motherboard guides cover the rest of the parts list.
What to skip while waiting for RTX Spark
Skip any unified-memory machine under 64 GB; you'll outgrow it before RTX Spark ships and you'll regret the spend. Skip RTX 4070 / 4080 / 5070 / 5080 cards for LLM-focused builds: they're slower at inference per dollar than a Mac Mini and the VRAM ceiling is no better. Skip Qualcomm Snapdragon X Elite laptops marketed as 'AI PCs'; the 45 TOPS NPU is for Windows Recall, not for running 70B local models. Skip no-name prebuilt RTX 5090 towers on Amazon; the assembly quality is uneven and the brand support is thin.
FAQ.
Three different machines whose names overlap on purpose. DGX Spark (already shipping) is a $3,000 desktop AI workstation built around NVIDIA's GB10 Grace Blackwell Superchip. Sold by NVIDIA, ASUS, Dell, and HP direct. Single-user, 128 GB unified memory. Targets data scientists. RTX Spark (announced for fall 2026) is the consumer-PC version of the same idea: 20-core Grace CPU plus Blackwell GPU plus up to 128 GB unified, in laptops and compact desktops from Dell, HP, Lenovo, Microsoft, ASUS, MSI. Targets creators, developers, and Windows AI agent users. DGX Station for Windows (also announced at Computex 2026) is an enterprise deskside supercomputer for trillion-parameter models, bigger than both Sparks, targets dev teams. NVIDIA reused 'Spark' twice for different audiences. The branding is confusing on purpose; the products are real.
For pure token-generation speed: yes. Mac Studio Ultra runs at 800 GB/s memory bandwidth. Press estimates put RTX Spark closer to 300 GB/s, which sits in Strix Halo territory. Bandwidth governs how fast tokens stream out of a loaded model. For raw AI compute (training, fine-tuning, image and video diffusion), RTX Spark's Blackwell tensor cores at FP4 will beat Apple Silicon on per-operation throughput. The answer depends on workload. Inference-heavy: Mac stays ahead. Compute-heavy: NVIDIA pulls ahead.
Yes, with caveats. RTX 4090 has 24 GB VRAM. That fits any 13B-class model at full precision, 30B-class at Q4, and Mixtral 8x7B at Q4. Llama 3.3 70B fits at Q2 with quality compromises, or at Q4 with partial CPU offloading. Image and video diffusion workloads (Stable Diffusion 3, FLUX, Wan, Hunyuan) all fit cleanly in 24 GB. The 24 GB ceiling does mean newer 70B reasoning models at FP8 or higher won't fit; you quantize aggressively or wait for newer cards. Otherwise: install Ollama or LM Studio, point it at a model file, and you're running local AI tonight.
NVIDIA didn't disclose pricing at the Computex 2026 reveal. Press coverage from Tom's Hardware, VideoCardz, and Yahoo Finance estimates RTX Spark laptops will start above $1,500 and flagship configurations (Microsoft Surface Laptop Ultra, Dell XPS 16 Creator Edition) will push considerably higher. NVIDIA confirmed 30+ laptops and 10+ desktops from six OEMs (Dell, HP, Lenovo, Microsoft, ASUS, MSI) with Acer and GIGABYTE 'to follow.' Compare against today's prices: a comparable 128 GB Strix Halo mini PC is $1,899 (Beelink GTR9 Pro), a 128 GB Mac Studio is around $4,000.
Depends on how much local AI you actually do today. If you're running local LLMs daily for work and you've outgrown your current machine: don't wait, buy one of the four above. Four months is a long time to be running 30B models on a laptop while RTX Spark sits in fall. If you've never run local AI and you're not sure you'd use it weekly: yes, wait. RTX Spark laptops will be widely reviewed within two weeks of launch and you'll have a clearer picture of price-vs-performance against this lineup. Our position is somewhere in between: the curious beginner pick (Mac Mini M4 Pro at $2,000) is the cheapest way to find out if you'd actually use local AI, with a full refund window if you decide you wouldn't.
For most LLM users: the Mac Studio M4 Max.
For unified-memory AI on Windows now: the GMKtec EVO-X2 or Beelink GTR9 Pro.
For the curious beginner: the Mac Mini M4 Pro.
NVIDIA RTX Spark looks like the most interesting Windows PC platform of the decade and it ships four months from now. The four machines above let you start running local 70B models tonight on hardware that already exists. The Mac Studio M4 Max remains the LLM enthusiast's default; the Strix Halo mini PCs (GMKtec EVO-X2 and Beelink GTR9 Pro) are the closest architectural analog to RTX Spark you can buy today; the Mac Mini M4 Pro is the curious beginner's entry point. For image and video diffusion specifically, build your own RTX 5090 tower from the GPU Buying Guide rather than buying an unbranded Amazon prebuild. When RTX Spark laptops land in fall 2026, we'll review one against this lineup and update the rankings. For now: pick by budget and ecosystem.