Updated Oct 10, 2026· 13 min read

Key takeaways

  • Best overall for big models: Apple Mac Studio (M5 Max or M5 Ultra). Up to 614 GB/s on M5 Max and 1.2 TB/s on M5 Ultra, with up to 512GB of memory on the Ultra.
  • Fastest for models that fit in VRAM: custom desktop with a GeForce RTX 5090 (32GB GDDR7, 1,792 GB/s).
  • Best for CUDA developers: NVIDIA DGX Spark (GB10, 128GB, or the new 64GB OEM version).
  • Best x86 unified-memory box: Ryzen AI Max+ 395 mini PCs such as the Framework Desktop, GMKtec EVO-X2 and Minisforum MS-S1 Max.
  • Most memory in a mini PC: Ryzen AI Max+ PRO 495 (Gorgon Halo) systems with 192GB.
  • Best portable AI PC: RTX Spark laptops, led by the Surface Laptop Ultra, shipping from October 16.
  • Best budget AI PC: a mid-range desktop with an RTX 5070 Ti or RTX 5060 Ti 16GB.

The best AI PC for running local LLMs depends on one question: do you want speed on models that fit in a graphics card, or capacity for very large models? For most people who want the fastest experience on 8B to 32B models, a desktop with an RTX 5090 is the pick. If you want to run 100B-plus models on one quiet box, a Mac Studio with M5 Max or M5 Ultra leads on bandwidth, while the NVIDIA DGX Spark and Ryzen AI Max (Strix Halo) mini PCs offer large unified memory with Linux or Windows. Below, we compare seven options, explain what each one actually runs well, and flag the catches that spec sheets leave out.

Verdict at a glance

  • Best overall for big models: Apple Mac Studio (M5 Max or M5 Ultra). Up to 614 GB/s on M5 Max and 1.2 TB/s on M5 Ultra, with up to 512GB of memory on the Ultra.
  • Fastest for models that fit in VRAM: custom desktop with a GeForce RTX 5090 (32GB GDDR7, 1,792 GB/s).
  • Best for CUDA developers: NVIDIA DGX Spark (GB10, 128GB, or the new 64GB OEM version).
  • Best x86 unified-memory box: Ryzen AI Max+ 395 mini PCs such as the Framework Desktop, GMKtec EVO-X2 and Minisforum MS-S1 Max.
  • Most memory in a mini PC: Ryzen AI Max+ PRO 495 (Gorgon Halo) systems with 192GB.
  • Best portable AI PC: RTX Spark laptops, led by the Surface Laptop Ultra, shipping from October 16.
  • Best budget AI PC: a mid-range desktop with an RTX 5070 Ti or RTX 5060 Ti 16GB.

What makes a PC good at local AI

Local LLMs care about three numbers more than anything else:

  • Memory capacity available to the accelerator. This decides which models you can load at all. A 4-bit 8B model needs about 5GB, a 32B model about 19 to 20GB, and a 70B model about 40GB before counting context.
  • Memory bandwidth. This decides how fast text is generated. Token generation streams the active model weights for every token, so tokens per second scales almost directly with bandwidth.
  • Compute and software. This decides how fast long prompts are read (prefill) and which tools work out of the box. NVIDIA’s CUDA has the broadest support; Apple’s MLX and AMD’s ROCm and Vulkan paths keep improving.

Every machine on this list is a different trade between those three. Discrete GPUs have huge bandwidth but limited capacity. Unified-memory machines have huge capacity but less bandwidth. Nothing on the consumer market maximizes all three at once.

Comparison table

AI PC Max memory for AI Bandwidth OS Best at Price tier
Mac Studio M5 Ultra Up to 512GB unified 1.2 TB/s macOS Largest models, fastest unified memory Premium to ultra-premium
Mac Studio M5 Max Up to 128GB unified 614 GB/s macOS 70B-class and MoE models, quiet desk Premium
RTX 5090 desktop 32GB VRAM 1,792 GB/s Windows / Linux Speed on 8B to 32B models, gaming too Premium
NVIDIA DGX Spark 128GB or 64GB unified 273 GB/s Linux (DGX OS) CUDA development, prefill, clustering Premium
Ryzen AI Max+ 395 mini PC 128GB (96GB GPU via BIOS) 256 GB/s Windows / Linux Large MoE models on x86 Upper mid-range to premium
Ryzen AI Max+ PRO 495 mini PC 192GB (up to 160GB GPU) LPDDR5X-8533, 256-bit Windows / Linux Most memory in a mini PC Ultra-premium
RTX Spark laptop 24GB to 128GB unified Up to 300 GB/s Windows on Arm Portable local AI plus gaming Premium
RTX 5070 Ti / 5060 Ti 16GB desktop 16GB VRAM 896 / 448 GB/s Windows / Linux 7B to 14B models, gpt-oss-20b Mid-range

1. Apple Mac Studio (M5 Max and M5 Ultra): best overall for big models

Apple announced the new Mac Studio on August 25, 2026, and most configurations shipped in September. The M5 Max version tops out at 128GB of unified memory and up to 614 GB/s, up from 546 GB/s on the M4 Max. The M5 Ultra joins two M5 Max dies and reaches 1.2 TB/s, which Apple says is 50 percent more bandwidth than the M3 Ultra it replaces. The 256GB and 512GB options are tied to the top M5 Ultra chip, and Apple lists 512GB configurations for late October.

Performance reports back up the specs. MacStories measured a median of 108 tokens per second on a MoE model in MLX on the M5 Ultra, 54 percent more than the M3 Ultra, and prompt processing of roughly 2,000 to 2,800 tokens per second across very long contexts. Apple credits the prefill jump to new Neural Accelerators in each GPU core. Ars Technica reported just over 50 tokens per second on a 27B dense model at 4-bit in LM Studio.

Reasons to pick: the highest bandwidth of any unified-memory machine; up to 512GB holds models that need several GPUs elsewhere; near-silent operation; MLX and LM Studio make setup easy.

Reasons to skip: no CUDA, so some research code and training tools will not run; memory is fixed at purchase; the largest configurations sit at the very top of the price range.

The catch: an RTX 5090 still reads long prompts faster; MacStories found the 5090 well ahead on a 6,000-token prompt, so agent-heavy coding work may feel snappier on NVIDIA if the model fits.

2. Custom RTX 5090 desktop: fastest for models that fit

If your models fit in 32GB, nothing on this list generates text faster. The RTX 5090’s 1,792 GB/s is about 1.78 times the RTX 4090’s bandwidth, and published llama.cpp results put it about 1.5 to 1.8 times ahead of the 4090 in single-user generation. OpenBenchmarking.org recorded 159 tokens per second on an 8B model at Q8_0, and LMSYS measured 205 tokens per second decode on gpt-oss-20b in Ollama.

A 5090 desktop also doubles as the best gaming PC money can buy, which matters if one machine has to do both jobs. Pair it with a Ryzen 9 9950X or similar, 64GB of system RAM for MoE offloading, and a quality 1,000W-class power supply.

Reasons to pick: fastest generation and prefill; full CUDA ecosystem; upgradeable; excellent for gaming.

Reasons to skip: 32GB caps model size; 575W card plus a high-end CPU means a loud, hot tower under load; street prices have run far above launch MSRP in 2026.

The catch: a 70B model at 4-bit is about 40GB, so it will not fit on one 5090 without dropping to a lower-quality quant or offloading to slower system RAM.

3. NVIDIA DGX Spark: best for CUDA developers

The DGX Spark is a small desktop built around NVIDIA’s GB10 chip, with a 20-core Arm CPU, a Blackwell GPU and 128GB of LPDDR5X at 273 GB/s. It ships with NVIDIA’s DGX OS (Linux) and the full CUDA stack, and two units can be linked over the built-in ConnectX-7 200Gb networking. On October 2, 2026, NVIDIA added a 64GB version sold only through partners including Acer, ASUS, Dell, Gigabyte, HP and MSI, available from October 23. ServeTheHome and VideoCardz report that 128GB pricing has also moved up because of memory costs.

Performance has improved a lot with software. In llama.cpp maintainer results, gpt-oss-120b generation rose from 38.6 tokens per second at launch to 58.7 tokens per second on a February 2026 build, with prefill around 2,444 tokens per second. That prefill lead over Strix Halo comes from Blackwell tensor cores.

Reasons to pick: CUDA in a small, efficient box; strong prefill; easy two-unit clustering for bigger models; good for fine-tuning experiments.

Reasons to skip: 273 GB/s limits decode on dense models; LMSYS measured just 2.7 tokens per second on Llama 3.1 70B at FP8; Linux only.

The catch: it is a developer appliance, not a general PC; if you want Windows apps, gaming or a desktop you can upgrade, look elsewhere.

4. Ryzen AI Max+ 395 mini PCs: best x86 unified-memory box

AMD’s Strix Halo chip combines 16 Zen 5 cores, a 40-CU Radeon 8060S GPU and up to 128GB of LPDDR5X-8000 on a 256-bit bus for 256 GB/s. Up to 96GB can be assigned to the GPU in BIOS, and AMD’s own guide shows a Linux kernel setting raising that to about 120GB on a 128GB Framework Desktop. Systems include the Framework Desktop, GMKtec EVO-X2, Minisforum MS-S1 Max, Beelink GTR9 Pro and AMD’s own Ryzen AI Halo developer box.

Community llama.cpp results put gpt-oss-120b at roughly 49 tokens per second on ROCm and 54.5 on Vulkan, close to the DGX Spark, though prefill is several times slower. AMD’s own figures claim small leads over the Spark on several MoE models.

Reasons to pick: runs both Windows and Linux; handles large MoE models; doubles as a capable compact PC with decent integrated graphics.

Reasons to skip: slow prefill; ROCm support for the gfx1151 target is still listed as preview; memory-shortage pricing has roughly doubled 128GB configurations since launch, per ComputingForGeeks’ price tracking.

The catch: on dense 70B models the 256 GB/s ceiling makes generation slow, so this is a MoE machine first.

5. Ryzen AI Max+ PRO 495 (Gorgon Halo): most memory in a mini PC

AMD’s Ryzen AI Max 400 refresh, announced in May 2026, keeps the Zen 5 and RDNA 3.5 design but raises memory to 192GB, with up to 160GB usable as VRAM according to AMD’s slides. The Register described it as essentially a factory-overclocked Strix Halo with much more memory. Systems on sale include the Minisforum MS-S1 MAX-P495, GMKtec Evo-X5, HP ZBook Ultra G3a and a 192GB Framework Desktop, all at the very top of the mini PC price range.

Reasons to pick: 192GB holds models up to roughly 300B parameters at low quants, per AMD; same software as Strix Halo.

Reasons to skip: bandwidth barely moves over Strix Halo, so speed per token is similar; pricing is close to a Mac Studio M5 Ultra.

The catch: more memory without much more bandwidth means bigger models run, but slowly; check that the models you want are MoE.

6. RTX Spark laptops: best portable AI PC

NVIDIA and Microsoft announced RTX Spark on May 31, 2026. It pairs a 20-core Grace CPU with a Blackwell RTX GPU of up to 6,144 CUDA cores and up to 128GB of LPDDR5X, linked by NVLink-C2C at up to 300 GB/s. NVIDIA rates it at up to one petaflop of FP4 AI performance and says it can run 120B-parameter models. Microsoft’s Surface Laptop Ultra ships October 16, and more than 30 laptop designs and about 10 desktop designs are planned from ASUS, Dell, HP, Lenovo, MSI and others.

Reasons to pick: big unified memory in a laptop; CUDA on Windows; NVIDIA also positions it for 1440p gaming.

Reasons to skip: Windows on Arm still relies on the Prism emulator for many x86 apps; the base configuration uses a smaller 5,120-core GPU with 24GB.

The catch: independent local-LLM measurements were not yet available at the time of writing, so treat NVIDIA’s claims as claims until reviews land.

7. Mid-range RTX desktop: best budget AI PC

A desktop with a 16GB card is the cheapest way to get a genuinely useful local AI setup. ComputingForGeeks measured Qwen2.5 14B at Q4_K_M at 73.2 tokens per second on an RTX 5070 Ti and 42.3 on an RTX 5060 Ti 16GB. Both run gpt-oss-20b, which OpenAI designed to fit in 16GB, and every 7B to 14B model you are likely to use day to day.

Reasons to pick: lowest entry tier; also a strong gaming PC; easy upgrade path to a bigger card later.

Reasons to skip: 16GB rules out 27B-plus dense models at good quality; the 5060 Ti has half the bandwidth of the 5070 Ti.

The catch: ComputingForGeeks found the 5070 Ti cheaper per token per second than the 5060 Ti despite the higher sticker, so the cheaper card is not automatically the better value.

How to choose

Start with the models, not the hardware.

  • Mostly 7B to 14B models for chat, coding help and summaries: a 16GB GPU desktop is enough. Spend extra on the 5070 Ti tier for speed.
  • 27B to 32B dense models: you need 24GB or more. An RTX 5090 is fastest; a used RTX 3090 is the value route; a Mac Studio M5 Max also handles these comfortably.
  • 70B dense models: plan for about 40GB plus context. A Mac Studio with 64GB or more, two 24GB cards, or a 128GB unified-memory machine. Expect modest speeds on everything except the M5 Ultra.
  • 100B-plus MoE models such as gpt-oss-120b: a 128GB DGX Spark, Strix Halo or Mac Studio. The Spark is best for prefill; the Mac is best for decode.
  • You also game: a Windows desktop with an RTX card is the only option that is excellent at both.
  • You build or fine-tune with CUDA tools: RTX desktop or DGX Spark.

Setup tips

  • Start with LM Studio or Ollama. Both handle model downloads and GPU detection for you. Move to llama.cpp, MLX or vLLM directly when you want more control.
  • Use 4-bit quants as your default. Q4_K_M in GGUF or 4-bit MLX keeps most quality at about a quarter of the memory of FP16.
  • Leave headroom for context. Long conversations and documents grow the KV cache; budget a few extra gigabytes.
  • On Strix Halo, set the GPU memory split. The default BIOS allocation is often small; raise it to 64GB or 96GB, or use the Linux kernel parameters in AMD’s guide.
  • Update software often. The DGX Spark gained about 50 percent in gpt-oss-120b generation from software in four months, per the llama.cpp results.

Common mistakes

  • Buying capacity you will not use. 512GB sounds exciting, but most daily work runs on 8B to 32B models.
  • Ignoring bandwidth. A 192GB box with 256 GB/s will run a dense 70B model slower than a Mac with 614 GB/s and half the memory.
  • Trusting launch-day benchmarks. ServeTheHome’s early DGX Spark result of 14.5 tokens per second on gpt-oss-120b was about a quarter of later llama.cpp numbers.
  • Forgetting the memory shortage. Prices on memory-heavy systems have climbed sharply through 2026; check live pricing before you commit.

How we compared

We compared these using manufacturer specifications, independent lab measurements from the sources named above, and owner reports. We did not bench-test these units ourselves. Where results came from different software builds, we noted the build and gave ranges rather than single numbers.

Sources

  • Apple Newsroom, new Mac Studio with M5 Max and M5 Ultra
  • MacStories, M5 Ultra Mac Studio review
  • Ars Technica, Mac Studio M5 Ultra coverage
  • NVIDIA Newsroom and product pages, DGX Spark and RTX Spark
  • VideoCardz, DGX Spark 64GB and RTX Spark laptop launch coverage
  • ServeTheHome, DGX Spark 64GB launch and Ryzen AI Halo developer system review
  • LMSYS Org, NVIDIA DGX Spark in-depth review
  • llama.cpp GitHub discussion #16578, DGX Spark performance
  • AMD, Ryzen AI Max+ 395 technical articles and Ryzen AI Halo product page
  • Tom’s Hardware, Ryzen AI Max 400 Gorgon Halo announcement
  • The Register, Gorgon Halo coverage
  • ComputingForGeeks, Ryzen AI Max+ 395 mini PC comparison and RTX 5060 Ti vs 5070 Ti results
  • OpenBenchmarking.org, llama.cpp results
  • AIMultiple, DGX Spark alternatives

Frequently Asked Questions

What is the best AI PC for running LLMs locally?

For speed on models up to about 32B, an RTX 5090 desktop. For the largest models on one box, a Mac Studio with M5 Max or M5 Ultra. For CUDA development in a small form factor, the DGX Spark.

Is a Copilot+ PC with an NPU good for local LLMs?

Not on its own. NPUs in typical laptops are built for small, efficient AI features. Running chat-sized LLMs well needs lots of fast memory feeding a GPU, which is why the machines on this list focus on VRAM or unified memory.

How much RAM do I need for a local AI PC?

For 7B to 14B models, a 16GB GPU plus 32GB of system RAM is fine. For 70B models, plan on about 48GB or more available to the accelerator. For 100B-plus MoE models, 128GB of unified memory is the practical starting point.

DGX Spark or Mac Studio?

Pick the Spark if you need CUDA, Linux and fast prompt processing. Pick the Mac Studio if you want faster token generation, more memory options and a machine that also works as an everyday computer.

Can a Strix Halo mini PC game as well as run AI?

Yes, its 40-CU Radeon 8060S is one of the strongest integrated GPUs available, but it is not a match for a discrete mid-range card in demanding games.

Are RTX Spark laptops worth waiting for?

They are the most interesting portable option on paper, with up to 128GB of unified memory. Wait for independent reviews of local-LLM speed and app compatibility on Windows on Arm before buying.

Why are AI PCs so expensive in 2026?

Memory makers have shifted capacity to high-bandwidth memory for data centers, which has pushed up the cost of LPDDR5X and DDR5. Memory-heavy machines like 128GB mini PCs and the DGX Spark have risen in price as a result.

G
GizmoPC Editorial Team
We compare specs, materials and verified owner reviews before a product earns a spot. Rankings are never paid.
Affiliate disclosure. As an Amazon Associate we earn from qualifying purchases at no extra cost to you. Prices accurate as of the date shown.