Local AI buying guide

Choose hardware for the workload, not the sticker.

Compatibility, memory, thermals, model fit, and official purchase sources for local runtimes and agent gateways.

Last checked: 2026-08-26 · Manufacturer links only · no fabricated benchmarks

Keep/Cut WeeklyMethodology

Start with the architecture

Ollama and LM Studio run models; Open WebUI provides a UI; Hermes Agent and OpenClaw orchestrate tools and channels. A gateway can run on a small low-power host while inference runs on a separate machine or cloud provider. Do not buy a GPU just to host a lightweight gateway.

Local-engine watch (checked 2026-09-02): Perplexity’s official blog describes Lily as a specialized Apple silicon engine for Qwen3.6-35B-A3B used in Perplexity Computer hybrid compute. It is an open-source demo, not a directory SKU. Official demo notes Apple GPU family 10 or later (M5 and newer). See the Change Radar and official Lily post.

Cross-stack decision matrix

HardwareBest useMemory guidanceEvidence
Mac mini with M6
compact desktop
Best compact Apple choice for a quiet always-on agent gateway plus modest local inference; choose 32GB if local models matter more than basic gateway work.16GB standard; configure 24GB or 32GB for local models and concurrent services.View evidence →
Mac mini with M5 Pro
compact pro desktop
The most compelling new Mac mini tier for local models plus always-on agent services, if the 48GB/64GB configurations fit the budget.Up to 64GB unified memory according to Apple announcement; prefer 48GB or 64GB for larger local models and multiple services.View evidence →
Mac Studio with M5 Max or M5 Ultra
high-memory pro desktop
Choose only when large local models or heavy concurrent workloads justify the cost; otherwise a 32GB/64GB Mac mini or discrete-GPU PC is easier to justify.Apple announcement: up to 128GB M5 Max and 512GB M5 Ultra. Configure for the model class and concurrency actually planned.View evidence →
NVIDIA RTX 5090 desktop workstation
discrete-GPU desktop
Best choice for CUDA-heavy local AI and generation workflows when power, heat, and desktop size are acceptable.32GB dedicated VRAM; system RAM of 64GB is a sensible workstation target for local AI, containers, datasets, and browser/agent overhead.View evidence →
Raspberry Pi 5 (8GB or 16GB) gateway
low-power ARM gateway
Buy for an inexpensive always-on gateway, not for serious local model inference.4GB is enough for a simple gateway; 8GB is a comfortable target; 16GB is unnecessary for most gateway-only deployments.View evidence →

How to use this guide

  1. Identify whether you need a gateway, local inference, or both.
  2. Choose the exact model family and quantization before selecting RAM or VRAM.
  3. Leave headroom for the operating system, context/KV cache, containers, browser automation, and concurrent services.
  4. Check official runtime support for the OS, chip, driver, and version.
  5. Buy only after testing a representative task or reserving a return path.

Test Ollama locallySecure the stack