Mac mini with M6
Best compact Apple choice for a quiet always-on agent gateway plus modest local inference; choose 32GB if local models matter more than basic gateway work.
Compatibility and purchase evidence →Local AI buying guide
Compatibility, memory, thermals, model fit, and official purchase sources for local runtimes and agent gateways.
Last checked: 2026-08-26 · Manufacturer links only · no fabricated benchmarks
Ollama and LM Studio run models; Open WebUI provides a UI; Hermes Agent and OpenClaw orchestrate tools and channels. A gateway can run on a small low-power host while inference runs on a separate machine or cloud provider. Do not buy a GPU just to host a lightweight gateway.
Local-engine watch (checked 2026-09-02): Perplexity’s official blog describes Lily as a specialized Apple silicon engine for Qwen3.6-35B-A3B used in Perplexity Computer hybrid compute. It is an open-source demo, not a directory SKU. Official demo notes Apple GPU family 10 or later (M5 and newer). See the Change Radar and official Lily post.
Best compact Apple choice for a quiet always-on agent gateway plus modest local inference; choose 32GB if local models matter more than basic gateway work.
Compatibility and purchase evidence →The most compelling new Mac mini tier for local models plus always-on agent services, if the 48GB/64GB configurations fit the budget.
Compatibility and purchase evidence →Choose only when large local models or heavy concurrent workloads justify the cost; otherwise a 32GB/64GB Mac mini or discrete-GPU PC is easier to justify.
Compatibility and purchase evidence →Best choice for CUDA-heavy local AI and generation workflows when power, heat, and desktop size are acceptable.
Compatibility and purchase evidence →Buy for an inexpensive always-on gateway, not for serious local model inference.
Compatibility and purchase evidence →| Hardware | Best use | Memory guidance | Evidence |
|---|---|---|---|
| Mac mini with M6 compact desktop | Best compact Apple choice for a quiet always-on agent gateway plus modest local inference; choose 32GB if local models matter more than basic gateway work. | 16GB standard; configure 24GB or 32GB for local models and concurrent services. | View evidence → |
| Mac mini with M5 Pro compact pro desktop | The most compelling new Mac mini tier for local models plus always-on agent services, if the 48GB/64GB configurations fit the budget. | Up to 64GB unified memory according to Apple announcement; prefer 48GB or 64GB for larger local models and multiple services. | View evidence → |
| Mac Studio with M5 Max or M5 Ultra high-memory pro desktop | Choose only when large local models or heavy concurrent workloads justify the cost; otherwise a 32GB/64GB Mac mini or discrete-GPU PC is easier to justify. | Apple announcement: up to 128GB M5 Max and 512GB M5 Ultra. Configure for the model class and concurrency actually planned. | View evidence → |
| NVIDIA RTX 5090 desktop workstation discrete-GPU desktop | Best choice for CUDA-heavy local AI and generation workflows when power, heat, and desktop size are acceptable. | 32GB dedicated VRAM; system RAM of 64GB is a sensible workstation target for local AI, containers, datasets, and browser/agent overhead. | View evidence → |
| Raspberry Pi 5 (8GB or 16GB) gateway low-power ARM gateway | Buy for an inexpensive always-on gateway, not for serious local model inference. | 4GB is enough for a simple gateway; 8GB is a comfortable target; 16GB is unnecessary for most gateway-only deployments. | View evidence → |