Local AI Capacity
Planner
Plan Qwen3.8, Qwen3.5, and Gemma 4 deployments on Apple Silicon (M1–M6, including M5 Ultra) and NVIDIA consumer GPUs. Estimate weight footprint, KV cache, safe concurrency, decode ceiling, and whole-system memory per dollar before you commit.
🤖 Model
🖥️ Hardware
⚖️ Quantization
Qwen3.8 uses exact source-artifact sizes. GGUF matches llama.cpp/Ollama; MLX requires a converted artifact whose final size must be checked. Official model card. A decode estimate appears only when the hardware, runtime, and represented artifact path match.
🔀 Concurrency & Context
Capacity and decode values are conservative planning estimates. Unified memory is shared with macOS and applications; measured results vary by artifact, runtime, context, batching, and thermal state. KV uses official text configurations: token-growing full-attention cache plus Gemma sliding cache capped at 1,024 tokens (conservatively storing separate K and V). Qwen recurrent state, vision projectors, MTP, and runtime allocations remain outside the model-plus-KV total.
Breakdown
Agent Slots
💡 Tips
2026 Quick Reference — Qwen3.8-27B Q4_K_M
| Hardware | Memory | Bandwidth | US price | GB / $1K | 2 × 16k fit | Est. decode |
|---|---|---|---|---|---|---|
| M6 mini · 16 GB | 16 GB | 153 GB/s | $899 | 17.80 | ❌ OOM | ~8 t/s |
| M6 mini · 24 GB | 24 GB | 170 GB/s | $1,099 | 21.84 | ✅ Tight | ~9 t/s |
| M6 mini · 32 GB | 32 GB | 170 GB/s | $1,299 | 24.63 | ✅ Good | ~9 t/s |
| M5 Ultra · 96 GB | 96 GB | 1,200 GB/s | $5,499 | 17.46 | ✅ Ample | ~63 t/s |
| M5 Ultra · 256 GB | 256 GB | 1,200 GB/s | $9,499 | 26.95 | ✅ Ample | ~63 t/s |
| M5 Ultra · 512 GB | 512 GB | 1,200 GB/s | TBD | N/A | ✅ Ample | ~63 t/s |
| M4 Max · 128 GB | 128 GB | 546 GB/s | Not tracked | N/A | ✅ Ample | ~29 t/s |
Sources: Qwen architecture, observed GGUF artifact sizes, and Apple's Aug. 25 launch release.
Prices are complete-system US prices before tax, checked Aug. 30, 2026; the 512 GB M5 Ultra price is unpublished and availability begins in late October.
Decode figures are single-device memory-bandwidth planning estimates at 87% assumed llama.cpp efficiency, not measured benchmarks. Multi-GPU decode is unavailable because PCIe/NVLink topology prevents a defensible sum of per-card bandwidth. Figures exclude the optional vision projector and MTP file. Apple's launch AI results measure prompt processing and are not used here.
Need help designing a production local-AI stack for your team?
Talk to Dataxad