The short answer
Add the model's weights to its context cache and compare the total with the memory your Mac's GPU may use. Weights: about 0.6 GB per billion parameters at Q4_K_M, 1.1 GB at Q8_0, 2 GB at 16-bit, or simply the download size. Context: about 1 GB per 8K tokens for an 8B model, more for bigger ones. The GPU may use roughly two thirds to three quarters of the Mac's memory. Then subtract what your open apps hold. If the model only fits with everything else closed, pick a smaller quantisation or a shorter context.
On this page
The arithmetic
needed = weights + context cache + about 0.5 GB of working buffers
fits if: needed < GPU limit, and needed < memory your open apps leave free| Quantisation | GB per billion parameters | 14B model |
|---|---|---|
| Q4_K_M | about 0.6 | about 9 GB |
| Q5_K_M | about 0.7 | about 10.5 GB |
| Q8_0 | about 1.1 | about 15 GB |
| F16 / BF16 | 2 | 28 GB |
Q4_K_M
- GB per billion parameters
- about 0.6
- 14B model
- about 9 GB
Q5_K_M
- GB per billion parameters
- about 0.7
- 14B model
- about 10.5 GB
Q8_0
- GB per billion parameters
- about 1.1
- 14B model
- about 15 GB
F16 / BF16
- GB per billion parameters
- 2
- 14B model
- 28 GB
The download size is the most reliable figure for the weights, because it's the weights. A GGUF file of 9.0 GB needs about 9 GB of memory before any context.
The context is the part people forget
Runtimes reserve memory for the whole context length when they load a model. For an 8B Llama model that's about 128 KB per token: 1 GB at 8K, 4 GB at 32K, 16 GB at 128K. Bigger models with more layers need more per token. A model that fits easily at 8K can stop fitting at 32K.
Find your Mac's GPU limit
macOS lets the GPU use a large share of unified memory, but not all of it. Metal reports the figure as the device's recommended working-set size, and runtimes such as llama.cpp, Ollama and MLX plan around it. It's typically between two thirds and three quarters of the Mac's memory.
There's no built-in command that prints it, but the runtimes report it:
grep -i recommendedMaxWorkingSetSize ~/.ollama/logs/server.log | tail -1python3 -c "import mlx.core as mx; print(mx.metal.device_info()['max_recommended_working_set_size'] / 1e9, 'GB')"Common models at Q4_K_M and 8K context
| Model | Needs, roughly | 16 GB Mac | 24 GB Mac | 36 GB Mac | 64 GB Mac |
|---|---|---|---|---|---|
| 7–8B | 6 GB | Fits | Fits | Fits | Fits |
| 14B | 10–11 GB | Tight: close other apps | Fits | Fits | Fits |
| 32B | 22 GB | No | No | Fits | Fits |
| 70B | 46 GB | No | No | No | Tight: may not fit |
7–8B
- Needs, roughly
- 6 GB
- 16 GB Mac
- Fits
- 24 GB Mac
- Fits
- 36 GB Mac
- Fits
- 64 GB Mac
- Fits
14B
- Needs, roughly
- 10–11 GB
- 16 GB Mac
- Tight: close other apps
- 24 GB Mac
- Fits
- 36 GB Mac
- Fits
- 64 GB Mac
- Fits
32B
- Needs, roughly
- 22 GB
- 16 GB Mac
- No
- 24 GB Mac
- No
- 36 GB Mac
- Fits
- 64 GB Mac
- Fits
70B
- Needs, roughly
- 46 GB
- 16 GB Mac
- No
- 24 GB Mac
- No
- 36 GB Mac
- No
- 64 GB Mac
- Tight: may not fit
“Fits” assumes your other apps leave the memory free. A browser with many tabs, Docker and an iOS simulator can hold 10 GB or more between them, and that comes out of the same pool. Check memory pressure after the model loads: if it turns yellow or red, it didn't really fit.
When it doesn't fit
- Use a smaller quantisation: Q4_K_M instead of Q8_0 roughly halves the weights, for a small loss in quality.
- Shorten the context: 8K instead of 32K frees several GB.
- Unload other models, and quit apps holding a lot of memory.
- Pick the next size down: a 14B model at Q8 is often better than a 32B model that only runs half on the CPU.
The faster way: “Will it fit?” in iKnowMyMac
In iKnowMyMac's AI models section, “Will it fit?” takes an installed model or a size you type, such as 14B Q4. It reads the model's own metadata where the file is on your Mac, adds the context you choose, and compares the total with free memory and your Mac's GPU memory limit. It works offline.