The short answer
Run ollama ps. The SIZE column is the memory each loaded model holds right now: roughly the model's download size plus the context cache, so a 14B model at Q4_K_M (a 9 GB download) holds about 10 GB at an 8K context. The PROCESSOR column should read 100% GPU; a CPU/GPU split means it didn't fit in the memory the GPU may use. The UNTIL column says when Ollama unloads it, 5 minutes after the last request by default. To give the memory back now, run ollama stop with the model's name.
On this page
Check what's loaded with ollama ps
Ollama loads a model the first time you send it a request and keeps it in memory for a while afterwards, so the next request starts instantly. ollama ps lists what's loaded now:
ollama psNAME ID SIZE PROCESSOR CONTEXT UNTIL
qwen2.5-coder:14b 9ec8897f747e 10 GB 100% GPU 8192 4 minutes from now| Column | Meaning |
|---|---|
| SIZE | Memory the loaded model holds: weights, context cache and working buffers. |
| PROCESSOR | 100% GPU means it all runs on the GPU. A split such as 38%/62% CPU/GPU means part of it didn't fit in the GPU's share of memory and runs slower on the CPU. |
| CONTEXT | The context length it was loaded with (recent versions). A longer context needs more memory. |
| UNTIL | When Ollama will unload it if nothing else asks for it. |
SIZE
- Meaning
- Memory the loaded model holds: weights, context cache and working buffers.
PROCESSOR
- Meaning
100% GPUmeans it all runs on the GPU. A split such as38%/62% CPU/GPUmeans part of it didn't fit in the GPU's share of memory and runs slower on the CPU.
CONTEXT
- Meaning
- The context length it was loaded with (recent versions). A longer context needs more memory.
UNTIL
- Meaning
- When Ollama will unload it if nothing else asks for it.
The same figures come from Ollama's local API, which is what other tools read:
curl -s http://localhost:11434/api/psWhy it holds more than the download size
A model in memory is three things added together:
- The weights. About the download size. Quantisation decides it: Q4_K_M stores roughly 0.6 GB per billion parameters, Q8_0 about 1.1 GB, and full 16-bit weights 2 GB.
- The context cache. Memory for every token of context the model can see, reserved up front for the context length it was loaded with. For an 8B Llama model it's about 1 GB at 8K tokens and grows in step with the context: 16 GB at 128K.
- Working buffers. Usually a few hundred MB.
| Model size | Download | Loaded at 8K context, roughly |
|---|---|---|
| 7–8B | 4.5–5 GB | 6 GB |
| 14B | 9 GB | 10–11 GB |
| 32B | 20 GB | 22 GB |
| 70B | 43 GB | 46 GB |
7–8B
- Download
- 4.5–5 GB
- Loaded at 8K context, roughly
- 6 GB
14B
- Download
- 9 GB
- Loaded at 8K context, roughly
- 10–11 GB
32B
- Download
- 20 GB
- Loaded at 8K context, roughly
- 22 GB
70B
- Download
- 43 GB
- Loaded at 8K context, roughly
- 46 GB
The GPU's share of memory
An Apple silicon Mac has one pool of unified memory shared by the CPU and GPU. macOS lets the GPU use a large share of it, but not all: Metal reports a recommended working-set size for each Mac, typically somewhere between two thirds and three quarters of its memory. A model that fits under that limit runs fully on the GPU. One that doesn't is split, and Ollama runs the rest on the CPU, much more slowly.
So on a 16 GB Mac, a 14B model at Q4_K_M is near the limit before any other app has asked for memory. On a 36 GB Mac it fits with room to spare. The will this model fit guide shows how to work it out for your Mac.
How to give the memory back
ollama stop qwen2.5-coder:14bTo change how long models stay loaded, set the keep-alive. Per request, keep_alive accepts a duration such as "10m", 0 to unload straight after answering, or -1 to keep the model loaded. For every model, set OLLAMA_KEEP_ALIVE in the environment Ollama's server runs in.
curl http://localhost:11434/api/generate -d '{"model": "qwen2.5-coder:14b", "keep_alive": 0}'To use less memory while it's loaded, pick a smaller quantisation (Q4 instead of Q8) or a shorter context. The context length is set per request with num_ctx, or for the server with OLLAMA_CONTEXT_LENGTH in recent versions.
When a model makes the whole Mac slow
- Memory pressure turns yellow or red when a model loads on top of a browser, Docker and an IDE. macOS compresses memory and swaps, and everything slows down. Unload the model, or quit something large first.
- The PROCESSOR column shows a CPU/GPU split. The model is bigger than the GPU's share of memory. A smaller model or quantisation fixes it; a faster CPU doesn't.
- Two models are loaded at once. Ollama can keep several loaded, and each holds its own memory.
ollama pslists them all.
The faster way: AI models in iKnowMyMac
iKnowMyMac's AI models section, under Projects, asks Ollama, LM Studio, llama.cpp's server and MLX servers on 127.0.0.1 which models are loaded, and shows the memory each one holds against the GPU memory limit macOS sets for your Mac, with when Ollama will unload it. Nothing leaves your Mac, and prompts are never read.