Skip to content

Guide

How much memory does Ollama use on a Mac?

An Ollama model holds about its download size in memory, plus room for the context, for as long as it stays loaded. Here's how to see the real figure, why it's bigger than you expect, and how to give the memory back.

By the iKnowMyMac team · Updated · 7 min read

The short answer

Run ollama ps. The SIZE column is the memory each loaded model holds right now: roughly the model's download size plus the context cache, so a 14B model at Q4_K_M (a 9 GB download) holds about 10 GB at an 8K context. The PROCESSOR column should read 100% GPU; a CPU/GPU split means it didn't fit in the memory the GPU may use. The UNTIL column says when Ollama unloads it, 5 minutes after the last request by default. To give the memory back now, run ollama stop with the model's name.

On this page

Check what's loaded with ollama ps

Ollama loads a model the first time you send it a request and keeps it in memory for a while afterwards, so the next request starts instantly. ollama ps lists what's loaded now:

Models Ollama has loaded
ollama ps
Example output
NAME                 ID              SIZE     PROCESSOR    CONTEXT    UNTIL
qwen2.5-coder:14b    9ec8897f747e    10 GB    100% GPU     8192       4 minutes from now
  • SIZE

    Meaning
    Memory the loaded model holds: weights, context cache and working buffers.
  • PROCESSOR

    Meaning
    100% GPU means it all runs on the GPU. A split such as 38%/62% CPU/GPU means part of it didn't fit in the GPU's share of memory and runs slower on the CPU.
  • CONTEXT

    Meaning
    The context length it was loaded with (recent versions). A longer context needs more memory.
  • UNTIL

    Meaning
    When Ollama will unload it if nothing else asks for it.
What the columns mean

The same figures come from Ollama's local API, which is what other tools read:

Loaded models as JSON (size and size_vram are in bytes)
curl -s http://localhost:11434/api/ps

Why it holds more than the download size

A model in memory is three things added together:

  1. The weights. About the download size. Quantisation decides it: Q4_K_M stores roughly 0.6 GB per billion parameters, Q8_0 about 1.1 GB, and full 16-bit weights 2 GB.
  2. The context cache. Memory for every token of context the model can see, reserved up front for the context length it was loaded with. For an 8B Llama model it's about 1 GB at 8K tokens and grows in step with the context: 16 GB at 128K.
  3. Working buffers. Usually a few hundred MB.
  • 7–8B

    Download
    4.5–5 GB
    Loaded at 8K context, roughly
    6 GB
  • 14B

    Download
    9 GB
    Loaded at 8K context, roughly
    10–11 GB
  • 32B

    Download
    20 GB
    Loaded at 8K context, roughly
    22 GB
  • 70B

    Download
    43 GB
    Loaded at 8K context, roughly
    46 GB
Typical sizes of Ollama models at Q4_K_M (download size; add the context on top)

The GPU's share of memory

An Apple silicon Mac has one pool of unified memory shared by the CPU and GPU. macOS lets the GPU use a large share of it, but not all: Metal reports a recommended working-set size for each Mac, typically somewhere between two thirds and three quarters of its memory. A model that fits under that limit runs fully on the GPU. One that doesn't is split, and Ollama runs the rest on the CPU, much more slowly.

So on a 16 GB Mac, a 14B model at Q4_K_M is near the limit before any other app has asked for memory. On a 36 GB Mac it fits with room to spare. The will this model fit guide shows how to work it out for your Mac.

How to give the memory back

Unload one model now
ollama stop qwen2.5-coder:14b

To change how long models stay loaded, set the keep-alive. Per request, keep_alive accepts a duration such as "10m", 0 to unload straight after answering, or -1 to keep the model loaded. For every model, set OLLAMA_KEEP_ALIVE in the environment Ollama's server runs in.

Unload through the API (what keep-alive 0 does)
curl http://localhost:11434/api/generate -d '{"model": "qwen2.5-coder:14b", "keep_alive": 0}'

To use less memory while it's loaded, pick a smaller quantisation (Q4 instead of Q8) or a shorter context. The context length is set per request with num_ctx, or for the server with OLLAMA_CONTEXT_LENGTH in recent versions.

When a model makes the whole Mac slow

  • Memory pressure turns yellow or red when a model loads on top of a browser, Docker and an IDE. macOS compresses memory and swaps, and everything slows down. Unload the model, or quit something large first.
  • The PROCESSOR column shows a CPU/GPU split. The model is bigger than the GPU's share of memory. A smaller model or quantisation fixes it; a faster CPU doesn't.
  • Two models are loaded at once. Ollama can keep several loaded, and each holds its own memory. ollama ps lists them all.

The faster way: AI models in iKnowMyMac

iKnowMyMac's AI models section, under Projects, asks Ollama, LM Studio, llama.cpp's server and MLX servers on 127.0.0.1 which models are loaded, and shows the memory each one holds against the GPU memory limit macOS sets for your Mac, with when Ollama will unload it. Nothing leaves your Mac, and prompts are never read.

FAQ

Questions

Something else? support@iknowmymac.com

How much RAM do I need to run Ollama on a Mac?

Enough for the model plus everything else you have open. As a rule of thumb, a model needs about its download size plus 1–2 GB at an 8K context, and it should stay under the GPU's share of memory, roughly two thirds to three quarters of the Mac's total. A 16 GB Mac runs 7–8B models comfortably; 14B models want 24 GB or more; 32B models want 36 GB or more.

Does Ollama keep using memory after I close the chat?

For a while. A model stays loaded for 5 minutes after the last request by default, then Ollama unloads it. ollama ps shows when, and ollama stop with the model's name unloads it straight away.

Why does ollama ps show more than the model's download size?

The figure includes the context cache, which is reserved for the whole context length when the model loads, plus working buffers. A longer context means a bigger number.

Is ollama ps the same as the memory in Activity Monitor?

No. Ollama runs models in a separate runner process, and GPU memory is counted differently from ordinary app memory, so Activity Monitor's Memory column usually doesn't match. Ollama's own figure is the one to go by for the model.

Know why your Mac is slow.

macOS 15 Sequoia or later · Apple silicon. A 4.2 MB native app, notarized by Apple. No account, no analytics, no telemetry.

One-time purchase · every feature · 14-day refund