Skip to content

Guide

Will this model fit on my Mac?

A local model fits when its weights plus its context fit in the memory the GPU may use, with room left for your other apps. Here's the arithmetic, a table of common sizes, and how to find your Mac's own limit.

By the iKnowMyMac team · Updated · 6 min read

The short answer

Add the model's weights to its context cache and compare the total with the memory your Mac's GPU may use. Weights: about 0.6 GB per billion parameters at Q4_K_M, 1.1 GB at Q8_0, 2 GB at 16-bit, or simply the download size. Context: about 1 GB per 8K tokens for an 8B model, more for bigger ones. The GPU may use roughly two thirds to three quarters of the Mac's memory. Then subtract what your open apps hold. If the model only fits with everything else closed, pick a smaller quantisation or a shorter context.

On this page

The arithmetic

What a model needs
needed  = weights + context cache + about 0.5 GB of working buffers
fits if: needed < GPU limit, and needed < memory your open apps leave free
  • Q4_K_M

    GB per billion parameters
    about 0.6
    14B model
    about 9 GB
  • Q5_K_M

    GB per billion parameters
    about 0.7
    14B model
    about 10.5 GB
  • Q8_0

    GB per billion parameters
    about 1.1
    14B model
    about 15 GB
  • F16 / BF16

    GB per billion parameters
    2
    14B model
    28 GB
Weights per billion parameters, by quantisation (rule of thumb)

The download size is the most reliable figure for the weights, because it's the weights. A GGUF file of 9.0 GB needs about 9 GB of memory before any context.

The context is the part people forget

Runtimes reserve memory for the whole context length when they load a model. For an 8B Llama model that's about 128 KB per token: 1 GB at 8K, 4 GB at 32K, 16 GB at 128K. Bigger models with more layers need more per token. A model that fits easily at 8K can stop fitting at 32K.

Find your Mac's GPU limit

macOS lets the GPU use a large share of unified memory, but not all of it. Metal reports the figure as the device's recommended working-set size, and runtimes such as llama.cpp, Ollama and MLX plan around it. It's typically between two thirds and three quarters of the Mac's memory.

There's no built-in command that prints it, but the runtimes report it:

From Ollama's log, once it has loaded a model
grep -i recommendedMaxWorkingSetSize ~/.ollama/logs/server.log | tail -1
From MLX in Python
python3 -c "import mlx.core as mx; print(mx.metal.device_info()['max_recommended_working_set_size'] / 1e9, 'GB')"

Common models at Q4_K_M and 8K context

  • 7–8B

    Needs, roughly
    6 GB
    16 GB Mac
    Fits
    24 GB Mac
    Fits
    36 GB Mac
    Fits
    64 GB Mac
    Fits
  • 14B

    Needs, roughly
    10–11 GB
    16 GB Mac
    Tight: close other apps
    24 GB Mac
    Fits
    36 GB Mac
    Fits
    64 GB Mac
    Fits
  • 32B

    Needs, roughly
    22 GB
    16 GB Mac
    No
    24 GB Mac
    No
    36 GB Mac
    Fits
    64 GB Mac
    Fits
  • 70B

    Needs, roughly
    46 GB
    16 GB Mac
    No
    24 GB Mac
    No
    36 GB Mac
    No
    64 GB Mac
    Tight: may not fit
Roughly what each needs, against a GPU share of two thirds to three quarters of memory

“Fits” assumes your other apps leave the memory free. A browser with many tabs, Docker and an iOS simulator can hold 10 GB or more between them, and that comes out of the same pool. Check memory pressure after the model loads: if it turns yellow or red, it didn't really fit.

When it doesn't fit

  1. Use a smaller quantisation: Q4_K_M instead of Q8_0 roughly halves the weights, for a small loss in quality.
  2. Shorten the context: 8K instead of 32K frees several GB.
  3. Unload other models, and quit apps holding a lot of memory.
  4. Pick the next size down: a 14B model at Q8 is often better than a 32B model that only runs half on the CPU.

The faster way: “Will it fit?” in iKnowMyMac

In iKnowMyMac's AI models section, “Will it fit?” takes an installed model or a size you type, such as 14B Q4. It reads the model's own metadata where the file is on your Mac, adds the context you choose, and compares the total with free memory and your Mac's GPU memory limit. It works offline.

FAQ

Questions

Something else? support@iknowmymac.com

Can a 16 GB Mac run a 14B model?

Only just, at Q4_K_M with a short context and few other apps open. The weights alone are about 9 GB and the GPU's share of 16 GB is around 10–12 GB. A 7–8B model is the comfortable choice on 16 GB.

What size model can a 36 GB Mac run?

Up to about 32B parameters at Q4_K_M with an 8K context, which needs around 22 GB against a GPU share of about 27 GB. A 70B model doesn't fit.

Does quantisation change whether a model fits?

Yes, more than anything else. The same 14B model needs about 9 GB at Q4_K_M, 15 GB at Q8_0 and 28 GB at 16-bit.

Why does a model that fits still make my Mac slow?

Because the model and your apps share one pool of memory. If the model fits under the GPU limit but your apps need the rest, macOS compresses memory and swaps to disk. Watch memory pressure, not just the model's size.

Know why your Mac is slow.

macOS 15 Sequoia or later · Apple silicon. A 4.2 MB native app, notarized by Apple. No account, no analytics, no telemetry.

One-time purchase · every feature · 14-day refund