Creatos Logo
Compare
Buy License

Free tool · checked against real runs

Ollama memory calculator for Mac

Pick a model, a context length and your Mac's memory to see how much memory Ollama will take and whether the model stays 100% on the GPU. The formula matches what ollama ps showed on a 16 GB M4 Mac mini within 0.1 GB.

Context length

Will my PDF fit?

pages

About 2,560 tokens. With your question and the answer, a 4K context holds it.

qwen3:4b at 32K on a 16 GB Mac

≈ 7.7 GB

100% GPU

GPU limit 11.5 GB
Weights
2.5 GB
KV cache
4.8 GB
Overhead
0.3 GB

Fits in the 11.5 GB macOS lets the GPU use, so the whole model runs on the GPU.

Measured with ollama ps on our 16 GB M4 Mac mini: 7.6 GB. See the test

How the calculator works

Ollama's memory is three parts:

memory   ≈ weights + KV cache + overhead
KV cache = 2 × layers × KV heads × head size × bytes × context
overhead ≈ 0.1 GB + 7.5 KB × context

Weights are the download size of the model's default tag. The KV cache holds the context, and Ollama reserves it for the whole window when the model loads, so it grows with the setting, not with your prompt. It takes 2 bytes a value by default (f16). Overhead is the compute buffer; we fitted it to our measurements.

The GPU limit is the share of memory macOS lets the GPU use by default: about two thirds on Macs with up to 36 GB, three quarters above. Whatever doesn't fit under it runs on the CPU, which is much slower. On our 16 GB Mac that limit is about 11 GB, which matches the GPU part of the splits ollama ps reported.

Gemma 3 is the exception. Most of its layers use a short sliding window, so its memory hardly grows with the context. For it the calculator shows what we measured.

Calculator vs. measured

Every run from our context-length test on an M4 Mac mini with 16 GB and Ollama 0.34.4, read from ollama ps. At 128K, the test measured 49% on the GPU for qwen3:4b and 61% for llama3.2:3b.

ModelContextMeasured (GB)Calculator (GB)GPU (calc.)
qwen3:4b4K3.23.2100%
qwen3:4b8K3.93.9100%
qwen3:4b32K7.67.7100%
qwen3:4b128K2322.949%
llama3.2:3b4K2.52.6100%
llama3.2:3b8K33.1100%
llama3.2:3b32K66.1100%
llama3.2:3b128K1818.163%

If the model spills to the CPU

  • Lower the context length. 8K to 32K is plenty for most chats and documents. In the Ollama app it is under Settings › Context length; from the terminal, set OLLAMA_CONTEXT_LENGTH; over the API, pass num_ctx.
  • Check it while it answers. Run ollama ps. You want 100% GPU in the PROCESSOR column.
  • Quantize the KV cache if you really need a long context: OLLAMA_FLASH_ATTENTION=1 and OLLAMA_KV_CACHE_TYPE=q8_0.
  • Pick a model built for long context, such as gemma3:4b, which stayed at 3.5 GB and 100% GPU at 128K.

FAQ

How much memory do I need to run Ollama on a Mac?

The model's download size plus its KV cache, which grows with the context length. To stay fast it must fit under the GPU limit: about two thirds of memory on Macs with up to 36 GB, about 11 GB on a 16 GB Mac. There, 4B to 8B models fit at 8K to 32K, and 14B models only at short contexts such as 8K.

Why does a longer context use so much more memory?

Ollama reserves the KV cache for the whole context window when it loads the model, whether your prompt uses it or not. For qwen3:4b that is about 0.15 MB per token: 1.2 GB at 8K, 19 GB at 128K.

Is VRAM the same as RAM on a Mac?

Apple Silicon Macs have unified memory: the CPU and GPU share it. macOS lets the GPU use only part of it by default, so a model can spill onto the CPU before your memory is full. ollama ps shows the split in its PROCESSOR column.

Should I quantize the KV cache?

If you need a long context, it is the cheapest saving. Start Ollama with OLLAMA_FLASH_ATTENTION=1 and OLLAMA_KV_CACHE_TYPE=q8_0: Ollama's docs describe q8_0 as about half the memory of f16 with a very small loss in precision, and q4_0 as about a quarter with a more noticeable loss at long contexts. We have not measured either yet.

The tests behind it

Why Ollama is slow on a Mac: one setting has all 12 runs and the method. More model tests. We ran them in Creatos, the desktop app we make, which puts local and cloud models side by side on your own files.