Creatos Logo
Compare
Buy License

How-to test · one run each

Why Ollama is slow on a Mac: context length, measured

On a 16 GB M4 Mac mini, a 128K context length pushed Qwen 3 4B and Llama 3.2 3B partly onto the CPU: up to 39% slower and up to 3.7x longer to the first token.

Tested
Updated

TL;DR

  • At 128K, qwen3:4b needed 23 GB, more than the 16 GB Mac has, so 51% of it ran on the CPU. Speed fell from 32 to 19.4 tokens/s and the first token took 27.9 s instead of 7.5–9.1 s.
  • llama3.2:3b followed the same pattern: 18 GB at 128K with 39% on the CPU, 31.4 tokens/s instead of about 41, and 14.8 s to the first token instead of 5.6 s.
  • gemma3:4b stayed at 3.5–4.0 GB and about 35 tokens/s at every setting, even 128K.
  • With a 2,500-token prompt, 4K, 8K and 32K ran at the same speed. On a 16 GB Mac, an 8K–32K context length is enough unless a document really needs more.
Grid of qwen3:4b, gemma3:4b and llama3.2:3b at 4K, 8K, 32K and 128K context on a 16 GB Mac mini M4, with memory, tokens per second and time to the first token. At 128K, qwen3:4b uses 23 GB with 51% on the CPU and drops to 19.4 tok/s; llama3.2:3b uses 18 GB with 39% on the CPU and drops to 31.4 tok/s; gemma3:4b stays near 4 GB and 35 tok/s.
All 12 runs at a glance: only the 128K runs of Qwen 3 and Llama 3.2 spill onto the CPU.

Setup

App
None: requests went straight to the Ollama API (/api/generate)
Machine
Mac mini, Apple M4, 16 GB unified memory, macOS 26.3.1
Runtime
Ollama 0.34.4
Context lengths tested
4K, 8K, 32K and 128K (num_ctx)
Temperature
0
Seed
42
Max output tokens
400
Cold start before each run
Yes
Qwen 3 thinking
Off (think: false)

Task

Prompt, identical for every model

In 3 short bullets, what is this document about? Then: which lander is planned for Artemis III, and which for Artemis V?

Source: Artemis Human Landing System Update (NASA Technical Reports Server 20240009596)

Each request sent the text of the NASA PDF followed by this question: 2,568 prompt tokens for qwen3:4b, 2,466 for gemma3:4b and 2,481 for llama3.2:3b (each model splits text into tokens differently).

Measurements

Memory used (GB), by context length

Model4K8K32K128K
qwen3:4b3.23.97.623
gemma3:4b3.73.94.03.5
llama3.2:3b2.53.06.018

Generation speed (tokens/s), by context length

Model4K8K32K128K
qwen3:4b31.832.032.019.4
gemma3:4b35.035.335.335.4
llama3.2:3b40.441.341.431.4

Seconds to the first token (cold start, includes loading)

Model4K8K32K128K
qwen3:4b9.17.58.027.9
gemma3:4b9.76.86.87.0
llama3.2:3b6.85.65.614.8

All 12 runs

ModelContextMemory (GB)Ran onLoad (s)Prompt (s)First token (s)Speed (tokens/s)Output tokensTotal (s)
qwen3:4b4K3.2100% GPU2.36.89.131.840021.7
qwen3:4b8K3.9100% GPU0.86.77.532.040020.0
qwen3:4b32K7.6100% GPU1.36.78.032.040020.6
qwen3:4b128K2351% CPU / 49% GPU10.517.327.919.440048.5
gemma3:4b4K3.7100% GPU4.15.79.735.012413.3
gemma3:4b8K3.9100% GPU1.35.46.835.312410.3
gemma3:4b32K4.0100% GPU1.35.46.835.312410.3
gemma3:4b128K3.5100% GPU1.65.47.035.412410.5
llama3.2:3b4K2.5100% GPU2.14.86.840.415010.6
llama3.2:3b8K3.0100% GPU0.84.85.641.31509.2
llama3.2:3b32K6.0100% GPU0.94.85.641.41509.3
llama3.2:3b128K1839% CPU / 61% GPU5.98.914.831.415019.6

How we tested

We sent the same prompt straight to Ollama's API (/api/generate), once per model and context length: 12 runs in all.

  • Prompt: the text of NASA's Artemis Human Landing System Update followed by one question, 2,466–2,568 prompt tokens depending on the model.
  • Context length: set per request with options.num_ctx = 4,096, 8,192, 32,768 and 131,072.
  • Cold start: each model was unloaded first (keep_alive: 0), so the time to the first token includes loading the model.
  • Sampling: temperature 0, seed 42, at most 400 output tokens. Qwen 3 ran with think: false but still wrote its reasoning out as text, so it always used all 400 tokens; Gemma 3 and Llama 3.2 stopped on their own after 124 and 150 tokens.
  • Measured: memory and the CPU/GPU split from ollama ps during each run, and speed as eval_count / eval_duration from Ollama's response.

The test changes only how much context is allocated, not how long the prompt is. Longer prompts would make large contexts slower still, and the 128K results depend on what else is using memory at the time.

Source

Artemis Human Landing System Update (NASA Technical Reports Server 20240009596)

Credit:
Alicia Dwyer Cianciolo, NASA Langley Research Center, 2024
License:
NASA work, not subject to copyright in the US. NASA does not endorse Creatos.

Compare local and cloud models on your own files

Creatos is a desktop app for macOS and Windows. Add your own PDFs, notes and web pages as sources, ask GPT, Claude, Gemini and local Ollama models the same question, and read the answers side by side.

$69 once. No subscription.

Cloud model usage is billed by your provider. Local models run on your own machine.

FAQ

How do I change Ollama's context length on a Mac?

In the Ollama app, open Settings and move the Context length slider. If you start Ollama from the terminal, set OLLAMA_CONTEXT_LENGTH, for example OLLAMA_CONTEXT_LENGTH=8192 ollama serve. Apps that call the API can also pass num_ctx with each request, which is what this test did.

How can I tell if a model is running partly on the CPU?

Run ollama ps while the model is answering. The PROCESSOR column says 100% GPU when everything fits, or a split such as 48%/52% CPU/GPU when it does not. Recent versions also show the context the model was loaded with. Check it: some versions of the desktop app have loaded models with a larger context than the slider showed (ollama#16896).

What context length should I use on a 16 GB Mac?

8K to 32K is a good default. Here 4K, 8K and 32K ran at the same speed, while 128K pushed two of the three models partly onto the CPU. Raise it only when a document needs it: the 32-page NASA deck in this test was only about 2,500 tokens of text. To check a model and context length before you load it, use our Ollama memory calculator.

Why didn't Gemma 3 slow down at 128K?

We saw it but did not investigate why. Google describes Gemma 3 as using mostly sliding-window (local) attention layers, which keep the memory needed for long contexts small.

All model tests

Model names are used for identification only. Creatos is not affiliated with OpenAI, Anthropic, Google, Meta, Alibaba or Ollama.