How-to test · one run each
Ollama KV cache quantization on Mac: q8_0 vs q4_0, measured
On a 16 GB M4 Mac mini, a q8_0 KV cache halved Ollama's cache with no loss in our answer check and kept 128K mostly on the GPU. q4_0 saved more but got a fact wrong.
- Tested
- Updated
- Answers
- 4
TL;DR
- q8_0 halved the KV cache with no loss we could see: qwen3:4b at 32K went from 7.6 to 5.3 GB at the same 32 tokens/s and still got both facts right, and llama3.2:3b gave the same answer, word for word, with every cache type.
- At 128K, q8_0 took qwen3:4b from 23 GB with 51% on the CPU (17.4 tokens/s) to 14 GB with 19% on the CPU (26.5 tokens/s). llama3.2:3b went from 18 GB, 39% on the CPU, to 10 GB fully on the GPU at 42.9 tokens/s.
- q4_0 quartered the cache: qwen3:4b at 128K fit on the GPU in 8.6 GB at full speed. But it named the wrong lander for Artemis V, which f16 and q8_0 got right. One run, so treat q4_0 as a risk, not a verdict.
- Flash attention, which a quantized cache needs, was already on by default here. Forcing it off doubled qwen3:4b at 32K to 16 GB, with 29% on the CPU.
- On a 16 GB Mac, set OLLAMA_KV_CACHE_TYPE=q8_0 when you need a long context. Try q4_0 only if q8_0 still spills, and check its answers on your own documents.

Setup
- App
- None: requests went straight to the Ollama API (/api/generate)
- Machine
- Mac mini, Apple M4, 16 GB unified memory, macOS 26.3.1
- Runtime
- Ollama 0.34.4
- KV cache types tested
- f16 (the default), q8_0 and q4_0 (OLLAMA_KV_CACHE_TYPE)
- Flash attention
- On (OLLAMA_FLASH_ATTENTION=1), plus one run with it off
- Context lengths tested
- 32K and 128K (num_ctx)
- Temperature
- 0
- Seed
- 42
- Max output tokens
- 400 for the speed runs, 2,048 for the Qwen 3 answer check
- Cold start before each run
- Yes
- Qwen 3 thinking
- Off (think: false)
Task
Prompt, identical for every model
In 3 short bullets, what is this document about? Then: which lander is planned for Artemis III, and which for Artemis V?
Source: Artemis Human Landing System Update (NASA Technical Reports Server 20240009596)
Each request sent the text of the NASA PDF followed by this question: 2,568 prompt tokens for qwen3:4b and 2,481 for llama3.2:3b. The document plans SpaceX's Starship HLS for Artemis III and Blue Origin's Blue Moon for Artemis V.
The answers, side by side
Same source, same prompt, one run each
qwen3:4b · f16 cache
LocalOllama
Here's the concise response:
In 3 short bullets, what this document is about:
- It details the specific landing requirements and constraints for Artemis III at the lunar south pole (e.g., latitude, slope, lighting, communication).
- It analyzes challenges of landing near the south pole due to extreme lighting, terrain, and orbital mechanics.
- It outlines future operations for Artemis IV–V, including the Human Landing System (HLS) and Gateway station.
Lander plans:
- Artemis III: SpaceX’s Starship (as explicitly stated: "NASA has awarded SpaceX a contract to develop its HLS Starship for use on Artemis III").
- Artemis V: Blue Origin’s Blue Moon (referenced in the "Artemis V Concept of Operations" diagram as the "Single Stage Landers" for lunar missions).
Right on both landers. 32K context, up to 2,048 output tokens. Qwen 3 wrote 758 tokens, most of them its reasoning as plain text even with thinking off; only the final answer after it is shown.
qwen3:4b · q8_0 cache
LocalOllama
Here's the concise response:
In 3 short bullets, what this document is about:
- It details NASA's Artemis III lunar landing requirements for the South Pole region, emphasizing strict constraints like latitude limits, slope thresholds, Earth visibility, and lighting conditions.
- It analyzes technical challenges of landing near the lunar south pole, including orbital mechanics, surface terrain, and communication needs for a safe crewed mission.
- It outlines the operational concept for Artemis V, involving Blue Origin’s Blue Moon lander and a cislunar refueling system to support sustained lunar exploration.
Lander assignments:
- Artemis III: SpaceX’s Starship (as confirmed by NASA’s award for HLS development).
- Artemis V: Blue Origin’s Blue Moon (explicitly listed under "Artemis V" in the document’s concept of operations).
Right on both landers. Same settings; 782 tokens in all, final answer shown.
qwen3:4b · q4_0 cache
LocalOllama
Here's the concise response:
In 3 short bullets:
- This document details NASA's Artemis Human Landing System (HLS) updates, focusing on landing requirements for Artemis III (first lunar landing) and Artemis V (next crewed mission), with emphasis on lunar south pole targeting challenges.
- It outlines technical constraints for Artemis III, including precise landing near the lunar south pole (within 6° latitude), steep slope limits (<10° for landing, <20° for EVA), direct Earth communication, and lighting conditions.
- It shows Artemis V operations will reuse the same HLS system with refueling and transporter support for sustained lunar exploration, avoiding new lander development.
Lander specifics:
- Artemis III: SpaceX Starship (as explicitly stated: "NASA has awarded SpaceX a contract to develop its HLS Starship for use on Artemis III")
- Artemis V: Same SpaceX Starship-derived HLS system (with upgrades for refueling and transporter operations, per the "Artemis V Concept of Operations" diagram showing reused lander/transporter)
Wrong on Artemis V: the document plans Blue Origin's Blue Moon, not a Starship-derived lander. Same settings; 1,083 tokens in all, final answer shown.
llama3.2:3b · all six settings
LocalOllama
Here are 3 short bullets summarizing the document:
- The document discusses the Artemis Human Landing System (HLS) and its development for NASA's Artemis program, which aims to return humans to the Moon by 2025.
- The document highlights the challenges of landing on the Lunar South Pole, including the need for direct communication with Earth, suitable lighting conditions, and precise landing accuracy.
- The document also outlines the concept of operations for Artemis III and Artemis V, including the use of the HLS Starship and the Gateway space station.
According to the document, the Human Landing System (HLS) Starship is planned for Artemis III, while the Blue Origin Blue Moon lander is planned for Artemis V.
The same answer, word for word, with f16, q8_0 and q4_0 at both 32K and 128K. Right on both landers.
Measurements
Memory used (GB), by KV cache type
| Model | Context | f16 | q8_0 | q4_0 |
|---|---|---|---|---|
| qwen3:4b | 32K | 7.6 | 5.3 | 4.1 |
| qwen3:4b | 128K | 23 | 14 | 8.6 |
| llama3.2:3b | 32K | 6.0 | 4.3 | 3.3 |
| llama3.2:3b | 128K | 18 | 10 | 7.0 |
Share of the model on the GPU
| Model | Context | f16 | q8_0 | q4_0 |
|---|---|---|---|---|
| qwen3:4b | 32K | 100% | 100% | 100% |
| qwen3:4b | 128K | 49% | 81% | 100% |
| llama3.2:3b | 32K | 100% | 100% | 100% |
| llama3.2:3b | 128K | 61% | 100% | 100% |
Generation speed (tokens/s)
| Model | Context | f16 | q8_0 | q4_0 |
|---|---|---|---|---|
| qwen3:4b | 32K | 32.0 | 32.6 | 32.7 |
| qwen3:4b | 128K | 17.4 | 26.5 | 32.9 |
| llama3.2:3b | 32K | 40.8 | 42.7 | 42.3 |
| llama3.2:3b | 128K | 29.1 | 42.9 | 42.7 |
Seconds to the first token (cold start, includes loading)
| Model | Context | f16 | q8_0 | q4_0 |
|---|---|---|---|---|
| qwen3:4b | 32K | 10.1 | 9.6 | 9.3 |
| qwen3:4b | 128K | 34.3 | 14.9 | 8.2 |
| llama3.2:3b | 32K | 7.6 | 7.2 | 6.9 |
| llama3.2:3b | 128K | 14.9 | 6.4 | 5.9 |
KV cache size (GB, from Ollama's server log)
| Model | Context | f16 | q8_0 | q4_0 |
|---|---|---|---|---|
| qwen3:4b | 32K | 4.8 | 2.6 | 1.4 |
| qwen3:4b | 128K | 19.3 | 10.3 | 5.5 |
| llama3.2:3b | 32K | 3.8 | 2.0 | 1.1 |
| llama3.2:3b | 128K | 15.0 | 7.9 | 4.2 |
Flash attention on vs. off (qwen3:4b, 32K, f16 cache)
| Flash attention | Memory (GB) | Ran on | Speed (tokens/s) | First token (s) | Total (s) |
|---|---|---|---|---|---|
| On (the default here) | 7.6 | 100% GPU | 32.0 | 10.1 | 22.6 |
| Off | 16 | 29% CPU / 71% GPU | 26.1 | 10.8 | 26.1 |
All 13 runs
| Model | Context | KV cache | Flash attention | Memory (GB) | Ran on | Load (s) | Prompt (s) | First token (s) | Speed (tokens/s) | Output tokens | Total (s) |
|---|---|---|---|---|---|---|---|---|---|---|---|
| qwen3:4b | 32K | f16 | On | 7.6 | 100% GPU | 3.3 | 6.7 | 10.1 | 32.0 | 400 | 22.6 |
| qwen3:4b | 32K | q8_0 | On | 5.3 | 100% GPU | 2.8 | 6.8 | 9.6 | 32.6 | 400 | 21.9 |
| qwen3:4b | 32K | q4_0 | On | 4.1 | 100% GPU | 2.6 | 6.7 | 9.3 | 32.7 | 400 | 21.5 |
| qwen3:4b | 32K | f16 | Off | 16 | 29% CPU / 71% GPU | 3.1 | 7.7 | 10.8 | 26.1 | 400 | 26.1 |
| qwen3:4b | 128K | f16 | On | 23 | 51% CPU / 49% GPU | 11.1 | 23.1 | 34.3 | 17.4 | 400 | 57.4 |
| qwen3:4b | 128K | q8_0 | On | 14 | 19% CPU / 81% GPU | 3.6 | 11.2 | 14.9 | 26.5 | 400 | 30.0 |
| qwen3:4b | 128K | q4_0 | On | 8.6 | 100% GPU | 1.3 | 6.9 | 8.2 | 32.9 | 400 | 20.4 |
| llama3.2:3b | 32K | f16 | On | 6.0 | 100% GPU | 2.8 | 4.8 | 7.6 | 40.8 | 150 | 11.3 |
| llama3.2:3b | 32K | q8_0 | On | 4.3 | 100% GPU | 2.2 | 5.0 | 7.2 | 42.7 | 150 | 10.8 |
| llama3.2:3b | 32K | q4_0 | On | 3.3 | 100% GPU | 2.1 | 4.8 | 6.9 | 42.3 | 150 | 10.5 |
| llama3.2:3b | 128K | f16 | On | 18 | 39% CPU / 61% GPU | 5.9 | 9.0 | 14.9 | 29.1 | 150 | 20.0 |
| llama3.2:3b | 128K | q8_0 | On | 10 | 100% GPU | 1.6 | 4.8 | 6.4 | 42.9 | 150 | 9.9 |
| llama3.2:3b | 128K | q4_0 | On | 7.0 | 100% GPU | 1.1 | 4.8 | 5.9 | 42.7 | 150 | 9.4 |
How we tested
We sent the same prompt straight to Ollama's API (/api/generate), restarting the server for each cache setting: 13 runs in all, plus an answer check.
- Prompt: the text of NASA's Artemis Human Landing System Update followed by one question, the same as in our context-length test: 2,568 prompt tokens for qwen3:4b and 2,481 for llama3.2:3b.
- Cache settings: the server ran with
OLLAMA_FLASH_ATTENTION=1andOLLAMA_KV_CACHE_TYPEset tof16,q8_0orq4_0. One extra run usedOLLAMA_FLASH_ATTENTION=0with f16. Ollama's server log confirmed each setting and reported the KV cache size. - Context length: set per request with
options.num_ctx= 32,768 and 131,072. - Cold start: each model was unloaded first (
keep_alive: 0), so the time to the first token includes loading the model. - Sampling: temperature 0, seed 42, at most 400 output tokens. Qwen 3 ran with
think: falsebut still wrote its reasoning out as text, so its 400-token runs ended before the final answer. - Answer check: llama3.2:3b's answers were compared word for word across all six settings. qwen3:4b ran once more at 32K with up to 2,048 output tokens per cache type, and we checked the two landers against the document.
- Measured: memory and the CPU/GPU split from
ollama psduring each run, and speed aseval_count / eval_durationfrom Ollama's response.
The qwen3:4b run at 128K with f16 was slower than the same setting in the context-length test the day before (17.4 vs. 19.4 tokens/s, 34.3 vs. 27.9 s to the first token): runs that spill onto the CPU vary more. Runs that stayed on the GPU repeated within about 1 token/s.
Source
Artemis Human Landing System Update (NASA Technical Reports Server 20240009596)
- Credit:
- Alicia Dwyer Cianciolo, NASA Langley Research Center, 2024
- License:
- NASA work, not subject to copyright in the US. NASA does not endorse Creatos.
Compare local and cloud models on your own files
Creatos is a desktop app for macOS and Windows. Add your own PDFs, notes and web pages as sources, ask GPT, Claude, Gemini and local Ollama models the same question, and read the answers side by side.
$69 once. No subscription.
Cloud model usage is billed by your provider. Local models run on your own machine.
FAQ
How do I set OLLAMA_KV_CACHE_TYPE on a Mac?
If you use the Ollama app, run launchctl setenv OLLAMA_KV_CACHE_TYPE q8_0 in Terminal, then quit and reopen Ollama. This lasts until you restart the Mac. If you start Ollama from the terminal, use OLLAMA_FLASH_ATTENTION=1 OLLAMA_KV_CACHE_TYPE=q8_0 ollama serve. Then load a model and run ollama ps: the size should be smaller. The setting applies to every model the server loads.
Does q8_0 make answers worse?
Not in this test. llama3.2:3b gave the same answer, word for word, and qwen3:4b got both facts right, as with f16. Ollama's docs describe q8_0 as about half the memory of f16 with a very small loss of precision, and q4_0 as about a quarter with a more noticeable loss at long contexts. Here q4_0 was where qwen3:4b got a fact wrong.
Do I need flash attention?
Yes. A quantized KV cache only works with flash attention on. In Ollama 0.34.4 the default is auto, and on this M4 Mac mini it was on: the default run used the same 7.6 GB as forcing it on. Forcing it off doubled qwen3:4b at 32K to 16 GB, because the attention buffers grew from about 0.9 GB to 8.6 GB.
Should I quantize the cache or shorten the context?
If your documents fit in 8K to 32K, a shorter context is the simpler fix and costs nothing. Quantize when you really need a long context: with q8_0, llama3.2:3b ran 128K fully on the GPU of a 16 GB Mac. Our Ollama memory calculator shows both options for your model and Mac.
Does this apply to Gemma 3?
We did not test it here. Gemma 3 uses mostly sliding-window attention, so its cache stays small anyway: it used 3.5 to 4.0 GB at every context length in our context-length test.
Related tests
All model tests- How-to testWhy Ollama is slow on a Mac: context length, measuredOn a 16 GB M4 Mac mini, a 128K context length pushed Qwen 3 4B and Llama 3.2 3B partly onto the CPU: up to 39% slower and up to 3.7x longer to the first token.
- Model comparisonQwen 3 vs Gemma 3 vs Llama 3.2 on a 16 GB Mac: NASA PDF testThree small local models read the same 32-page NASA PDF in Creatos on a 16 GB M4 Mac mini. All three named the right Artemis landers; only Gemma 3 followed the format.
Model names are used for identification only. Creatos is not affiliated with OpenAI, Anthropic, Google, Meta, Alibaba or Ollama.