Model comparison · one run each
Qwen 3 vs Gemma 3 vs Llama 3.2 on a 16 GB Mac: NASA PDF test
Three small local models read the same 32-page NASA PDF in Creatos on a 16 GB M4 Mac mini. All three named the right Artemis landers; only Gemma 3 followed the format.
- Tested
- Updated
- Models
- 3
TL;DR
- All three models found the right answer in the 32-page NASA deck: SpaceX Starship for Artemis III and Blue Origin's Blue Moon for Artemis V.
- Only Gemma 3 4B followed the requested format (3 bullets, then the landers). Qwen 3 4B wrote a paragraph, and Llama 3.2 3B ran its 3 bullets together on one line.
- With all three running at once on a 16 GB M4 Mac mini, answer text appeared after 43 s (Qwen), 55 s (Gemma) and 63 s (Llama). All three were done 79 s after the first question was sent.
- Cost: $0. Every model ran offline through Ollama, with an 8K context window.


Setup
- App
- Creatos 0.2 (preview build)
- Machine
- Mac mini, Apple M4, 16 GB unified memory, macOS 26.3.1
- Runtime
- Ollama 0.34.4
- Context window
- 8,192 tokens
- Temperature
- 0.7
- Max output tokens
- 2,048
- System prompt
- None
- Models running
- All three at once
Task
Prompt, identical for every model
In 3 short bullets, what is this document about? Then: which lander is planned for Artemis III, and which for Artemis V?
Source: Artemis Human Landing System Update (NASA Technical Reports Server 20240009596)
Each model also received the text Creatos extracted from the PDF (8,657 characters from 32 pages). No system prompt and no other instructions.
The answers, side by side
Same source, same prompt, one run each
qwen3:4b
LocalOllama
- Cost$0
This document is a NASA presentation from Langley Research Center about the Artemis Human Landing System updates, focusing on mission concepts and technical requirements for lunar landings. It details the Artemis program's evolution, including the Human Landing System (HLS) and specific mission planning for Artemis III and Artemis V. The document particularly addresses the challenges of landing at the lunar south pole for Artemis III with precise landing requirements.
- Artemis III lander: SpaceX Starship
- Artemis V lander: Blue Origin Blue Moon (single stage lander)
Named both landers correctly but ignored the format: a paragraph instead of 3 bullets. Answer text appeared 43.0 s after sending.
gemma3:4b
LocalOllama
- Cost$0
Here’s a summary of the document in 3 short bullets:
- This document provides an update on the NASA’s Artemis human lunar exploration program, focusing on the Artemis III mission.
- It details the key technologies, missions, and operational requirements for establishing a sustainable cislunar ecosystem.
- The document outlines the challenges and considerations for landing at the lunar South Pole, a critical target for Artemis III.
Lander for Artemis III: The Human Landing System (HLS), currently being developed by SpaceX under the Starship program.
Lander for Artemis V: Blue Origin’s Blue Moon Single Stage Lander.
Named both landers correctly and followed the format: 3 bullets, then the landers. Answer text appeared 54.5 s after sending.
llama3.2:3b
LocalOllama
- Cost$0
Here are 3 short bullets summarizing the document:
• The document discusses the Artemis program, a NASA mission aimed at returning humans to the Moon and establishing a sustainable presence on the lunar surface. • The document provides details on the challenges of landing at the Lunar South Pole for Artemis III, including the need for precise landing accuracy, suitable surface conditions, and reliable communication. • The document also outlines the concept of operations for Artemis III and Artemis V, including the roles of the Gateway, Orion, and landers in supporting the mission objectives.
According to the document, the planned lander for Artemis III is the Human Landing System (HLS) Starship, developed by SpaceX.
The planned lander for Artemis V is the Blue Moon Single Stage Landers, developed by Blue Origin.
Named both landers correctly. Wrote 3 bullets but ran them together on one line. Answer text appeared 62.9 s after sending.
Measurements
Results (all three models ran at the same time on one Mac)
| Model | Artemis III lander | Artemis V lander | Followed the format | Answer visible after (s) |
|---|---|---|---|---|
| qwen3:4b | Correct: Starship | Correct: Blue Moon | No, wrote a paragraph | 43.0 |
| gemma3:4b | Correct: Starship | Correct: Blue Moon | Yes | 54.5 |
| llama3.2:3b | Correct: Starship | Correct: Blue Moon | Partly, bullets on one line | 62.9 |
How we tested
- We added the PDF to a Creatos canvas as one source. Creatos extracted its text: 8,657 characters from 32 pages.
- We connected that source to three AI nodes, one per model, all served by Ollama on the same Mac.
- We typed the same question into each node and sent the three about 2 seconds apart, so the models ran at the same time and shared the machine. Running alone, each would be faster: in our context-length test, the same models at 8K took 5.6–7.5 s to the first token.
- No system prompt and no hidden instructions. Temperature 0.7 (the Creatos default), up to 2,048 output tokens and an 8,192-token context window.
- Times come from the screen recording's log. "Answer visible after" is when at least 150 characters of the answer were on screen, counted from when that model's question was sent, so it is a little later than the first token.
We checked the landers against the document itself: its mission slides list SpaceX Starship for Artemis III and IV, and Blue Origin's Blue Moon for Artemis V. Creatos did not save the raw output of this run, so the answers above are copied from the screen recording: the words are unedited, and list bullets and bold text are shown as the app displayed them.
Source
Artemis Human Landing System Update (NASA Technical Reports Server 20240009596)
- Credit:
- Alicia Dwyer Cianciolo, NASA Langley Research Center, 2024
- License:
- NASA work, not subject to copyright in the US. NASA does not endorse Creatos.
Run this comparison on your own files
Creatos is a desktop app for macOS and Windows. Add your own PDFs, notes and web pages as sources, ask GPT, Claude, Gemini and local Ollama models the same question, and read the answers side by side.
$69 once. No subscription.
Cloud model usage is billed by your provider. Local models run on your own machine.
FAQ
Can a 16 GB Mac answer questions about a long PDF with local models?
Yes, for a document like this one. All three 3–4B models answered the factual question correctly from a 32-page deck, which came to about 2,500 tokens of text. Keep Ollama's context length at 8K–32K: at 128K, Qwen 3 4B and Llama 3.2 3B need more memory than a 16 GB Mac has and slow down. See our context-length test.
Which of the three should I use?
On this task all three were accurate, and only Gemma 3 4B followed the requested format. If you need answers in a fixed shape (bullets, a table, a template), test instruction-following on your own documents: one run on one question is a single data point.
Did this cost anything?
No. All three models ran offline through Ollama, so there were no API fees, and the PDF never left the Mac.
Related tests
All model testsModel names are used for identification only. Creatos is not affiliated with OpenAI, Anthropic, Google, Meta, Alibaba or Ollama.