Model tests
AI model tests on real files
One source, one task, several AI models: GPT, Claude, Gemini and local Ollama models answer side by side in Creatos, with timings, speed and cost. Each page is one run on a stated date and setup, with the full answers so you can judge them yourself.
- Model comparisonQwen 3 vs Gemma 3 vs Llama 3.2 on a 16 GB Mac: NASA PDF testThree small local models read the same 32-page NASA PDF in Creatos on a 16 GB M4 Mac mini. All three named the right Artemis landers; only Gemma 3 followed the format.qwen3:4bgemma3:4bllama3.2:3bTested
- How-to testWhy Ollama is slow on a Mac: context length, measuredOn a 16 GB M4 Mac mini, a 128K context length pushed Qwen 3 4B and Llama 3.2 3B partly onto the CPU: up to 39% slower and up to 3.7x longer to the first token.Tested