# Creatos > Creatos is a desktop app (macOS and Windows) for working with AI on a visual canvas. You put sources on the canvas (YouTube videos, PDFs, web pages, notes), connect them to AI nodes, and compare cloud models and local models side by side on the same material. Website: https://creatos.io. Short index: https://creatos.io/llms.txt. ## Facts - Price: $69 once, a lifetime license with lifetime updates. No subscription. - There is no free trial: the app needs a license to open. Purchases are generally non-refundable except where the law requires. - Models: bring your own API keys for cloud providers such as OpenAI, Anthropic and Google Gemini, and pay them directly for usage. Local models run through Ollama with no key. - Data stays on your device: projects are stored locally. - Systems: macOS 12 or later on Apple Silicon (M1 or newer), and Windows 10 or later, 64-bit (beta). - Support: hello@creatos.io. ## FAQ ### Why not use ChatGPT or Claude in the browser? Browser chat suits single prompts. Creatos is for work that spans several sources and steps: research, captured pages and AI answers stay on one canvas, and you can run the same sources through several models to compare them. ### Can I use local models? Yes. Creatos talks to Ollama on your machine and needs no API key for it. ### Can I move my license to a new machine? Yes. Deactivate on the old machine, then activate on the new one with the same purchase email and license key. ## Pages - [Home](https://creatos.io): what Creatos does - [Pricing](https://creatos.io/pricing): one-time lifetime license - [Download](https://creatos.io/download): macOS and Windows installers - [Research canvas example](https://creatos.io/use-cases/research-canvas): a worked example ## Guide - [AI Nodes](https://creatos.io/docs/ai-nodes): Generate content with AI using the chat interface - [Browser Nodes](https://creatos.io/docs/browser-nodes): Browse pages inside Creatos and extract readable content - [Content Nodes](https://creatos.io/docs/content-nodes): Import, write, and edit content in a single node type - [Flow Editor](https://creatos.io/docs/flow-editor): Build, import, organize, and export canvases in Creatos - [Welcome to Creatos](https://creatos.io/docs): Official guide for the Creatos desktop app - [Saved Outputs](https://creatos.io/docs/output-history): Review, search, and manage generated content - [Quick Start](https://creatos.io/docs/quick-start): Create your first flow in a few minutes - [Content Library](https://creatos.io/docs/source-library): Save and reuse content across flows - [Templates](https://creatos.io/docs/templates): Reuse proven flow structures ## Model tests Each test runs the same task through several models or settings on our own machine and reports what we measured. ### Ollama KV cache quantization on Mac: q8_0 vs q4_0, measured URL: https://creatos.io/tests/ollama-kv-cache-quantization-mac Tested: 2026-10-06. Updated: 2026-10-06. Machine: Mac mini, Apple M4, 16 GB unified memory, macOS 26.3.1. On a 16 GB M4 Mac mini, a q8_0 KV cache halved Ollama's cache with no loss in our answer check and kept 128K mostly on the GPU. q4_0 saved more but got a fact wrong. - q8_0 halved the KV cache with no loss we could see: qwen3:4b at 32K went from 7.6 to 5.3 GB at the same 32 tokens/s and still got both facts right, and llama3.2:3b gave the same answer, word for word, with every cache type. - At 128K, q8_0 took qwen3:4b from 23 GB with 51% on the CPU (17.4 tokens/s) to 14 GB with 19% on the CPU (26.5 tokens/s). llama3.2:3b went from 18 GB, 39% on the CPU, to 10 GB fully on the GPU at 42.9 tokens/s. - q4_0 quartered the cache: qwen3:4b at 128K fit on the GPU in 8.6 GB at full speed. But it named the wrong lander for Artemis V, which f16 and q8_0 got right. One run, so treat q4_0 as a risk, not a verdict. - Flash attention, which a quantized cache needs, was already on by default here. Forcing it off doubled qwen3:4b at 32K to 16 GB, with 29% on the CPU. - On a 16 GB Mac, set OLLAMA_KV_CACHE_TYPE=q8_0 when you need a long context. Try q4_0 only if q8_0 still spills, and check its answers on your own documents. ### Why Ollama is slow on a Mac: context length, measured URL: https://creatos.io/tests/why-ollama-is-slow-on-mac-context-length Tested: 2026-10-05. Updated: 2026-10-06. Machine: Mac mini, Apple M4, 16 GB unified memory, macOS 26.3.1. On a 16 GB M4 Mac mini, a 128K context length pushed Qwen 3 4B and Llama 3.2 3B partly onto the CPU: up to 39% slower and up to 3.7x longer to the first token. - At 128K, qwen3:4b needed 23 GB, more than the 16 GB Mac has, so 51% of it ran on the CPU. Speed fell from 32 to 19.4 tokens/s and the first token took 27.9 s instead of 7.5–9.1 s. - llama3.2:3b followed the same pattern: 18 GB at 128K with 39% on the CPU, 31.4 tokens/s instead of about 41, and 14.8 s to the first token instead of 5.6 s. - gemma3:4b stayed at 3.5–4.0 GB and about 35 tokens/s at every setting, even 128K. - With a 2,500-token prompt, 4K, 8K and 32K ran at the same speed. On a 16 GB Mac, an 8K–32K context length is enough unless a document really needs more. ### Qwen 3 vs Gemma 3 vs Llama 3.2 on a 16 GB Mac: NASA PDF test URL: https://creatos.io/tests/qwen3-vs-gemma3-vs-llama32-16gb-mac-nasa-pdf Tested: 2026-10-05. Updated: 2026-10-05. Machine: Mac mini, Apple M4, 16 GB unified memory, macOS 26.3.1. Three small local models read the same 32-page NASA PDF in Creatos on a 16 GB M4 Mac mini. All three named the right Artemis landers; only Gemma 3 followed the format. - All three models found the right answer in the 32-page NASA deck: SpaceX Starship for Artemis III and Blue Origin's Blue Moon for Artemis V. - Only Gemma 3 4B followed the requested format (3 bullets, then the landers). Qwen 3 4B wrote a paragraph, and Llama 3.2 3B ran its 3 bullets together on one line. - With all three running at once on a 16 GB M4 Mac mini, answer text appeared after 43 s (Qwen), 55 s (Gemma) and 63 s (Llama). All three were done 79 s after the first question was sent. - Cost: $0. Every model ran offline through Ollama, with an 8K context window. ## Free tools ### Ollama Memory Calculator for Mac (VRAM, Measured) URL: https://creatos.io/tools/ollama-memory-calculator How much memory will an Ollama model use on your Mac? Pick a model, context length and RAM to see the GB used and the GPU/CPU split, checked against real runs. ## How-to guides ### Excalidraw MCP in Claude: setup, a real run, and the alternatives URL: https://creatos.io/guides/excalidraw-mcp Add the official Excalidraw MCP to Claude in a minute, see what one real run drew and cost, and compare it with the community server, tldraw and Miro.