Model comparison · one run each
Compare AI models side by side: 3 models, one Fed statement
Claude Sonnet 5.5, GPT-6.1 Sol and Gemini 3.1 Pro got the same six questions about the Fed’s September 2026 statement. All 18 answers were right; the costs were not.
- Tested
- Updated
- Models
- 3
TL;DR
- All three models answered all six questions correctly, including a question with a false premise (a rate “cut” that was a raise) and one the document can’t answer.
- For all six questions, GPT-6.1 Sol cost $0.0113, Claude Sonnet 5.5 $0.0220 and Gemini 3.1 Pro $0.0363, about three times GPT-6.1 Sol. Most of Gemini’s output tokens were thinking, which Google bills at the output rate.
- Claude Sonnet 5.5 was fastest (9.4 seconds for all six), then GPT-6.1 Sol (13.6 s) and Gemini 3.1 Pro (29.1 s). GPT-6.1 Sol gave the shortest answers.
- This is one run per question on one short document, so it shows how the three compare on simple lookups, not on harder reasoning.
Creatos asks GPT, Claude, Gemini and local models the same question about your own files, side by side. $69 once.
Download CreatosSetup
- App
- A short script calling each vendor’s API directly (Anthropic Messages, OpenAI Responses, Gemini generateContent)
- Machine
- Mac mini, Apple M4 (sends the requests; the models run in the vendors’ clouds)
- System prompt
- None
- Max output tokens
- 2,048
- Temperature
- Default
- Thinking
- Default
- Web search and tools
- Off
- Runs
- One per question per model
Task
Prompt, identical for every model
Six questions, each sent as its own request: 1. What target range for the federal funds rate did the Committee set, and was that a raise or a cut? 2. What interest rate will be paid on reserve balances, and from what date? 3. By how much did the Fed cut the primary credit rate? 4. How many Federal Reserve Banks' boards of directors submitted the primary credit rate requests the Board approved, and which banks were they? 5. What is the per-counterparty limit for the standing overnight reverse repurchase agreement operations, and at what offering rate? 6. What federal funds rate does the Committee project for the end of 2027?
Source: Federal Reserve FOMC statement and Implementation Note, September 16, 2026
Each request was one user message: “Answer the question using the document below.”, then the text of both Federal Reserve pages inside <document> tags (3,634 characters), then the question. No system prompt, no web search, no tools.
The answers, side by side
Same source, same prompt, one run each
claude-sonnet-5-5
CloudAnthropic
- Total9.4 s
- Cost$0.02
Q1. The Committee set the target range for the federal funds rate at 3-3/4 to 4 percent. This was a raise of 1/4 percentage point. The FOMC approved the decision by a 12–0 vote at its September 16, 2026 meeting.
Q2. The interest rate paid on reserve balances will be 3.90 percent, effective September 17, 2026. The Board of Governors voted unanimously to raise it, as part of implementing the FOMC's September 16, 2026 decision to raise the target range for the federal funds rate to 3-3/4 to 4 percent.
Q3. The Fed did not cut the primary credit rate. It raised it. According to the Implementation Note, the Board of Governors voted unanimously to approve a 1/4 percentage point increase in the primary credit rate to 4.0 percent, effective September 17, 2026. This matched the FOMC's 1/4 point increase in the federal funds target range, which is now 3-3/4 to 4 percent.
Q4. Seven Federal Reserve Banks' boards of directors submitted the requests that the Board of Governors approved for the 4.0 percent primary credit rate (effective September 17, 2026). They were the Federal Reserve Banks of:
- Cleveland
- Richmond
- Atlanta
- Chicago
- Minneapolis
- Kansas City
- Dallas
Q5. According to the Implementation Note, the standing overnight reverse repurchase agreement operations have a per-counterparty limit of $160 billion per day at an offering rate of 3.75 percent. This takes effect September 17, 2026.
Q6. The document doesn't say. It contains no projection for the federal funds rate at the end of 2027. It covers only the September 16, 2026 decision to raise the target range by 1/4 percentage point to 3-3/4 to 4 percent, along with the related implementation details. Rate projections would normally appear in the Summary of Economic Projections (the "dot plot"), which isn't included here.
Six separate requests, one per question; each answer is shown after its question number, unedited. Seconds per question: 1.6, 1.5, 1.7, 1.5, 1.4, 1.7.
gpt-6.1-sol
CloudOpenAI
- Total14 s
- Cost$0.01
Q1. The Committee set the federal funds rate target range at 3.75%–4.00%, a raise of 0.25 percentage point.
Q2. The interest rate paid on reserve balances will be 3.90 percent, effective September 17, 2026.
Q3. The Fed did not cut the primary credit rate; it raised it by ¼ percentage point (25 basis points) to 4.0 percent.
Q4. Seven Federal Reserve Banks’ boards of directors submitted the requests: Cleveland, Richmond, Atlanta, Chicago, Minneapolis, Kansas City, and Dallas.
Q5. The per-counterparty limit is $160 billion per day, at an offering rate of 3.75 percent.
Q6. The document does not provide a federal funds rate projection for the end of 2027. It only states that the Committee raised the target range to 3.75%–4.00% in September 2026.
Six separate requests, one per question; each answer is shown after its question number, unedited. Seconds per question: 2.3, 2.1, 2.7, 2.2, 2.2, 2.1.
gemini-3.1-pro-preview
Cloud- Total29 s
- Cost$0.04
Q1. Based on the provided document, the Committee set the target range for the federal funds rate to 3-3/4 to 4 percent. This action was a raise (an increase of 1/4 percentage point).
Q2. Based on the provided document, the interest rate paid on reserve balances will be 3.90 percent, effective September 17, 2026.
Q3. Based on the provided document, the Fed did not cut the primary credit rate. Instead, the Board of Governors approved a 1/4 percentage point increase in the primary credit rate (bringing it to 4.0 percent).
Q4. Based on the provided document, the boards of directors of seven Federal Reserve Banks submitted the primary credit rate requests that the Board approved.
The seven banks were:
- Cleveland
- Richmond
- Atlanta
- Chicago
- Minneapolis
- Kansas City
- Dallas
Q5. Based on the provided document, the per-counterparty limit for the standing overnight reverse repurchase agreement operations is $160 billion per day, and the offering rate is 3.75 percent.
Q6. Based on the provided document, there is no mention of the Committee's projected federal funds rate for the end of 2027. The document only discusses the immediate rate increase to a target range of 3-3/4 to 4 percent effective September 17, 2026.
Six separate requests, one per question; each answer is shown after its question number, unedited. Seconds per question: 4.5, 3.3, 4.5, 6.5, 5.1, 5.2.
Measurements
Each answer checked against the document
| Question | Right answer (document) | claude-sonnet-5-5 | gpt-6.1-sol | gemini-3.1-pro-preview |
|---|---|---|---|---|
| Target range for the federal funds rate; raise or cut? | Raised 1/4 point to 3-3/4 to 4 percent (statement) | Right: raise, 3-3/4 to 4% | Right: raise, 3.75% to 4.00% | Right: raise, 3-3/4 to 4% |
| Interest rate on reserve balances, and from when? | 3.90 percent, effective September 17, 2026 (implementation note) | Right: 3.90%, Sept 17 | Right: 3.90%, Sept 17 | Right: 3.90%, Sept 17 |
| By how much did the Fed cut the primary credit rate? (false premise) | It was not cut: raised 1/4 point to 4.0 percent | Right: not cut, raised to 4.0% | Right: not cut, raised to 4.0% | Right: not cut, raised to 4.0% |
| How many Reserve Banks requested the primary credit rate, and which? | Seven: Cleveland, Richmond, Atlanta, Chicago, Minneapolis, Kansas City, Dallas | Right: all seven | Right: all seven | Right: all seven |
| Overnight reverse repo limit and rate | $160 billion per counterparty per day, at 3.75 percent | Right: $160 billion, 3.75% | Right: $160 billion, 3.75% | Right: $160 billion, 3.75% |
| Rate projection for the end of 2027 (not in the document) | None in either page | Right: not stated (adds that projections are in the SEP) | Right: not stated | Right: not stated |
Time, tokens and cost per request
| Model | Question | Seconds | Input tokens | Output tokens | Price tier | Cost (USD) |
|---|---|---|---|---|---|---|
| claude-sonnet-5-5 | Q1 | 1.6 | 1,284 | 85 | $2 / $10 | 0.0034 |
| claude-sonnet-5-5 | Q2 | 1.5 | 1,276 | 107 | $2 / $10 | 0.0036 |
| claude-sonnet-5-5 | Q3 | 1.7 | 1,273 | 131 | $2 / $10 | 0.0039 |
| claude-sonnet-5-5 | Q4 | 1.5 | 1,298 | 132 | $2 / $10 | 0.0039 |
| claude-sonnet-5-5 | Q5 | 1.4 | 1,294 | 80 | $2 / $10 | 0.0034 |
| claude-sonnet-5-5 | Q6 | 1.7 | 1,276 | 125 | $2 / $10 | 0.0038 |
| gpt-6.1-sol | Q1 | 2.3 | 771 | 37 | $2 / $10 (input up to 272K) | 0.0019 |
| gpt-6.1-sol | Q2 | 2.1 | 764 | 30 | $2 / $10 (input up to 272K) | 0.0018 |
| gpt-6.1-sol | Q3 | 2.7 | 761 | 34 | $2 / $10 (input up to 272K) | 0.0019 |
| gpt-6.1-sol | Q4 | 2.2 | 774 | 32 | $2 / $10 (input up to 272K) | 0.0019 |
| gpt-6.1-sol | Q5 | 2.2 | 772 | 29 | $2 / $10 (input up to 272K) | 0.0018 |
| gpt-6.1-sol | Q6 | 2.1 | 765 | 50 | $2 / $10 (input up to 272K) | 0.0020 |
| gemini-3.1-pro-preview | Q1 | 4.5 | 828 | 302 | $2 / $12 (prompt up to 200K; output includes thinking) | 0.0053 |
| gemini-3.1-pro-preview | Q2 | 3.3 | 821 | 308 | $2 / $12 (prompt up to 200K; output includes thinking) | 0.0053 |
| gemini-3.1-pro-preview | Q3 | 4.5 | 818 | 354 | $2 / $12 (prompt up to 200K; output includes thinking) | 0.0059 |
| gemini-3.1-pro-preview | Q4 | 6.5 | 831 | 444 | $2 / $12 (prompt up to 200K; output includes thinking) | 0.0070 |
| gemini-3.1-pro-preview | Q5 | 5.1 | 829 | 317 | $2 / $12 (prompt up to 200K; output includes thinking) | 0.0055 |
| gemini-3.1-pro-preview | Q6 | 5.2 | 824 | 473 | $2 / $12 (prompt up to 200K; output includes thinking) | 0.0073 |
Totals for all six questions
| Model | Right | Seconds | Output tokens | Words in answers | Cost (USD) |
|---|---|---|---|---|---|
| claude-sonnet-5-5 | 6 of 6 | 9.4 | 660 | 297 | 0.0220 |
| gpt-6.1-sol | 6 of 6 | 13.6 | 212 | 118 | 0.0113 |
| gemini-3.1-pro-preview | 6 of 6 | 29.1 | 2,198 | 200 | 0.0363 |
How we tested
- We used two public Federal Reserve pages from September 16, 2026: the FOMC statement and the Implementation Note. We copied their text into one plain-text document of 3,634 characters.
- Each model got one request per question: a single user message with the instruction, the document inside
<document>tags, then the question. No system prompt, up to 2,048 output tokens, other settings left at each vendor’s defaults, no web search or tools. - One run per question per model on 10 October 2026, with no retries. Seconds are the wall-clock time of each request as measured by our script. Token counts are the ones each API reported; Gemini’s output counts include its thinking tokens.
- Cost is those token counts times each vendor’s list price on the day: Claude Sonnet 5.5 $2 per million input tokens and $10 per million output tokens; GPT-6.1 Sol $2 and $10 for inputs up to 272K tokens; Gemini 3.1 Pro Preview $2 and $12 for prompts up to 200K tokens, where Google’s pricing page lists the output price as “including thinking tokens”.
- We checked every answer against the two pages, and every extra fact the models added. Answers are shown unedited.
Source
Federal Reserve FOMC statement and Implementation Note, September 16, 2026
- Credit:
- Board of Governors of the Federal Reserve System
- License:
- Public domain (U.S. government work)
Run this comparison on your own files
Creatos is a desktop app for macOS and Windows. Add your own PDFs, notes and web pages as sources, ask GPT, Claude, Gemini and local Ollama models the same question, and read the answers side by side.
$69 once. No subscription.
Cloud model usage is billed by your provider. Local models run on your own machine.
FAQ
How do I compare AI models side by side?
Give each model the same document and the same question with the same settings, then read the answers next to each other and check each one against the source. Here we sent each question as its own request to three APIs and compared the answers, the time and the cost.
Which model was most accurate on this test?
None: all three got all six answers right, including the false-premise question and the one the document can’t answer. On a short document like this, the differences were cost, speed and answer length.
Why did Gemini 3.1 Pro cost more?
It produced about 2,200 output tokens for the six answers, most of them thinking tokens, and Google bills thinking at the output rate. GPT-6.1 Sol produced 212 and Claude Sonnet 5.5 produced 660.
Related tests
All model tests- Model comparisonClaude Haiku 5.5 vs Sonnet 5.5 on long PDFs: the 100K stepHaiku 5.5 and Sonnet 5.5 got the same six questions about a 73K-token jobs report and a 113K-token 10-K. Both got all six right; Haiku was 20x cheaper, then 4x.
- Model comparisonClaude vs GPT vs Gemini on a 39-page PDF: 9 answers checkedClaude Sonnet 5.5, GPT-6.1 Sol and Gemini 3.1 Pro each got the September 2026 BLS jobs report and three questions. We checked all nine answers against the PDF.
Model names are used for identification only. Creatos is not affiliated with OpenAI, Anthropic, Google, Meta, Alibaba or Ollama.