Creatos Logo
Compare
Buy License

Model comparison · one run each

Claude vs GPT vs Gemini on a 39-page PDF: 9 answers checked

Claude Sonnet 5.5, GPT-6.1 Sol and Gemini 3.1 Pro each got the September 2026 BLS jobs report and three questions. We checked all nine answers against the PDF.

Tested
Updated
Models
3

TL;DR

  • All nine answers were right: U-6 at 7.6% seasonally adjusted (not the 7.3% on the same row), $23.90 for leisure and hospitality hourly earnings, and no October forecast in the report.
  • Claude and Gemini named the tables and added a second earnings figure from Table B-8 ($21.34, also right). GPT gave the shortest answers and named one table.
  • GPT was the fastest in this run: 14.8 s for all three questions, against 21.0 s for Gemini and 26.5 s for Claude.
  • Each request carried the whole report, about 68,000 to 74,000 input tokens. Three questions cost $0.22 on Gemini, which reused a cached copy for questions 2 and 3, $0.44 on Claude and $0.51 on GPT, whose bill included cache writes.

Creatos asks GPT, Claude, Gemini and local models the same question about your own files, side by side. $69 once.

Download Creatos
The same 39-page jobs report and the same three questions in ChatGPT, Claude and Gemini, with every answer checked against the page it came from. It plays from YouTube when you press play. Watch on YouTube

Setup

App
A short script calling each vendor’s API directly (PDF text extracted with Creatos’s PDF parser)
Machine
Mac mini, Apple M4 (sends the requests; the models run in each vendor’s cloud)
System prompt
None
Max output tokens
8,192
Temperature
Provider default
Reasoning effort
Provider default
Web search and tools
Off
Runs
One per question per model

Task

Prompt, identical for every model

Three questions, each sent as its own request: 1. What was the U-6 rate in September 2026, seasonally adjusted? 2. What were average hourly earnings in leisure and hospitality in September 2026? 3. What unemployment rate does BLS forecast for October 2026?

Source: The Employment Situation, September 2026 (USDL-26-1549)

Each request was one user message: "<document>", the text extracted from the PDF with the same parser Creatos uses (175,300 characters from 39 pages, with no page numbers), "</document>", a blank line, then the question. No system prompt, no web search, no tools.

The answers, side by side

Same source, same prompt, one run each

claude-sonnet-5-5

Cloud

Anthropic

  • Total27 s
  • Cost$0.44

Q1. The seasonally adjusted U-6 rate in September 2026 was 7.6 percent, from Table A-15. That was down from 7.7 percent in August 2026 and 7.9 percent in July 2026.

U-6 counts the total unemployed, all people marginally attached to the labor force, and people employed part time for economic reasons. It is expressed as a percent of the civilian labor force plus all people marginally attached to the labor force.

Q2. In September 2026 (preliminary, seasonally adjusted), average hourly earnings in leisure and hospitality were $23.90 for all employees on private nonfarm payrolls (Table B-3). That was up from $23.76 in August 2026 and $23.65 in July 2026.

For production and nonsupervisory employees in the industry, average hourly earnings were $21.34 (Table B-8).

Q3. The document doesn't contain a BLS forecast for the October 2026 unemployment rate. The release only reports data for September 2026 and earlier, and BLS doesn't publish forecasts in the Employment Situation.

The relevant facts the document does give:

  • September 2026 unemployment rate: 4.2 percent, which it describes as little changed. The rate has stayed in a narrow range of 4.1 to 4.3 percent since March.
  • Next release: The Employment Situation for October 2026 is scheduled for Friday, November 6, 2026, at 8:30 a.m. (ET). That release will report the actual October figure.

Three separate requests, one per question; each answer is shown after its question number, unedited. Seconds per question: 5.3, 8.6, 12.6. The line that BLS doesn’t publish forecasts in the Employment Situation is true but is not in this PDF.

gpt-6.1-sol

Cloud

OpenAI

  • Total15 s
  • Cost$0.51

Q1. The seasonally adjusted U-6 rate in September 2026 was 7.6 percent.

Q2. Average hourly earnings in leisure and hospitality were $23.90 in September 2026 for all employees, seasonally adjusted (Table B-3).

Q3. The document does not provide a BLS forecast for October 2026’s unemployment rate. It reports September’s rate at 4.2% and states that the October Employment Situation report is scheduled for November 6, 2026, at 8:30 a.m. ET.

Three separate requests, one per question; each answer is shown after its question number, unedited. Seconds per question: 4.6, 6.9, 3.3.

gemini-3.1-pro-preview

Cloud

Google

  • Total21 s
  • Cost$0.22

Q1. Based on Table A-15 (Alternative measures of labor underutilization), the seasonally adjusted U-6 rate in September 2026 was 7.6 percent.

Q2. Based on the provided document, the average hourly earnings in leisure and hospitality in September 2026 were:

  • $23.90 for all employees (from Table B-3)
  • $21.34 for production and nonsupervisory employees (from Table B-8)

Q3. Based on the provided document, the Bureau of Labor Statistics (BLS) does not forecast or project the upcoming unemployment rate. The news release only reports actual historical data for September 2026 (which was 4.2 percent).

The document does note that the employment situation data for October 2026 is scheduled to be published on Friday, November 6, 2026.

Three separate requests, one per question; each answer is shown after its question number, unedited. Seconds per question: 6.8, 7.8, 6.4.

Measurements

Each answer checked against the PDF

QuestionRight answer (PDF page)claude-sonnet-5-5gpt-6.1-solgemini-3.1-pro-preview
U-6 rate, Sept 2026, seasonally adjusted7.6% (Table A-15, p. 26; the same row shows 7.3% not seasonally adjusted)Right: 7.6%, names A-15Right: 7.6%Right: 7.6%, names A-15
Average hourly earnings, leisure and hospitality, Sept 2026$23.90, preliminary (Table B-3, p. 33)Right: $23.90 (B-3), plus $21.34 (B-8)Right: $23.90 (B-3)Right: $23.90 (B-3), plus $21.34 (B-8)
BLS forecast for October 2026None in the report; the October release is scheduled for 6 Nov 2026 (p. 3)Right: no forecast; gives 6 NovRight: no forecast; gives 6 NovRight: no forecast; gives 6 Nov

Time, tokens and cost per request

ModelQuestionSecondsInput tokensOutput tokensCost (USD)
claude-sonnet-5-5Q15.372,9101430.1472
claude-sonnet-5-5Q28.672,9101340.1472
claude-sonnet-5-5Q312.672,9062100.1479
gpt-6.1-solQ14.668,174250.1707
gpt-6.1-solQ26.968,172580.1710
gpt-6.1-solQ33.368,170670.1711
gemini-3.1-pro-previewQ16.874,1954410.1537
gemini-3.1-pro-previewQ27.874,1947970.0327
gemini-3.1-pro-previewQ36.474,1924710.0287

How we tested

  1. We downloaded the BLS release as a PDF (39 pages, 580,018 bytes) and extracted its text with the PDF parser Creatos uses (pdf.js): 175,300 characters, every page in order, with no page numbers.
  2. Each model got one request per question: a single user message with the text inside <document> tags, then the question. No system prompt, up to 8,192 output tokens, temperature and reasoning left at each provider’s default, no web search or tools.
  3. One run per question per model on 7 October 2026, with no retries. Seconds are the wall-clock time of each request as measured by our script. Token counts are the ones each API reported; Gemini’s output count includes its thinking tokens.
  4. Cost is those token counts times each vendor’s list price on the day: $2 per million input tokens for all three, and $10 (Claude Sonnet 5.5, GPT-6.1 Sol) or $12 (Gemini 3.1 Pro) per million output tokens. OpenAI cached the report automatically on every request (about 68,170 tokens written each time, billed at 1.25x the input price, never read back in these three runs). Gemini read 69,611 tokens from its implicit cache on questions 2 and 3, billed at $0.20 per million; any cache storage charge is not included. Claude used no cache.
  5. We checked every answer against the PDF itself: the U-6 row of Table A-15 on page 26, the leisure and hospitality row of Table B-3 on page 33 and of Table B-8 on page 38, and the line on page 3 that the October release “is scheduled to be published on Friday, November 6, 2026”.

Page numbers are the PDF viewer’s. These are lookups with one right answer each; published tests that ask for summaries of long reports find more mistakes.

Source

The Employment Situation, September 2026 (USDL-26-1549)

Credit:
U.S. Bureau of Labor Statistics
License:
Public domain (U.S. government work)

Run this comparison on your own files

Creatos is a desktop app for macOS and Windows. Add your own PDFs, notes and web pages as sources, ask GPT, Claude, Gemini and local Ollama models the same question, and read the answers side by side.

$69 once. No subscription.

Cloud model usage is billed by your provider. Local models run on your own machine.

FAQ

Which AI reads PDFs best: ChatGPT, Claude or Gemini?

On these three lookups, none stood out: all three models were right every time, including the seasonal-adjustment trap and the question the PDF can’t answer. The differences were in length, speed and cost. Tests that ask for a summary of a long report find more errors, so check any figure you plan to repeat.

Why didn’t the models give page numbers?

The text they received has no page numbers in it, so they could only name the tables. In the Claude app, Anthropic’s help page says to use the page numbers your PDF viewer shows, not the ones printed on the page.

How much does it cost to ask one question about a 39-page PDF through an API?

Between 3 and 17 cents per question here. The whole report (68,000 to 74,000 tokens) goes with every question, so the first question costs the most; Gemini’s second and third questions cost about 3 cents each because it reused a cached copy of the report.

All model tests

Model names are used for identification only. Creatos is not affiliated with OpenAI, Anthropic, Google, Meta, Alibaba or Ollama.