Model comparison · one run each
Claude Haiku 5.5 vs Sonnet 5.5 on long PDFs: the 100K step
Haiku 5.5 and Sonnet 5.5 got the same six questions about a 73K-token jobs report and a 113K-token 10-K. Both got all six right; Haiku was 20x cheaper, then 4x.
- Tested
- Updated
- Models
- 2
TL;DR
- Both models answered all six questions correctly, including a seasonal-adjustment trap, a margin that needs figures from two pages, and two questions the documents can’t answer.
- On the 73K-token jobs report, a question cost $0.0074 on Haiku 5.5 and $0.147 on Sonnet 5.5, about 20 times more.
- On the 113K-token 10-K, Haiku’s prompts were over 100,000 tokens and paid its higher rate: about $0.057 a question, against $0.229 on Sonnet, about 4 times more.
- Haiku was faster in this run: 13.2 s for all six questions, against 20.2 s for Sonnet.
Creatos asks GPT, Claude, Gemini and local models the same question about your own files, side by side. $69 once.
Download CreatosSetup
- App
- A short script calling Anthropic’s API directly (PDF text extracted with Creatos’s PDF parser)
- Machine
- Mac mini, Apple M4 (sends the requests; the models run in Anthropic’s cloud)
- System prompt
- None
- Max output tokens
- 8,192
- Temperature
- Default
- Thinking
- Default
- Prompt caching
- Off
- Web search and tools
- Off
- Runs
- One per question per model
Task
Prompt, identical for every model
Six questions, each sent as its own request. About the BLS Employment Situation, September 2026: 1. What was the U-6 rate in September 2026, seasonally adjusted? 2. What were average hourly earnings in leisure and hospitality in September 2026? 3. What unemployment rate does BLS forecast for October 2026? About the Alphabet 10-K for 2025: 4. What were YouTube ads revenues in 2025? 5. What was Google Cloud's operating margin in 2025? 6. What total revenue does Alphabet forecast for 2026?
Source: Alphabet Inc. Form 10-K for fiscal year 2025
Each request was one user message: "<document>", the text extracted from the PDF with the same parser Creatos uses, "</document>", a blank line, then the question. With the question, the prompt came to about 72,900 tokens for the jobs report and about 113,000 tokens for the 10-K. No system prompt, no web search, no tools, no prompt caching.
The answers, side by side
Same source, same prompt, one run each
claude-haiku-5-5
CloudAnthropic
- Total13 s
- Cost$0.19
Q1. The U-6 measure of labor underutilization was 7.6 percent in September 2026, seasonally adjusted (Table A-15). That was down from 7.7 percent in August 2026.
Q2. Average hourly earnings in leisure and hospitality were $23.90 in September 2026, seasonally adjusted, for all employees on private nonfarm payrolls (Table B-3). That was up from $23.76 in August 2026.
Q3. The document doesn't include a forecast for October 2026. BLS reports only past and current data in this release, and it gives no projection for future months.
The most recent figure is the 4.2 percent unemployment rate for September 2026, which the release says changed little from August. The rate has stayed between 4.1 and 4.3 percent since March 2026. The October 2026 Employment Situation is scheduled for release on Friday, November 6, 2026, at 8:30 a.m. ET, and that report will provide the October unemployment rate.
Q4. YouTube ads revenues were $40,367 million (about $40.4 billion) in 2025, up from $36,147 million in 2024. That is an increase of $4.2 billion, or roughly 12%.
Q5. Google Cloud's operating margin in 2025 was about 23.7%. It had operating income of $13,910 million on revenues of $58,705 million (13,910 ÷ 58,705 ≈ 0.237).
Q6. The filing doesn't give a total revenue forecast for 2026. Alphabet's 10-K contains no revenue guidance. Its forward-looking statements are qualitative, and the 2025 figures are actuals.
Relevant forward-looking items in the document include:
- Capital expenditures: Alphabet expects to "significantly increase" its technical infrastructure investment in 2026 relative to 2025, when capex was $91.4 billion.
- Revenue backlog: As of December 31, 2025, remaining performance obligations were $242.8 billion, primarily related to Google Cloud. Just over 50% is expected to be recognized over the next 24 months.
- Qualitative trends: Management expects AI offerings, growth in non-advertising revenue, and rising infrastructure costs to affect future growth and margins, but it gives no numbers.
For a 2026 revenue outlook, check Alphabet's quarterly earnings releases or earnings call transcripts, since companies typically provide guidance there rather than in the 10-K.
Six separate requests, one per question; each answer is shown after its question number, unedited. Q1 to Q3 are about the BLS jobs report, Q4 to Q6 about the Alphabet 10-K. Seconds per question: 2.3, 1.8, 2.6, 1.4, 1.7, 3.4.
claude-sonnet-5-5
CloudAnthropic
- Total20 s
- Cost$1.13
Q1. The seasonally adjusted U-6 rate in September 2026 was 7.6 percent (Table A-15). That was down from 7.7 percent in August 2026 and 7.9 percent in July 2026.
U-6 counts total unemployed people, all people marginally attached to the labor force, and those employed part time for economic reasons. It is shown as a percent of the civilian labor force plus all people marginally attached to the labor force.
Q2. In September 2026, average hourly earnings for all employees in leisure and hospitality on private nonfarm payrolls were $23.90 (seasonally adjusted, preliminary), per Table B-3. This was up from $23.76 in August 2026 and $23.65 in July 2026.
For production and nonsupervisory employees in the same industry, average hourly earnings were $21.34 in September 2026 (Table B-8).
Q3. The release doesn't contain a BLS forecast for the October 2026 unemployment rate. BLS reports measured data and doesn't publish projections in this report.
The document says only that the Employment Situation for October 2026 is scheduled for release on Friday, November 6, 2026, at 8:30 a.m. (ET).
For context, the latest figures are:
- September 2026 unemployment rate: 4.2 percent (7.1 million unemployed)
- Recent range: The rate has stayed between 4.1 and 4.3 percent since March.
- Recent monthly readings (seasonally adjusted): 4.3 percent in May, 4.2 in June, 4.1 in July and August, and 4.2 in September.
The release also notes that the 90-percent confidence interval for monthly changes in the unemployment rate is about ±0.3 percentage point (at a 6.0 percent rate), so small month-to-month movements are often not statistically significant.
Q4. YouTube ads revenues in 2025 were $40,367 million (about $40.4 billion). That is up from $36,147 million in 2024, an increase of about $4.2 billion, or roughly 12%.
Alphabet said the growth came from its direct response advertising products, followed by its brand advertising products. Both benefited from higher advertiser spending.
Q5. Google Cloud's operating margin in 2025 was about 23.7%.
- Google Cloud operating income: $13,910 million
- Google Cloud revenues: $58,705 million
- $13,910 ÷ $58,705 ≈ 23.7%
This is up from about 14.1% in 2024, when operating income was $6,112 million on revenues of $43,229 million.
Q6. The 10-K doesn't give a total revenue forecast for 2026. Alphabet doesn't issue revenue guidance in this filing. The forward-looking statements are qualitative and don't project a specific figure.
The filing does say a few things about 2026:
- Capital expenditures: Alphabet expects to "significantly increase, relative to 2025," its investment in technical infrastructure, including servers, network equipment, and data centers. Capex was $91.4 billion in 2025. No 2026 dollar amount is given.
- Infrastructure costs: Depreciation, energy, equipment, and network capacity costs are expected to increase significantly.
- Revenue mix: Revenues from cloud, subscriptions, platforms, and devices are expected to keep growing faster than advertising, with generally lower margins.
- Pending acquisitions: Wiz ($32.0 billion) and Intersect ($4.8 billion plus assumed debt) are both expected to close in 2026, subject to regulatory approvals.
- Revenue backlog: $242.8 billion of remaining performance obligations as of December 31, 2025, primarily Google Cloud. Just over 50% is expected to be recognized over the next 24 months.
For reference, 2025 revenue was $402.8 billion, up 15% from $350.0 billion in 2024. The filing cautions that past results may not indicate future performance.
If you need an actual 2026 revenue figure, you would have to look at analyst estimates or management commentary on earnings calls, since neither appears in this document.
Six separate requests, one per question; each answer is shown after its question number, unedited. Q1 to Q3 are about the BLS jobs report, Q4 to Q6 about the Alphabet 10-K. Seconds per question: 2.7, 2.8, 3.2, 3.3, 3.9, 4.3. Every extra fact in these answers is in the documents except two details in Q6: the filing says “certain” of the cloud, subscription, platform and device revenues grow faster than advertising, not all of them, and its wording is “our past results should not be relied on as an indication of our future performance”.
Measurements
Each answer checked against the PDF
| Question | Right answer (PDF page) | claude-haiku-5-5 | claude-sonnet-5-5 |
|---|---|---|---|
| Jobs report: U-6 rate, Sept 2026, seasonally adjusted | 7.6% (Table A-15, p. 26; the same row shows 7.3% not seasonally adjusted) | Right: 7.6%, names A-15 | Right: 7.6%, names A-15 |
| Jobs report: average hourly earnings, leisure and hospitality, Sept 2026 | $23.90, preliminary (Table B-3, p. 33) | Right: $23.90 (B-3) | Right: $23.90 (B-3), plus $21.34 (B-8) |
| Jobs report: BLS forecast for October 2026 | None in the report; the October release is scheduled for 6 Nov 2026 (p. 3) | Right: no forecast; gives 6 Nov | Right: no forecast; gives 6 Nov |
| 10-K: YouTube ads revenues, 2025 | $40,367 million (revenue table, p. 34; 2024 was $36,147 million) | Right: $40,367 million | Right: $40,367 million |
| 10-K: Google Cloud operating margin, 2025 | 23.7%: $13,910 million operating income (p. 37) on $58,705 million revenues (p. 34) | Right: 23.7%, shows the math | Right: 23.7%, shows the math, adds 14.1% for 2024 |
| 10-K: Alphabet’s total revenue forecast for 2026 | None in the filing | Right: not stated | Right: not stated |
Time, tokens and cost per request
| Model | Question | Seconds | Input tokens | Output tokens | Price tier | Cost (USD) |
|---|---|---|---|---|---|---|
| claude-haiku-5-5 | Q1 | 2.3 | 72,910 | 210 | $0.10 / $0.50 (prompt up to 100K) | 0.0074 |
| claude-haiku-5-5 | Q2 | 1.8 | 72,910 | 135 | $0.10 / $0.50 (prompt up to 100K) | 0.0074 |
| claude-haiku-5-5 | Q3 | 2.6 | 72,906 | 306 | $0.10 / $0.50 (prompt up to 100K) | 0.0074 |
| claude-haiku-5-5 | Q4 | 1.4 | 113,040 | 68 | $0.50 / $2.50 (prompt over 100K) | 0.0567 |
| claude-haiku-5-5 | Q5 | 1.7 | 113,042 | 166 | $0.50 / $2.50 (prompt over 100K) | 0.0569 |
| claude-haiku-5-5 | Q6 | 3.4 | 113,043 | 506 | $0.50 / $2.50 (prompt over 100K) | 0.0578 |
| claude-sonnet-5-5 | Q1 | 2.7 | 72,910 | 141 | $2 / $10 | 0.1472 |
| claude-sonnet-5-5 | Q2 | 2.8 | 72,910 | 140 | $2 / $10 | 0.1472 |
| claude-sonnet-5-5 | Q3 | 3.2 | 72,906 | 317 | $2 / $10 | 0.1490 |
| claude-sonnet-5-5 | Q4 | 3.3 | 113,040 | 119 | $2 / $10 | 0.2273 |
| claude-sonnet-5-5 | Q5 | 3.9 | 113,042 | 128 | $2 / $10 | 0.2274 |
| claude-sonnet-5-5 | Q6 | 4.3 | 113,043 | 517 | $2 / $10 | 0.2313 |
How we tested
- We used two public PDFs: the BLS release The Employment Situation, September 2026 (39 pages) and Alphabet’s Form 10-K for 2025 (99 pages). We extracted their text with the PDF parser Creatos uses (pdf.js), every page in order, with no page numbers.
- Each model got one request per question: a single user message with the text inside
<document>tags, then the question. No system prompt, up to 8,192 output tokens, other settings left at Anthropic’s defaults, no prompt caching, no web search or tools. - One run per question per model on 8 October 2026, with no retries. Seconds are the wall-clock time of each request as measured by our script. Token counts are the ones the API reported; Haiku’s output counts include its thinking tokens.
- Cost is those token counts times Anthropic’s list price on the day. Claude Sonnet 5.5 costs $2 per million input tokens and $10 per million output tokens at any length. Claude Haiku 5.5 costs $0.10 and $0.50 for a prompt up to 100,000 tokens, and $0.50 and $2.50 for a prompt over 100,000 tokens. Anthropic’s pricing page says “Claude Haiku 5.5 is priced by prompt length: a prompt of over 100,000 tokens pays higher prices”, so the 10-K prompts are priced at the higher rate in full.
- We checked every answer against the PDFs themselves, and every extra fact the models added against the extracted text.
Page numbers are the PDF viewer’s. These are lookups with one right answer each, so they say little about harder reasoning tasks.
Source
Alphabet Inc. Form 10-K for fiscal year 2025
- Credit:
- Alphabet Inc.
- License:
- Public SEC filing
Run this comparison on your own files
Creatos is a desktop app for macOS and Windows. Add your own PDFs, notes and web pages as sources, ask GPT, Claude, Gemini and local Ollama models the same question, and read the answers side by side.
$69 once. No subscription.
Cloud model usage is billed by your provider. Local models run on your own machine.
FAQ
Is Claude Haiku 5.5 as accurate as Sonnet 5.5 for questions about documents?
On these six lookups, yes: both got every answer right, including a trap on the same table row and two questions the documents can’t answer. Sonnet added more context to its answers. This was one run per question, so it doesn’t tell you how they compare on harder reasoning.
What happens to Claude Haiku 5.5’s price over 100K tokens?
A prompt of over 100,000 tokens pays $0.50 per million input tokens and $2.50 per million output tokens instead of $0.10 and $0.50, and the rate is set by the length of the whole prompt. Here, one question went from about $0.007 on the 73K-token report to about $0.057 on the 113K-token 10-K.
Is Haiku 5.5 still cheaper than Sonnet 5.5 over 100K tokens?
Yes. Sonnet 5.5 has one price at any length, $2 and $10 per million tokens. On the 10-K, Haiku cost about a quarter as much per question; on the shorter jobs report, about a twentieth.
Related tests
All model testsModel names are used for identification only. Creatos is not affiliated with OpenAI, Anthropic, Google, Meta, Alibaba or Ollama.