What Is Claude Opus 5.5? What Changed, What It Costs, and Where It Falls Down
Anthropic's Claude Opus 5.5 costs 20% less per token than Opus 5, and Anthropic says it performs at the level of Claude Fable 5.1 on most work. What changed, what it costs on the API and on TulexAI, where it falls down, and every way to get it.
TL;DR: Claude Opus 5.5, released September 22, 2026, is Anthropic's newest Opus: $4/$20 per million API tokens, 20% less per token than Opus 5. Anthropic says it performs "at the level of Claude Fable 5.1 on most work". On TulexAI: 10 PT per ~750 tokens of prompt plus reply, from Basic ($11/mo). The catch: thinking is always on.
What is Claude Opus 5.5?
Claude Opus 5.5 is Anthropic's current Opus model. The launch announcement calls it "the first model in our new Claude 5.5 family", and Anthropic's docs sum it up as "For long-running agentic coding and knowledge work". Sonnet 5.5 and Haiku 5.5 are announced, not released.
It succeeds Claude Opus 5, which Anthropic's Opus 5 page now calls "a legacy model". Legacy isn't deprecated: Opus 5 stays active, with retirement not sooner than July 24, 2027, and on TulexAI it stays available at 12 PT. Opus 5.5 is an addition, not a replacement.
The specs, from Anthropic's Opus 5.5 model page:
- API id:
claude-opus-5-5, with no dated snapshot. - Context window: 1,000,000 tokens, billed at the standard rate across the full window. Max output: 128,000 tokens.
- Input and output: text, images and PDFs in, text out. Knowledge cutoff: June 2026.
- Thinking: adaptive and always on. Effort runs from low to max, and the default is medium.
What changed from Claude Opus 5
| What | Claude Opus 5 | Claude Opus 5.5 |
|---|---|---|
| API price per 1M tokens (input / output) | $5 / $25 | $4 / $20 |
| Cache reads per 1M tokens | $0.50 | $0.20 |
| Cache writes per 1M tokens (5-minute / 1-hour) | $6.25 / $10 | $5 / $8 |
| Effort when a request sets none | high | medium |
| Status at Anthropic | Legacy, still active | Current Opus |
| On TulexAI, per ~750 tokens | 12 PT | 10 PT |
- 20% cheaper per token. Anthropic calls it "20% less than Opus 5", and cache reads drop 60% (Anthropic pricing).
- The 40% is not a price cut. Anthropic says its tests show "at default settings it will cost 40% less than Opus 5 on typical workloads", comparing Opus 5.5 at medium effort with Opus 5 at high. On TulexAI the gap is 10 PT against 12 PT per ~750 tokens.
- Lower default effort, more thinking. Per Anthropic's What's new page, Opus 5.5 thinks "more per turn at a given effort level". TulexAI sends no effort setting, so it runs at medium here.
- Faster tokens, per Anthropic. Anthropic says it generates output more than 30% faster than Opus 5. That's tokens per second, not faster replies (see the first-token catch below).
The benchmarks: Anthropic's table, and the effort behind it
Anthropic's launch table is vendor-reported. Its footnote: "Unless otherwise noted, all Claude Opus 5.5 results use adaptive thinking at max effort." Terminal-Bench 4.0 ran at xhigh for Opus 5.5 and high for GPT-6 Astra, whose figures OpenAI reported. Where safeguards stepped in, Opus 4.8 or Opus 5 finished the task, which Anthropic says "likely reduces" Opus 5.5's scores.
| Benchmark (Anthropic's table; Opus 5.5 at max effort unless noted) | Opus 5.5 | Fable 5.1 | Opus 5 | GPT-6 Astra |
|---|---|---|---|---|
| Terminal-Bench 4.0 | 66.4% (xhigh) | 55.8% | 52.3% | 57.9% (high) |
| FrontierCode v1.1 main | 54.4% | 50.3% | 48.0% | 53.3% |
| CursorBench 4.0 | 57.8% | 51.8% | 46.6% | not reported |
| Humanity's Last Exam (with tools) | 67.7% | 65.6% | 63.6% | 57.2% |
| OSWorld 2.0 (partial) | 81.8% | 80.7% | 74.0% | not reported |
| GDPval-AA v2.1 (Elo) | 1846 | 1735 | 1708 | 1542 |
| Terminal-Bench-Science 0.1 | 58.7% | 52.6% | 29.0% | 64.6% |
| AutomationBench (Zapier) | 40.0% | 31.4% | 26.9% | 41.4% |
Anthropic hedges its own table: at this level, margins are "a less reliable guide to real-world differences", and "the gap between Opus 5.5 and Claude Fable 5.1 is narrower than these scores suggest". Its claim is that Opus 5.5 performs at the level of Fable 5.1 on most work. Where Astra has no score, there's no win to claim.
What TulexAI runs. That table is max effort (xhigh on Terminal-Bench 4.0); TulexAI runs medium. Anthropic's only medium-effort scores are 54.6% on FrontierCode v1.1 main, which it calls "higher than all other models", and 52.5% on CursorBench 4.0, against 51.8% for Fable 5.1 and 46.6% for Opus 5 at max. They're the closest Anthropic figures to a TulexAI reply, and still Anthropic's own.
The independent numbers: Artificial Analysis
Artificial Analysis (AA) scored Opus 5.5 on launch day. Index scores here are Intelligence Index v4.3.2, read September 24, 2026, with ranks out of 212 models. Our September 15 Astra posts used v4.3, and AA re-anchored its GDPval-AA Elo on September 19, so don't mix the two.
| Model | Index v4.3.2, max effort | Index v4.3.2, medium effort |
|---|---|---|
| Claude Opus 5.5 | 58 (#1) | 51 (#8) |
| Claude Fable 5.1 | 53 (#4) | 49 (#15) |
| GPT-6 Astra | 53 (#6) | 50 (#14) |
| Claude Opus 5 | 51 (#11) | 45 (#26) |
At max effort Opus 5.5 leads "by several points", in AA's words. The medium column is TulexAI's setting for Opus 5.5 only (TulexAI runs Opus 5 at its default, high), and AA ran Opus 5.5 with Anthropic's default fallback on, which TulexAI doesn't use.
- Cost per task: $1.34 at medium against Opus 5's $2.19. At max, about the same ($5.98 vs $5.86): Opus 5.5 used about 1.6x the output tokens, offset by lower per-token and cache prices.
- Hallucination rate (wrong answers as a share of all non-correct ones, lower is better): 59% at max, below Opus 5 (61%) and Fable 5.1 (73%) but above GPT-6 Astra (51%). At medium it's 68%, level with Fable 5.1 (69%) and well above Astra (47%).
- Terminal-Bench 4.0: 59.6% at max effort on AA's harness, "level with the leader GPT-6 Astra (xhigh)". A different setup from Anthropic's 66.4%, so don't compare the two.
- Coding agents: Claude Code with Opus 5.5 at max effort is #1 of 20 on AA's Coding Agent Index v1.5 (66), with about 9% of attempts finished by a fallback model after a safety refusal.
Where Claude Opus 5.5 falls down
- Thinking eats short output budgets. It can't be switched off and shares the output budget with the answer. In our live check on Anthropic's API on September 22, a hard prompt capped at 900 output tokens spent 496 on thinking and hit the cap. On TulexAI, replies cap at 8,192 output tokens, and reasoning counts toward the tokens you're billed.
- Slow to start. AA partly backs Anthropic's speed claim: at medium effort on Anthropic's API, Opus 5.5 streamed 75.2 tokens per second to Opus 5's 56.6 (+33%). But its first answer token took 22.17 seconds against 4.84, and a full response 28.83 against 13.68. Those are AA's tests, not measured through TulexAI.
- Flagged cyber and biology prompts are declined on TulexAI. Anthropic says "most cybersecurity tasks will be re-routed to Opus 4.8", and the Claude apps send flagged biology requests to Opus 5. On the API a flagged request comes back as a refusal; re-routing happens only if the caller opts in, and the response names the model that served it (refusals doc). TulexAI doesn't opt in, so a flagged prompt is declined with a clear message. Ordinary bug finding and fixing stays allowed.
- No forced tool choice. Tools work, but Anthropic's API rejects a forced tool choice with an HTTP 400, as on Fable 5.1. On the TulexAI API it's served as auto, so the model decides whether to call.
- GPT-6 Astra wins two of Anthropic's rows: Terminal-Bench-Science 0.1, and Zapier's AutomationBench, where Zapier counted safeguard interventions as failures. AA says Opus 5.5 "remains behind on CritPt, AA-LCR, and GDP.pdf".
- The hardest work still goes to Fable 5.1. Anthropic's docs say "start with Claude Opus 5.5 for most workloads", but "Use Claude Fable 5.1 for demanding reasoning and long-horizon agentic work", or when Opus 5.5 at higher effort still falls short. On TulexAI you can't raise the effort, so Fable 5.1 is the step up.
Opus 5.5 vs Opus 5 vs Fable 5.1 vs GPT-6 Astra: who should pick which?
- Pick Claude Opus 5.5 when you want Anthropic's current Opus for coding, long documents and multi-step work, at 10 PT per ~750 tokens from Basic up.
- Pick Claude Opus 5 when your prompts are already tuned to it. It stays on TulexAI at 12 PT, Basic and up.
- Pick Claude Fable 5.1 when Opus 5.5 falls short on demanding reasoning or long agent runs. It's 25 PT, Pro and up, metered by the frontier allowance.
- Pick GPT-6 Astra when fewer hallucinations matter most (47% vs 68% at medium, per AA). Also 25 PT, Pro and up.
- Pick Claude Sonnet 5 when the job doesn't need Opus depth: 7 PT per ~750 tokens.
Head-to-heads: Claude Opus 5 vs Claude Opus 5.5 and Claude Fable 5.1 vs Claude Opus 5.5. The spec sheet and prompts live on the Claude Opus 5.5 model page.
How to get Claude Opus 5.5
Per Anthropic's pricing page, the Claude Code docs and each cloud's own docs, read September 24, 2026:
| Where | Price | What you get | Limits and catches |
|---|---|---|---|
| Claude Free | $0 | No Opus: Sonnet and Haiku only | - |
| Claude Pro | $20/mo, or $17/mo billed yearly | Opus in Claude; Claude Code defaults to Opus 5.5 | Five-hour and weekly limits |
| Claude Max | From $100/mo | 5x or 20x Pro usage per session | Higher output limits, priority access |
| Claude Team | $20/seat/mo yearly ($25 monthly); premium seats $100 ($125) | Opus on both seat types | 2 to 150 people |
| Claude Enterprise | US$20/seat/mo yearly, plus usage at API rates | Opus | Admins can switch individual models off |
| Anthropic API | $4 / $20 per 1M tokens | Every effort level, 1M context, 128K output | Forced tool choice rejected |
| Amazon Bedrock | Billed through AWS Marketplace | anthropic.claude-opus-5-5 | Geo and Global inference only |
| Google Cloud | - | claude-opus-5-5, generally available | US and EU multi-region, global endpoint |
| Microsoft Foundry | $4 / $20 per 1M tokens | Generally available, hosted on Azure | In Claude Code, pick your Opus 5.5 deployment: Foundry defaults to an older model |
| TulexAI | Basic $11/mo and up (also Pro, VIP, Elite) | Opus 5.5 in chat, plus tool calling through the TulexAI API | 10 PT per ~750 tokens; medium effort; 8,192-token replies; flagged cyber and biology prompts declined |
- Claude Code needs v2.1.280 or later (
claude update), and it shares your Claude plan's usage limits. - Limits went up, not away. Anthropic is "increasing five-hour usage limits on Pro, Max, Team, and seat-based Enterprise plans" without giving a figure; five-hour and weekly limits remain. On Claude Pro, Opus 5.5 is inside plan limits, while Fable 5.1 runs on usage credits.
What Claude Opus 5.5 costs
On the API, it's $4 in and $20 out per million tokens, with no promotional or end-date wording. Cache reads are $0.20, cache writes $5 (5-minute) or $8 (1-hour), batch $2 / $10, and US-only inference 1.1x. A 250-token prompt with a 500-token reply costs $0.001 + $0.01 = $0.011, before thinking tokens, which count as output. Fast mode ($8 / $40) is a separate research preview that TulexAI doesn't use.
On TulexAI, it's 10 PT per ~750 tokens of prompt plus reply, reasoning included, with a 10 PT minimum per request, so a long or reasoning-heavy answer costs more than 10 PT. That undercuts Claude Opus 5 at 12 PT and Claude Fable 5.1 at 25 PT.
It's premium tier, so free accounts can't run it. It's on Basic ($11/mo), Pro, VIP and Elite; on Basic it shares Claude Opus 5's small premium-text allowance before requests downgrade. The honest trade: medium effort with no dial, and 8,192-token replies instead of 128K. Need max effort or long outputs? Use the API. Want Anthropic's newest Opus next to GPT, Gemini and image, video and voice models on one plan? That's us. See the pricing page or our Claude alternative breakdown.
How to use Claude Opus 5.5 on TulexAI
- Give it the whole job up front: the goal, the constraints and what done looks like, in the first message.
- Leave room for thinking. Reasoning and answer share the 8,192-token cap, so ask for long deliverables one section at a time.
- Write tool rules into the prompt, since you can't force a call. Our agents docs cover the setup.
Frequently Asked Questions
How much does Claude Opus 5.5 cost?
On Anthropic's API, $4 per million input tokens and $20 per million output, 20% less per token than Claude Opus 5. On TulexAI, 10 PT per ~750 tokens of prompt plus reply, reasoning included, with a 10 PT minimum, on Basic ($11/mo) and every plan above it.
Is Claude Opus 5.5 better than Claude Opus 5?
It beats Opus 5 on every row of Anthropic's own table, run at max effort (xhigh on Terminal-Bench 4.0), and on Artificial Analysis Intelligence Index v4.3.2 (58 vs 51 at max, 51 vs 45 at medium). It's also cheaper per token. The trade-off is a slower first token: 22.17 seconds against 4.84 at medium in AA's test.
Is Claude Opus 5.5 better than Claude Fable 5.1?
Anthropic says it performs at the level of Fable 5.1 on most work, and that its table overstates the gap. AA's Index v4.3.2 has Opus 5.5 at 58 and Fable 5.1 at 53, both at max effort. Anthropic's docs still recommend Fable 5.1 for demanding reasoning and long-horizon agentic work.
Is Claude Opus 5 going away?
No. Anthropic lists Opus 5 as legacy but still active, with retirement not sooner than July 24, 2027. On TulexAI it stays available at 12 PT per ~750 tokens, Basic and up.
Can I turn off thinking or set the effort level?
Thinking can't be turned off: adaptive thinking is always on for Opus 5.5. On Anthropic's API you can set effort from low to max. TulexAI sends no effort setting, so Opus 5.5 runs at Anthropic's default, medium.
Why was my cybersecurity prompt declined?
Anthropic's safeguards flag some cyber and biology requests, and the Claude apps hand those to an older model. TulexAI doesn't use that fallback, so a flagged request is declined with a clear message. Finding and fixing bugs in normal development stays allowed.
Create your TulexAI account and run your hardest prompt through Claude Opus 5.5, then Claude Opus 5, on one subscription from $11/mo.
Ready to consolidate your AI tools?
40+ AI models - GPT-5.6, Claude Opus 5, Gemini, Flux, Sora & more. One subscription from $11/mo.
Try 1 Free PromptContinue reading
What Is GPT-6 Astra? OpenAI's New Flagship, and How to Actually Get It
OpenAI's GPT-6 Astra ties Claude Fable 5.1 on the neutral scoreboard, but ChatGPT Plus only gets it in Work and Codex. What it is, every way to access it, what it costs, and the catches.
Can GPT-6 Astra Run Agents? What We Found Testing It Live
OpenAI launched GPT-6 Astra on computer use and browsing, but on Chat Completions it has no function calling. The two 400 errors we hit on September 8, 2026, and which models to use for agents instead.