Claude Opus 5.5: the Opus to start with, and when to reach for a sibling
Anthropic released Claude Opus 5.5 on September 22, 2026, calling it "the first model in our new Claude 5.5 family", at $4 input and $20 output per million tokens, 20% below Opus 5 per token. On TulexAI it sits beside Opus 5 rather than replacing it, so this page covers which jobs to start here and when a sibling is the better call.
- Cost
- 10 PT
- Per
- ~750 tokens
- Tier
- premium
- From
- $11/mo
Claude Opus 5.5 is Anthropic’s current Opus, summed up in its docs as "For long-running agentic coding and knowledge work". On TulexAI it costs 10 PT per 750 tokens of prompt plus reply, below Opus 5’s 12, runs at Anthropic’s default medium effort and caps replies at 8,192 tokens. For the hardest reasoning, Anthropic’s docs point to Fable 5.1.
What Claude Opus 5.5 is best at
- Agentic coding plans and multi-step implementation, the job Anthropic’s docs name first. Anthropic reports that at medium effort, its default and the setting TulexAI runs, it scored 54.6% on FrontierCode v1.1 main and 52.5% on CursorBench 4.0, the second above Anthropic’s own max-effort figures for Fable 5.1 (51.8%) and Opus 5 (46.6%).
- Refactor plans that span a codebase: an inventory of every call site, a merge order that keeps each batch shippable, and a list of the cases that cannot be changed mechanically, all written before any code.
- Knowledge work built from material you paste: memos, briefs, analyses and reports, the second job Anthropic’s docs name. Artificial Analysis measured 1846 Elo on GDPval-AA v2.1 at max effort against 1735 for Fable 5.1 (Intelligence Index v4.3.2, read 2026-09-24).
- Tool-calling agents on TulexAI’s API at a lower rate than the frontier models: it supports tools, unlike GPT-6 Astra here, and costs 10 PT per 750 tokens against 25 for Claude Fable 5.1 and GPT-6 Astra.
Where it falls down
- Thinking is always on and cannot be turned off. It comes out of the same output budget as the answer and is billed with it, so a hard prompt spends part of every reply reasoning before a word of the answer appears.
- Replies are capped at 8,192 output tokens on TulexAI. Anthropic allows up to 128,000, but the platform clamp applies, so a long document or a multi-file change has to be requested in parts.
- There is no effort control on TulexAI. It runs at Anthropic’s default, medium, where Artificial Analysis scored it 51 on Intelligence Index v4.3.2 against 58 at max effort (read 2026-09-24). Most of Anthropic’s headline benchmarks were taken at max, so they describe the model’s ceiling, not a TulexAI reply.
- Flagged requests are declined, not re-routed. Anthropic’s safeguards flag some cybersecurity, biology and frontier LLM development work; Claude’s apps switch those requests to Opus 4.8 or Opus 5, and Anthropic’s API re-routes them to another model only when the caller opts in. TulexAI does not opt in, so here a flagged request comes back declined.
- A forced tool choice is not honoured. Anthropic rejects it for this model, and TulexAI’s API serves it as auto, so the model decides for itself whether to call the tool.
- It is slow to start. At medium effort Artificial Analysis measured 22.17 seconds to first token against 4.84 seconds for Opus 5 at the same effort, on Anthropic’s API rather than through TulexAI. Once it starts it streams faster than Opus 5.
- Its measured hallucination rate at medium effort is 68%, about level with Fable 5.1 (69%) and well above GPT-6 Astra (47%), per Artificial Analysis on Intelligence Index v4.3.2. The rate is lower at max effort (59%), which TulexAI does not run.
- It is not ahead everywhere. In Anthropic’s own table, with Opus 5.5 at max effort, GPT-6 Astra led on Terminal-Bench-Science 0.1 (64.6% to 58.7%, Astra’s score as OpenAI reports it) and on Zapier’s AutomationBench (41.4% to 40.0%, run without fallback models). Among the Intelligence Index v4.3.2 evals, Artificial Analysis has it behind on CritPt, AA-LCR and GDP.pdf.
Prompts for Claude Opus 5.5
Copy one, swap the subject, keep the structure - the constraint clauses are what make these land.
Turn a feature request into a plan a coding agent can execute
Claude Opus 5.5I will hand this plan to a coding agent one step at a time. Feature: add per-team usage limits to our billing service. Break it into numbered steps small enough to finish and test in one sitting. For each step give: the files it touches, the change in one or two sentences, the test that proves it works, and what must already be true before it starts. Flag every step that changes a database schema or a public API. Do not write code, and keep the whole plan under 1,500 words.
Replies stop at 8,192 tokens on TulexAI and thinking comes out of the same budget, so a word limit and a no-code rule keep the reply spent on the plan itself. A step with its own test and precondition is something an agent loop can run and a reviewer can check.
Plan a codebase-wide refactor before touching a file
Claude Opus 5.5Below are the call sites of our old logging helper, each with its file path and line number. We want every call moved to the new structured logger without changing behaviour. First list the call sites grouped by file, with the fields each one logs. Then propose an order to migrate the files in, so that each batch can merge on its own with the tests still passing. Name every call that cannot be migrated mechanically and say why. Do not rewrite any files yet.
The inventory and the merge order are the parts that need the whole codebase in view, and they fit in one reply. The rewrites do not fit in 8,192 tokens, so they follow file by file in later requests. Paste the approved plan into each one: on TulexAI’s web chat every message reaches the model on its own, without the earlier turns.
Turn a pile of pasted material into a decision memo
Claude Opus 5.5Below are notes from eight customer calls, our pricing table and last quarter’s churn summary. Write a one-page decision memo on whether to raise the price of our mid-tier plan. Structure: the decision, the three strongest points for and against with the source of each, what we do not know, and the cheapest way to find out. Quote the notes directly wherever you rely on them, and write "not in the material" instead of filling a gap.
Knowledge work is one of the two jobs Anthropic’s docs name, and its failure is a fluent memo resting on facts nobody supplied. At medium effort Artificial Analysis measured a 68% hallucination rate, so a source beside every point and an explicit "not in the material" are doing real work.
Get a long document in parts that stay consistent
Claude Opus 5.5Write a technical design document for moving our nightly CSV export to an event-driven sync, in four parts. Send only part 1 now: goals, non-goals and the data model. End part 1 with a numbered outline of parts 2 to 4, one line each, then stop. I will paste that outline into each later request, one part at a time: keep every name from it unchanged.
A reply here stops at 8,192 tokens, reasoning included, and each web message reaches the model on its own, without the earlier turns. Fixing the outline in the first reply, then pasting it into every later request for one part at a time, means no section is cut off halfway and the names stay consistent.
Steer a tool-calling agent without forcing a tool
Claude Opus 5.5You are a support triage agent with three tools: search_tickets, get_customer and add_note. For each new ticket, call get_customer first, then search_tickets for up to three similar past tickets, then call add_note once with the likely category, the matching ticket ids and a suggested first reply. Never reply to the customer directly. If any tool returns an error, add a note naming the tool that failed and stop.
Opus 5.5 takes tools on TulexAI’s API, but a forced tool choice is served as auto, so you cannot make it call a specific tool through the request. Writing the call order into the instructions is the reliable substitute, and a stop rule on errors keeps a failed loop from spending tokens.
Craft notes
- 01Start here and step up only when it falls short. That is the order Anthropic’s docs give, and on TulexAI the step to Fable 5.1 costs 25 PT per 750 tokens against 10.
- 02Ask for the plan before the code. A plan you can correct costs far less than a rewrite that runs into the 8,192-token reply cap halfway through a file.
- 03Put a length in the prompt. Thinking and answer share one output budget and are billed together, so a stated length keeps both in check.
- 04Write the tool order into your instructions instead of forcing a tool. A forced tool_choice runs as auto on this model, so the instructions are what actually steer the calls.
- 05Paste the relevant files, not the whole repository. What you send is billed as well as what comes back, and each chat message has its own length limit.
- 06Keep offensive security testing off this model. A flagged cyber request is declined on TulexAI rather than handed to another model, though ordinary bug finding and fixing is fine.
When to use something else
Claude Opus 5.5 is one of 30 text models on the same Platform Token budget, so switching costs nothing but the tokens.
Your API integration forces a specific tool with tool_choice, which Opus 5 honours and Opus 5.5 serves as auto, or your prompts and review checklists are already tuned on Opus 5. It stays on every plan at 12 PT per 750 tokens, and Anthropic lists it as legacy with retirement no sooner than July 24, 2027.
Opus 5.5 still falls short on demanding reasoning or long-horizon agentic work. Anthropic’s docs send that work to Fable 5.1, and since TulexAI does not raise Opus 5.5’s effort, Fable 5.1 is the step up here: 25 PT per 750 tokens, Pro plan and up, metered by the frontier allowance.
The task is ordinary drafting, extraction, summarising or code review you can check at a glance. Claude Sonnet 5 costs 7 PT per 750 tokens, is on every plan from Basic, and does not draw on the premium text allowance Opus 5.5 uses.
You want a model from a different vendor to check a plan or a claim and do not need tools. On Intelligence Index v4.3.2, Artificial Analysis measured GPT-6 Astra’s hallucination rate below Opus 5.5’s at medium (47% to 68%) and at max effort (51% to 59%). It costs 25 PT per 750 tokens, Pro plan and up.
Frequently asked questions
Anthropic’s current Opus, released on September 22, 2026 as the first of its Claude 5.5 family; Sonnet 5.5 and Haiku 5.5 are announced but not yet out. On Anthropic’s API it has a 1,000,000-token context window, up to 128,000 output tokens, a June 2026 knowledge cutoff and thinking that is always on. It takes text and images and replies in text.
It costs 10 Platform Tokens for every 750 tokens of prompt plus reply, reasoning included, with a 10 PT minimum per request, and it is on every plan from Basic at $11 a month. Claude Opus 5 costs 12 PT and Claude Fable 5.1 25 PT on the same basis. It is a premium-tier model, so free accounts cannot run it, and on Basic it shares a small premium text allowance with Opus 5 before requests downgrade.
In Anthropic’s own benchmark table it scores above Opus 5 on every row, with Opus 5.5 at max effort (xhigh on Terminal-Bench 4.0). Artificial Analysis scored it 58 to Opus 5’s 51 on Intelligence Index v4.3.2, both at max effort. Anthropic says that "at default settings it will cost 40% less than Opus 5 on typical workloads"; on TulexAI the difference is 10 PT against 12 per 750 tokens. On TulexAI, though, Opus 5.5 answers at medium effort, where Artificial Analysis scored it 51, the same as Opus 5 at max, so expect a smaller step than the max-effort gap suggests. Opus 5 stays available here.
Start with Opus 5.5. Anthropic says it "performs at the level of Claude Fable 5.1 on most work", and its docs send demanding reasoning and long-horizon agentic work to Fable 5.1. Artificial Analysis scored Opus 5.5 at 58 and Fable 5.1 at 53 on Intelligence Index v4.3.2, both at max effort. On TulexAI, Opus 5.5 costs 10 PT per 750 tokens from Basic; Fable 5.1 costs 25 PT from Pro and draws on the frontier allowance.
It comes back declined, with an error that says so, and no other model answers it. Anthropic’s safeguards flag some cybersecurity, biology and frontier LLM development requests. Claude’s apps switch those to Opus 4.8 or Opus 5, and Anthropic’s API re-routes them to another model only when the caller opts in, then names the model that answered. TulexAI does not opt in. Ordinary bug finding and fixing is not affected.
Its thinking is always on and runs before the answer. At medium effort, the setting TulexAI runs, Artificial Analysis measured 22.17 seconds to first token on Anthropic’s API with a 10,000-token input, then 75.2 output tokens per second. Anthropic says it generates output more than 30% faster than Opus 5, which is about speed once it starts, not the wait before. None of these figures was measured through TulexAI.
Claude Opus 5.5, on the same bill as everything else
10 PT per ~750 tokens, estimated before you send, then metered on the prompt plus the reply. Claude Opus 5.5 is on the $11/month plan and up.
Start free1 free prompt · no credit card · cancel any time