GPT-6 Astra: frontier reasoning for plans, logs and fact-checks, with the tools left behind
GPT-6 Astra reached general availability on September 4, 2026, and OpenAI calls it "the most intelligent and aligned model in the world". Much of what OpenAI claims for it, from computer use to browsing, runs through tools that only its Responses API offers, and TulexAI does not pass tools to Astra. This page is about the jobs that still hold up in plain chat.
- Cost
- 25 PT
- Per
- ~750 tokens
- Tier
- premium
- From
- $23/mo
Astra is OpenAI’s newest flagship. Measured at max effort by Artificial Analysis, it ties Claude Fable 5.1 at 53 on Intelligence Index v4.3. On TulexAI it runs as chat only: no tools, replies capped at 8,192 tokens. Use it for plans, automation designs, pasted logs and fact-checks against a source.
What OpenAI GPT-6 Astra is best at
- Long-horizon plans you will carry out yourself: a migration in phases, a roadmap with dependencies, a launch checklist with stop conditions. On AA-Briefcase, the long-horizon work eval, it sits about 90 Elo above GPT-5.6 Sol, measured at max effort by Artificial Analysis.
- Designing an automation before anyone builds it: the trigger, each step, failure handling and the human checkpoints, written out as a spec. It scored 68% on AutomationBench-AA to Claude Fable 5.1’s 59%, measured at max effort by Artificial Analysis. On TulexAI it writes the design; it does not run it.
- Reasoning over terminal output and ops logs you paste in: a failed deploy, a crash loop, a permissions error three commands deep. It led Fable 5.1 on Terminal-Bench v4.0, 59% to 52% measured at max effort by Artificial Analysis, though here it reads the output rather than running the commands.
- Checking claims against a source you paste, where an invented detail is expensive. Artificial Analysis measured its hallucination rate at 51% against GPT-5.6 Sol’s 92%, both at max effort. Lower is not zero, so the prompt should tie every answer to the text.
- Screenshots and diagrams: it takes image input, so an error dialog, a dashboard or an architecture sketch goes straight into the prompt instead of being described in words.
Where it falls down
- No tools on TulexAI. We call OpenAI’s Chat Completions API, and OpenAI’s docs are plain: "Chat Completions does not support function calling with GPT-6 Astra." A request that carries tools is refused with a clear error, so function calling, web search, file search, code interpreter and computer use are all unavailable here.
- Replies are capped at 8,192 output tokens on TulexAI. OpenAI allows up to 128,000, but the platform clamp applies, so a report-length answer has to be requested in parts.
- It is slow to start. Artificial Analysis measured 37 seconds to first token at high effort and 334 seconds at max, thinking included. TulexAI runs OpenAI’s default effort, which OpenAI does not document, so neither figure is a measurement of TulexAI.
- Its knowledge stops on April 30, 2026. Anything newer has to be pasted into the prompt, because with no browsing it cannot go and fetch it.
- You cannot inspect its reasoning. It uses a "recurrent depth" technique that obscures some or all of how it reaches an answer, so a written explanation is something to verify, not a trace of what happened.
- It draws on the frontier text allowance shared with Claude Fable 5 and Claude Fable 5.1, so a heavy week on Astra leaves less of that allowance for Fable.
- The benchmark scores on this page were measured at max or high effort by Artificial Analysis. TulexAI runs Astra at OpenAI’s default effort, so the scores describe the model at those settings, not what a TulexAI reply will score.
- It is not ahead everywhere. On Intelligence Index v4.3, Claude Fable 5.1 scored higher on SciCode (63% to 56%), Humanity’s Last Exam (59% to 55%), AA-Briefcase and GDPval-AA v2.
Prompts for OpenAI GPT-6 Astra
Copy one, swap the subject, keep the structure - the constraint clauses are what make these land.
Turn a vague project into a phased plan with stop conditions
OpenAI GPT-6 AstraWe are moving a 40-table Postgres database from a self-managed server to a managed service with no more than 15 minutes of downtime. Write the plan as numbered phases. For each phase give: the goal, what must already be true before it starts, the exact steps, how we verify it worked, and the condition under which we stop and roll back. Call out every dependency between phases. End with the three decisions I must make before phase one, and do not write any code.
Forcing a precondition, a check and a rollback trigger into every phase turns a long plan into something you can execute and abort step by step. Banning code keeps the 8,192-token reply spent on the plan itself rather than on scripts you would rewrite anyway.
Design an automation as a spec someone else will build
OpenAI GPT-6 AstraDesign an automation for supplier invoices that arrive as PDF email attachments and must end up approved in our accounting system. Do not assume any specific tool. Write it as a spec: the trigger, each step with its input and output, what happens when a step fails, where a human must approve, and what gets logged. Then list every point where an invoice could be paid twice, and the step that prevents it.
Asking for failure handling and a named double-payment check pulls out the part automation designs usually skip. Astra cannot run tools here, so a tool-neutral written spec is the useful deliverable, and it hands straight to whoever builds the workflow.
Diagnose a failure from pasted terminal output
OpenAI GPT-6 AstraThe output below is from a deploy that fails at the container start step. Do not suggest a fix yet. First list the three most likely causes, ranked, and quote the log lines that support each one. For each cause, give one read-only command I can run to confirm or rule it out, and say what output would confirm it. Only after that, describe the smallest fix for the top cause. The log follows this line.
Astra cannot execute commands on TulexAI, so the prompt makes you its hands: read-only checks with the expected result spelled out. Quoting log lines per hypothesis stops a confident diagnosis that the log does not actually support.
Fact-check a draft against its source, not against memory
OpenAI GPT-6 AstraBelow are a source document and a draft article based on it. For every factual claim in the draft, label it supported, contradicted, or not in source, and quote the sentence from the source that decides it. Treat anything the source does not say as not in source, even if you believe it is true. Do not rewrite the draft. The source comes first, then the draft.
Astra measured a far lower hallucination rate than GPT-5.6 Sol at max effort, but 51% is no licence to trust its memory. Restricting it to the pasted source and demanding a deciding quote makes every verdict checkable in seconds.
Check an architecture diagram against how the docs describe it
OpenAI GPT-6 AstraThe attached image is our architecture diagram, and the paragraph below is how our documentation describes the same system. List every component or connection that appears in one but not the other, and every arrow whose direction contradicts the text. Say where each item sits in the diagram so I can find it. Do not redraw or redesign anything.
Image input lets the diagram be the evidence rather than your summary of it. Asking where each mismatch sits in the picture gives you a way to check every finding, which matters on a model whose reasoning you cannot inspect.
Craft notes
- 01Draft the prompt on a faster model and send only the finished version to Astra. With a long wait before the first token, iterating on wording here is the slow way to work.
- 02Ask for long deliverables in parts. Replies stop at 8,192 output tokens, so request the first three phases, then continue in the same chat.
- 03Paste the source for anything after April 30, 2026. It has no browsing on TulexAI and nothing newer in its training data.
- 04Keep input under roughly 270,000 tokens. Past 272,000, OpenAI reprices the whole request from $10 input and $50 output per million tokens to $20 and $75, even though the context window runs to 1,050,000.
- 05Ask for evidence, not reasoning. Its reasoning is obscured, so quotes, log lines and commands you can run are the parts of an answer you can actually verify.
- 06Leave tool definitions out of API calls. A request that includes tools is refused rather than quietly ignored, so strip them or send that call to Claude Fable 5.1.
- 07Spend the frontier allowance deliberately. Astra, Fable 5 and Fable 5.1 draw on the same pool, so routine drafting belongs on a model outside it.
When to use something else
OpenAI GPT-6 Astra is one of 30 text models on the same Platform Token budget, so switching costs nothing but the tokens.
You need tool use or an agent loop through TulexAI, or the job looks like AA-Briefcase or GDPval-AA v2 agentic knowledge work, where Fable 5.1 scored higher than Astra on Intelligence Index v4.3. It costs the same 25 PT per 750 tokens.
You need function calling on an OpenAI model through TulexAI, or the task does not need Astra’s extra reasoning and should not draw on the frontier allowance. Artificial Analysis also measured Sol about 45 Elo ahead of Astra on GDPval-AA v2.
You need tool calls for agentic coding or knowledge work at a lower rate: Opus 5.5 supports tools on TulexAI, though a forced tool choice runs as auto, and costs 10 PT per 750 tokens against Astra’s 25, on every plan from Basic.
The work is keeping a long document or a codebase consistent and you want tools alongside it: Opus 5 supports tools on TulexAI, costs 12 PT per 750 tokens against Astra’s 25, and is on every plan from Basic.
Frequently asked questions
OpenAI’s newest flagship, generally available since September 4, 2026, after a limited preview the day before. It takes text and image input, replies in text, has a 1,050,000-token context window and a knowledge cutoff of April 30, 2026. OpenAI positions it for computer use, browsing, software engineering, cybersecurity, science and professional work.
On the Artificial Analysis Intelligence Index v4.3, published September 7, 2026, it scored 53 at max effort (rank 3 of 200) and 51 at high effort (rank 6 of 200). Claude Fable 5.1 scored 53 at max effort with fallback, and GPT-5.6 Sol 47 at max. TulexAI runs OpenAI’s default effort, so these are not scores for a TulexAI reply.
No. TulexAI reaches Astra through Chat Completions, which OpenAI says does not support function calling with this model, so any request carrying tools is refused with a clear error. For tool use or agent loops, Claude Fable 5.1 and GPT-5.6 Sol both support tools here.
Pro at $23 a month, VIP and Elite; it is not on Basic. It costs 25 Platform Tokens for every 750 tokens of prompt plus reply, reasoning included, with a 25 PT minimum per request, so a long answer costs a multiple of that. It is metered in the frontier text allowance it shares with Claude Fable 5 and Claude Fable 5.1.
Measured at max effort by Artificial Analysis, they tie at 53 on Intelligence Index v4.3 and at 62 on the Coding Agent Index. Astra led on AutomationBench-AA and Terminal-Bench v4.0; Fable 5.1 led on SciCode, Humanity’s Last Exam, AA-Briefcase and GDPval-AA v2. Astra used about 27k output tokens per task against Fable 5.1’s 78k. Both cost 25 PT per 750 tokens here, but of the two only Fable 5.1 supports tools on TulexAI.
It thinks before it streams. Artificial Analysis measured 37 seconds to first token at high effort and 334 seconds at max, thinking included, with output at 52.4 and 58.1 tokens per second respectively once it starts. OpenAI does not document the default effort TulexAI uses, so expect a wait rather than a specific number.
OpenAI GPT-6 Astra, on the same bill as everything else
25 PT per ~750 tokens, estimated before you send, then metered on the prompt plus the reply. OpenAI GPT-6 Astra is on the $23/month plan and up.
Start free1 free prompt · no credit card · cancel any time