GPT-6 Astra Explained: Price, Benchmarks, Limits
Frontier Models

GPT-6 Astra Explained: Price, Benchmarks, Limits

OpenAI released GPT-6 Astra on 3 September 2026 and called it the most intelligent and aligned model in the world. The model card and the independent benchmarks tell a more interesting story than the launch post.

By AITraining2U Editorial Team 2026-09-04 9 min read
Illuminated AI signage in a dark room — GPT-6 Astra frontier model release 2026

OpenAI named Astra publicly on 1 August 2026, inside a research post about mathematics rather than a product launch. An internal build had closed ten long-standing open problems in mathematics and theoretical computer science — among them an explicit construction of a non-sofic group and a disproof of Connes's rigidity conjecture, with the proofs formally checked in Lean 4. A month later, on 3 September, the same family shipped as GPT-6 Astra.

That gap between the research framing and the product framing is the useful thing to hold onto. Astra is a genuinely different model at the top of the difficulty curve. It is not a uniform upgrade over what you were using last month, and for a lot of Malaysian teams it will be the wrong default.

1. What actually shipped

The model card is short and worth reading in full. The numbers that matter:

  • Model ID: gpt-6-astra. One snapshot, no dated variants yet.
  • Context window: 1,050,000 tokens — 922,000 of which can be input.
  • Max output: 128,000 tokens per call.
  • Knowledge cutoff: 30 April 2026.
  • Pricing: US$10 per million input tokens, US$50 per million output. Cached input drops to US$1; cache writes cost US$12.50. Anything over 272K input tokens is billed at 2x input and 1.5x output.
  • Reasoning effort: low, medium, high, xhigh, max.
  • Tools: web search, file search, image generation, code interpreter, hosted shell, apply patch, skills, computer use, MCP, and tool search.

Rollout was staged. A limited set of organisations got it on day one; Business and Pro subscribers followed within roughly a day, Plus users hours after that, then the API and AWS. Enterprise administrators control workspace access and the model launches disabled by default — a detail worth checking if your team says "we don't have Astra yet."

2. The headline benchmarks

OpenAI's own numbers are strong: 98% on FrontierMath Tier 4, 99.9% on ARC-AGI-3, 100% on ExploitBench. On computer and browser use it scores 72.6% on OSWorld 2.0 while taking roughly 47% less time per task than GPT-5.6 Sol. Computer-use performance close to doubled, and the underlying optimisations gave existing Sol deployments something like a 60% speedup for free.

The cybersecurity result is the one that carries a policy tail. OpenAI designated Astra as crossing the "Critical" cybersecurity capability threshold under its Preparedness Framework — the highest severity level the company defines. That is not marketing. It changes what safeguards ship with the model and how enterprise access is gated.

3. Where the independent numbers disagree

Here is the part the launch coverage mostly skipped. Artificial Analysis scored Astra at 61 on its Intelligence Index — tied with GPT-5.6 Sol, and five points behind Claude Fable 5.1. On GDPval-AA v2, a benchmark adapted from OpenAI's own dataset of economically valuable tasks across 44 occupations, Astra dropped roughly 80 Elo points against its predecessor.

Token efficiency improved by about 10% at max effort. Price went up 2.5x. Net result on that index: roughly 75% more expensive per task than GPT-5.6 Sol, for no measured gain on generic knowledge work.

Both things are true at once. Astra is a step change on long-horizon agentic work, computer use, formal mathematics and security research. It is roughly flat on the kind of drafting, summarising and analysis that makes up most corporate AI usage. If your evaluation set looks like the second category, upgrading is a price increase with extra steps. We make the same argument at more length in our frontier AI models comparison.

4. The four things Astra changes in practice

What Astra actually changes

What Astra actually changes 1Computer & browser use
72.6% on OSWorld 2.0 at roughly 47% less time per task than GPT-5.6 Sol. Performance close to doubled on desktop and browser control.
2Context preservation
Searchable notes replace lossy summarisation across context windows. Requirements from hour one survive to hour eight.
3Formal mathematics
An internal build closed ten open problems, proofs verified in Lean 4 with a sorry count of zero. 98% on FrontierMath Tier 4.
4Security capability
100% on ExploitBench. First OpenAI model classified "Critical" for cybersecurity under the Preparedness Framework.

Filtering out the noise, four capability shifts are real enough to plan around.

5. What it costs on a real workload

List price is misleading for agentic work because agents burn output tokens. One complex early-access build reported by testers consumed roughly 8 million tokens — about US$17 for a single task. Multiply that across a team of ten engineers running long agent sessions daily and you are into five-figure ringgit monthly spend before anyone has shipped a feature.

Three levers keep that sane. Prompt caching cuts repeated input from US$10 to US$1 per million, which matters enormously for agents that re-read the same codebase. Batch processing runs at half price for anything that does not need a live answer. And reasoning effort is a dial, not a switch — most tasks do not need max, and the gap between medium and max on routine work is small relative to the token bill.

The discipline that saves money is boring: measure cost per solved task, not cost per token. A cheaper model that fails 30% of the time is not cheaper.

6. Who should actually switch

Our read for Malaysian teams, stated plainly so you can argue with it:

  • Switch now if you run long-horizon coding agents, browser automation, or security research. The computer-use and Codex improvements are the strongest part of this release, and the context-preservation change — searchable notes instead of lossy summarisation across context windows — is the first fix we have seen that meaningfully reduces agent drift on multi-hour tasks.
  • Stay on GPT-5.6 Terra or Luna if your workload is document drafting, meeting summaries, customer replies, or retrieval-augmented Q&A. Terra sits at US$2/US$12 per million tokens. Astra is five times the input price for work it does not measurably do better.
  • Run your own evaluation before deciding either way. Public benchmarks are proxies for someone else's job, not yours. We cover how to build a defensible internal eval set in our piece on LLM evaluation in production.

7. The governance question nobody is asking

A model classified "Critical" for cybersecurity capability, with hosted shell access, computer use and MCP tool calling enabled by default in the API, is a different governance object than a chat assistant. If you are in a BNM RMiT-regulated environment, or handling PDPA-covered personal data, the questions your risk committee should be asking changed on 3 September — specifically around what the model is permitted to execute, not just what it is permitted to read. Our AI governance guide for Malaysian organisations covers the control set; the short version is that agentic tool access needs its own approval path, separate from your existing LLM policy.

💡
HRDC SBL-KHAS claimable

AITraining2U runs HRD Corp-registered programmes on frontier model selection, evaluation and agent deployment — see AI Engineering and AI Automation, or read how to claim HRDC funding for the step-by-step.

Astra is the most capable model OpenAI has shipped and simultaneously the easiest one to waste money on. Both statements come from the same source: it is optimised for the hardest 5% of tasks, and most organisations spend their token budget on the other 95%.

About the author

AITraining2U Editorial Team →

HRDC-Certified · Practitioner-Led · Malaysia & SEA

The AITraining2U Editorial Team is a working group of practitioners — instructors, working consultants, and HRDC-certified trainers — who collectively deliver AI training to Malaysian organisations across financial services, technology, professional services, and the public sector. Articles attributed to the Editorial Team draw on consolidated learnings from live programmes, corporate engagements, and regional industry research.

Frequently Asked Questions

OpenAI released GPT-6 Astra on 3 September 2026, to a limited set of organisations first. Access expanded over the following days: Business and Pro subscribers within about a day, ChatGPT Plus users hours later, then the OpenAI API and AWS. The name had already appeared publicly on 1 August 2026 in an OpenAI research post about an internal Astra build solving ten open problems in mathematics and theoretical computer science.

US$10 per million input tokens and US$50 per million output tokens on the standard API tier. Cached input drops to US$1 per million and cache writes cost US$12.50. Batch processing runs at half price; a Fast mode is available at 2x the standard rate. Requests exceeding 272,000 input tokens are billed at 2x input and 1.5x output. That is roughly 2.5x the price of GPT-5.6 Sol.

It depends entirely on the task. Astra is clearly ahead on computer use (72.6% on OSWorld 2.0, at about 47% less time per task), long-horizon coding, formal mathematics and cybersecurity. But Artificial Analysis scored it at 61 on their Intelligence Index — tied with GPT-5.6 Sol — and measured roughly an 80 Elo drop on GDPval-AA v2, a benchmark of economically valuable knowledge work. For ordinary drafting and analysis, Astra costs more without measurably performing better.

1,050,000 tokens total, of which up to 922,000 can be input and 128,000 can be output in a single call. Knowledge cutoff is 30 April 2026. Note the pricing cliff: prompts above 272,000 input tokens are charged at double the input rate and 1.5x the output rate, so a genuinely million-token prompt costs considerably more than the headline figure suggests.

Only if your workload is agentic. Teams running long-horizon coding agents, browser automation, or security research will see real gains from the computer-use improvements and the new context-preservation behaviour in Codex. Teams doing document drafting, summarisation, customer replies or RAG-based Q&A should stay on GPT-5.6 Terra (US$2/US$12 per million) and spend the difference on evaluation and training instead. Run your own eval set before deciding.

Want to apply this in your organisation?

AITraining2U runs HRDC-claimable corporate AI training for Malaysian organisations — from leadership awareness to hands-on builder workshops. Talk to us about a programme tailored to your team.