It respected the frozen records.
Helmig describes Astra updating two eligible contacts while leaving frozen contacts untouched. Sol updated four, including the frozen ones. A concrete signal of constraint-following within this benchmark.
OpenAI’s numbers look impressive.
What happens when Astra
meets real work?
The meaningful shift is from generating an answer to carrying out a workflow.
Astra combines reasoning, coding and computer operation. It can work through browsers, terminals and professional software, accept steering, and continue across a multi-step job.
That is the product thesis. Whether the job finishes correctly, and what it costs to get there, is the evaluation.
Model documentation cited in the research ↗Not all evidence
answers the same question.
This report combines community experience, firsthand use and supplied background research, with reported material cut off on 4 September 2026. It is not a controlled benchmark run or a live availability check. Read the method and source differences ↗
Change the test.
Watch the story change.
A §2 · DR pp. 7, 12–13
The famous ARC-AGI-3 result
needs both numbers.
Provider-neutral interface.
The stronger common comparison.
Retained reasoning state + compaction.
A result for the model and its harness.
The adapter preserves reasoning state and compacts long conversations. DR reports comparable solved runs were roughly 3.66× faster and used 49% fewer tokens. Those figures describe overlapping runs, not every task or every Astra workload.
These two reported configurations also use different reasoning efforts. This is evidence that the full agent setup matters, not a clean experiment isolating a single setting.
The most convincing shift is what people are beginning to delegate.
Claire Vo demonstrates CRM operation and browser QA. Every reports hours inside Adobe Premiere. These are more meaningful signals than a polished launch montage, but neither gives us a reliable everyday failure rate.
OpenAI’s result, corroborated by benchmark owner Bryan Helmig in Source A.
“Better than the previous model” and “reliable enough to leave alone” are different claims.
Helmig describes Astra updating two eligible contacts while leaving frozen contacts untouched. Sol updated four, including the frozen ones. A concrete signal of constraint-following within this benchmark.
Captchas, changing interfaces, recovery loops and false completion claims remain under-documented. Successful demonstrations cannot be converted into a general real-world success percentage.
Agentic coding is a stronger case than universal intelligence.
Terminal work and token efficiency show real gains over Sol. DeepSWE is much closer. Fable 5.1 still leads Artificial Analysis’s coding composite.
Artificial Analysis · A §4 / DR pp. 12–13 ↗Astra used roughly one-third of Sol Max’s tokens in the Codex coding-agent harness. Approximately the same cost per task, despite the higher token price.
Model + agent comparisons. Harnesses and configurations matter; these results do not isolate model weights.
The price is a fact.
The value depends on the work.
Above 272K input tokens, the higher rate applies to the entire request. Tool charges are additional. Work/Codex credits use separate accounting; no standalone computer-use click fee is inferred. Pricing cited by A §5 / DR pp. 4–5 ↗
Fewer tokens offset the premium in this harness. That is an encouraging measured trade-off, not a promise for every repository.
At Max, Astra costs more for essentially the same composite score as Sol. The efficiency argument does not generalize.
OpenAI’s offline latency simulation, not live desktop timing. Latent Space’s preview ~33 tokens/second is a separate observation. Neither establishes a universal speed-up. A §5 · DR pp. 4–5
The people who used it.
The work. The limits.
Claire Vo reports getting roughly 90% of a ChatPRD feature in one run after repeated unsuccessful attempts with earlier models. Her chaptered demonstration also covers CRM, browser QA, hardware and Blender.
Watch the coding chapter · ~15:20Anecdotal comparison. Prompts, budgets and full failure counts are not controlled. A §4; DR pp. 12, 19.
Most of a first video cut reportedly produced in ~5 hours. In a journal-app comparison, the team preferred Fable’s simpler workflow over Astra’s more elaborate interface.
Probable firsthand use; DR relied on a transcript mirror because direct video fetching was throttled. Interventions undisclosed.
View the reported test ↗Shumer describes ambitious Unreal and multi-agent projects, but also plateaus, large token consumption and weaker visual taste. His custom Manager Loop helps keep long projects progressing.
Firsthand account with custom coordination, not a controlled model-only comparison.
Read the early-access account ↗The authors report a 20B+ token early-access programme across model training, pipelines, deployment and debugging. Some workloads use 20–50 simultaneous agents.
Self-reported scale across a programme, not one run. Failure rates are not fully disclosed; output-only “under $6/hour” excludes much of the total cost.
Read the engineering report ↗Mollick’s reported research run produced polished but uninteresting papers. The workflow completed; the intended quality of the research did not follow automatically.
A qualitative report about research taste, not a measured failure rate.
Read the research observation ↗Reports are paraphrased, not direct quotations. Waitlist reactions and uninspected “AGI” videos do not count as technical evidence. Sources A and DR overlap; a repeated account is not a second independent trial.
Astra is announced. The model entry exists. Initial access is limited.
Account, product and administrator settings determine access. One person’s Codex access does not establish universal availability.
OpenAI states only that availability expands over the coming days. It has not published a universal rollout completion date.
Astra reached a limited set of organisations first and expands from there to ChatGPT Plus, Pro, Business and Enterprise, the OpenAI API, Microsoft Azure and AWS Bedrock. Availability differs by account and by product while the rollout runs. Buying additional credits does not bring access forward, and enterprise administrators may need to switch Astra on: enterprise access was off by default at launch.
GPT-6 Pro, powered by Astra, is rolling out to eligible Pro $100, Pro $200, Business and Enterprise accounts. Availability remains gradual and may differ by account.
Astra is rolling out across Plus, Pro, Business and Enterprise. Plus and Business Standard receive limited Astra usage once access reaches the account. Pro $100, Pro $200 and Business Premium can use their existing full Work / Codex allowance for Astra.
GPT-6 Astra is available as gpt-6-astra, subject to API account access and rate limits.
Hours of activity are not hours of verified, correct autonomy.
Reported jobs running three to five hours. Better long-horizon analytical quality in AA-Briefcase. Strong ARC performance with preserved reasoning state.
Too few unedited trajectories, complete failure logs or disclosed interventions to estimate how often arbitrary long jobs finish correctly.
DR documents credential and safeguard boundary failures in safety simulations. Both sources report harder-to-monitor reasoning. More capable does not establish universally safe delegation.
Yes, for turning reasoning into sustained work. The evidence is much weaker for a universal intelligence leap or dependable unattended autonomy.
Both sources ultimately land on “major upgrade.” The case rests on interactive reasoning, agentic coding, token efficiency and early computer-use results. It does not require an AGI claim.
Terminal engineering, browser QA and long tool workflows where you can inspect the result.
Broad reasoning per dollar, automatic visual taste, or production work left unattended.
Controlled tasks, complete trajectories, intervention counts and cost per successful job.
A frozen research snapshot.
An explicitly provisional conclusion.
Manually reviewed X discussions, demonstrations and user reports about GPT-6 Astra, including experiences from Plus and Pro users.
Reviewed 4 September 2026. Community reactions, early hands-on impressions and differing accounts of access and usage, read and synthesized by Tenzy. This is community experience, not an independent benchmark organisation and not a reproduction of any published evaluation.
Firsthand observations from Tenzy’s own use of GPT-6 Astra, covering coding, frontend work, long-task behaviour and practical workflow.
Personal use rather than a controlled test. These observations are identified as firsthand where they appear and are not presented as measured results.
This report combines X community experience, Tenzy’s own hands-on use of Astra, supplied background research and linked original sources. Personal observations are identified separately from externally reported results. Citations marked DR refer to the supplied ChatGPT Deep Research used while preparing the report; “independent” and “verified access” describe how those sources classify their own evidence, which we have not reproduced. Links lead to original sources and may have changed since the cutoff. Nothing here is a controlled Astra benchmark.