Tenzy. REPORTS Independent evaluation / GPT-6 The evidence
Beyond the launch numbersResearch snapshot · 04 SEP 2026

How good is
Astra, actually?

The question behind GPT-6

OpenAI’s numbers look impressive.
What happens when Astra
meets real work?

Explore the investigation
An independent investigation. Not an OpenAI publication.Evidence first. Conclusions second.
The model, in perspectiveDocumented specs · A §1 / DR pp. 3–4

A model built
to stay on the job.

The meaningful shift is from generating an answer to carrying out a workflow.

Astra combines reasoning, coding and computer operation. It can work through browsers, terminals and professional software, accept steering, and continue across a multi-step job.

That is the product thesis. Whether the job finishes correctly, and what it costs to get there, is the evaluation.

Model documentation cited in the research ↗
Context window
1.05M
tokens
Maximum output
128K
tokens
Model interface
Text + image in.
Text out.
No native audio or video
Reasoning + execution
Think. Act.
Check. Continue.
Tools, steering and context management

Not all evidence
answers the same question.

OfficialWhat the provider reports.
IndependentWhat evaluators measured.
FirsthandWhat early users observed.
Our readingWhat the evidence justifies.

This report combines community experience, firsthand use and supplied background research, with reported material cut off on 4 September 2026. It is not a controlled benchmark run or a live availability check. Read the method and source differences ↗

Reading the numbersSelected comparisons · not an overall ranking

Impressive.
Uneven. Both can be true.

Change the test.
Watch the story change.

Coding & terminal work

Terminal-Bench 4.0

Score (%) · higher is better
GPT-6 Astra57.9%
Claude Fable 5.155.8%
GPT-5.6 Sol37.3%

A §2 · DR pp. 7, 12–13

The most important asteriskIndependent · ARC Prize

Same benchmark.
Different machinery.

The famous ARC-AGI-3 result
needs both numbers.

Standard harness · Max62.7%

Provider-neutral interface.
The stronger common comparison.

Provider Adapter · High99.9%

Retained reasoning state + compaction.
A result for the model and its harness.

What changed besides the score?

The adapter preserves reasoning state and compacts long conversations. DR reports comparable solved runs were roughly 3.66× faster and used 49% fewer tokens. Those figures describe overlapping runs, not every task or every Astra workload.

These two reported configurations also use different reasoning efforts. This is evidence that the full agent setup matters, not a clean experiment isolating a single setting.

When the model takes the controlsA §3 · DR pp. 10–12

Computer use is
the real story.

The most convincing shift is what people are beginning to delegate.

Claire Vo demonstrates CRM operation and browser QA. Every reports hours inside Adobe Premiere. These are more meaningful signals than a polished launch montage, but neither gives us a reliable everyday failure rate.

AutomationBench · simulated apps41.4%

A higher ceiling.
Still a low full-pass rate.

OpenAI’s result, corroborated by benchmark owner Bryan Helmig in Source A.

41.4% fully pass58.6% do not fully pass

“Better than the previous model” and “reliable enough to leave alone” are different claims.

Sol 18.1%Fable 5.1 31.4%Astra 41.4%
Read the benchmark owner’s report ↗
The encouraging detail

It respected the frozen records.

Helmig describes Astra updating two eligible contacts while leaving frozen contacts untouched. Sol updated four, including the frozen ones. A concrete signal of constraint-following within this benchmark.

The missing evidence

Messy, everyday autonomy.

Captchas, changing interfaces, recovery loops and false completion claims remain under-documented. Successful demonstrations cannot be converted into a general real-world success percentage.

Software engineeringIndependent + firsthand

Stronger at
the work.
Not every test.

Agentic coding is a stronger case than universal intelligence.

Terminal work and token efficiency show real gains over Sol. DeepSWE is much closer. Fable 5.1 still leads Artificial Analysis’s coding composite.

Artificial Analysis · A §4 / DR pp. 12–13 ↗
AA Coding Agent Index · points
GPT-6 Astra67
Claude Fable 5.170
~⅓

The tokens, in this test.

Astra used roughly one-third of Sol Max’s tokens in the Codex coding-agent harness. Approximately the same cost per task, despite the higher token price.

Model + agent comparisons. Harnesses and configurations matter; these results do not isolate model weights.

Speed & economicsPrices frozen at the research cutoff

More per token.
What about per task?

The price is a fact.
The value depends on the work.

Astra API rates

Price the configuration.

Input
$10
Cached input
$1
Cache write
$12.50
Output
$50
USD per 1 million tokens · API onlyStandard short-context rates are 2.5× Sol’s.

Above 272K input tokens, the higher rate applies to the entire request. Tool charges are additional. Work/Codex credits use separate accounting; no standalone computer-use click fee is inferred. Pricing cited by A §5 / DR pp. 4–5 ↗

Independent · AA Coding Agent Index≈ same

Cost per coding task vs Sol Max.

Fewer tokens offset the premium in this harness. That is an encouraging measured trade-off, not a promise for every repository.

Independent · AA Intelligence Index~75% more

Cost per broad intelligence task.

At Max, Astra costs more for essentially the same composite score as Sol. The efficiency argument does not generalize.

Speed needs a denominator

~75 min ~40 min

Sol → Astra · OSWorld task time

OpenAI’s offline latency simulation, not live desktop timing. Latent Space’s preview ~33 tokens/second is a separate observation. Neither establishes a universal speed-up. A §5 · DR pp. 4–5

Outside the launch deckEarly access, small sample, useful signals

Field notes.
Not a consensus.

The people who used it.
The work. The limits.

Every / Dan Shipper · DR pp. 11, 20–21Five hours in Premiere. A UX miss elsewhere.

Most of a first video cut reportedly produced in ~5 hours. In a journal-app comparison, the team preferred Fable’s simpler workflow over Astra’s more elaborate interface.

Probable firsthand use; DR relied on a transcript mirror because direct video fetching was throttled. Interventions undisclosed.

View the reported test ↗
Matt Shumer · A §7More ambition. Still needs a manager.

Shumer describes ambitious Unreal and multi-agent projects, but also plateaus, large token consumption and weaker visual taste. His custom Manager Loop helps keep long projects progressing.

Firsthand account with custom coordination, not a controlled model-only comparison.

Read the early-access account ↗
Latent Space · A §7 / DR p. 1320B+ tokens. Substantial scaffolding.

The authors report a 20B+ token early-access programme across model training, pipelines, deployment and debugging. Some workloads use 20–50 simultaneous agents.

Self-reported scale across a programme, not one run. Failure rates are not fully disclosed; output-only “under $6/hour” excludes much of the total cost.

Read the engineering report ↗
Ethan Mollick · Source A, long-task auditFinishing is not the same as succeeding.

Mollick’s reported research run produced polished but uninteresting papers. The workflow completed; the intended quality of the research did not follow automatically.

A qualitative report about research taste, not a measured failure rate.

Read the research observation ↗

Reports are paraphrased, not direct quotations. Waitlist reactions and uninspected “AGI” videos do not count as technical evidence. Sources A and DR overlap; a repeated account is not a second independent trial.

Release & availabilityResearch snapshot 04 SEP 2026 · availability reverified 05 SEP 2026

Announced
doesn’t mean
available to you.

03 SEP 2026

Announcement + API changelog.

Astra is announced. The model entry exists. Initial access is limited.

04 SEP 2026 · RESEARCH CUTOFF

Rollout is still incomplete.

Account, product and administrator settings determine access. One person’s Codex access does not establish universal availability.

COMING DAYS · NO VERIFIED UNIVERSAL DATE

Broader access is rolling out gradually.

OpenAI states only that availability expands over the coming days. It has not published a universal rollout completion date.

Who gets access?

Astra reached a limited set of organisations first and expands from there to ChatGPT Plus, Pro, Business and Enterprise, the OpenAI API, Microsoft Azure and AWS Bedrock. Availability differs by account and by product while the rollout runs. Buying additional credits does not bring access forward, and enterprise administrators may need to switch Astra on: enterprise access was off by default at launch.

Chat

GPT-6 Pro, powered by Astra, is rolling out to eligible Pro $100, Pro $200, Business and Enterprise accounts. Availability remains gradual and may differ by account.

Work & Codex

Astra is rolling out across Plus, Pro, Business and Enterprise. Plus and Business Standard receive limited Astra usage once access reaches the account. Pro $100, Pro $200 and Business Premium can use their existing full Work / Codex allowance for Astra.

API

GPT-6 Astra is available as gpt-6-astra, subject to API account access and rate limits.

Product availability documentation ↗
The unresolved testOur reading · A §7 / DR pp. 13–14, 21

Can it keep going?
Yes.
Can you stop watching?

Arbitrary long-task success rateUnknown.

Hours of activity are not hours of verified, correct autonomy.

What we have

Real signs of persistence.

Reported jobs running three to five hours. Better long-horizon analytical quality in AA-Briefcase. Strong ARC performance with preserved reasoning state.

What we don’t have

A reliable denominator.

Too few unedited trajectories, complete failure logs or disclosed interventions to estimate how often arbitrary long jobs finish correctly.

What should temper trust

Scope and supervision still matter.

DR documents credential and safeguard boundary failures in safety simulations. Both sources report harder-to-monitor reasoning. More capable does not establish universally safe delegation.

System-card evidence cited by DR ↗
So, is Astra actually good?Synthesis · provisional · medium confidence

A major upgrade.
An unfinished story.

Yes, for turning reasoning into sustained work. The evidence is much weaker for a universal intelligence leap or dependable unattended autonomy.

Both sources ultimately land on “major upgrade.” The case rests on interactive reasoning, agentic coding, token efficiency and early computer-use results. It does not require an AGI claim.

The strongest case

Hard, supervised agent work.

Terminal engineering, browser QA and long tool workflows where you can inspect the result.

The weaker case

A blanket replacement.

Broad reasoning per dollar, automatic visual taste, or production work left unattended.

What would change the verdict

Repeated real-world proof.

Controlled tasks, complete trajectories, intervention counts and cost per successful job.

Follow the evidence
The reading roomTransparent by design

Research sources.

A frozen research snapshot.
An explicitly provisional conclusion.

Source A · X community research

Community and user experience.

Manually reviewed X discussions, demonstrations and user reports about GPT-6 Astra, including experiences from Plus and Pro users.

Reviewed 4 September 2026. Community reactions, early hands-on impressions and differing accounts of access and usage, read and synthesized by Tenzy. This is community experience, not an independent benchmark organisation and not a reproduction of any published evaluation.

Source B · Personal hands-on

Tenzy’s hands-on experience with Astra.

Firsthand observations from Tenzy’s own use of GPT-6 Astra, covering coding, frontend work, long-task behaviour and practical workflow.

Personal use rather than a controlled test. These observations are identified as firsthand where they appear and are not presented as measured results.

This report combines X community experience, Tenzy’s own hands-on use of Astra, supplied background research and linked original sources. Personal observations are identified separately from externally reported results. Citations marked DR refer to the supplied ChatGPT Deep Research used while preparing the report; “independent” and “verified access” describe how those sources classify their own evidence, which we have not reproduced. Links lead to original sources and may have changed since the cutoff. Nothing here is a controlled Astra benchmark.

Editorial decisions & source disagreements
  • Availability: product-specific OpenAI documentation takes precedence over broad launch wording. No universal rollout completion date has been published, so none is stated here.
  • Terminal-Bench: retain 57.9% for the launch comparison and separately identify DR’s 58.2% ± 2.8 maintainer result.
  • Humanity’s Last Exam: A’s summary says independent; its detailed table and DR say OpenAI-reported. Use the conservative attribution.
  • Computer use: A calls it “Good, trending Strong”; DR calls it “Strong.” Both identify a narrow sample. We emphasize that shared limit instead of inventing a consensus rating.
  • Subjective scores: A gives lower usefulness and coding ratings than DR, and assigns long-task scores where DR says insufficient evidence. We do not average subjective scores or turn them into measurements.
  • Tool fees: A mentions a generic call fee; DR identifies specific search fees. No standalone computer-use charge is established here.
  • Precision: AA’s broad index uses A’s precise table values; DR uses rounded approximations. The difference is disclosed beside the chart.