GPT-6 Astra Got Better at Doing. It Got Worse at Writing.

Sam Altman recently announced the debut of OpenAI’s newest frontier model for ChatGPT which is running under the name “Astra”. Historically each subsequent model is heralded as a tremendous, game-changing, step forward for LLM style AI technology. Also historically, the models tend to oversell and underdeliver in terms of how reality meets the hype. We’ve reached the iPhone era of AI where each new release upgrades or improves a number of features and functions but in terms of the “WOW!” stuff that consumers tend to look for when a new release drops, there’s little to be found. This is the incremental nature of technological progression, it’s a bit of a grind which doesn’t exactly fit the “revolutionary” narrative that AI bulls have been pushing for the past few years.

It also doesn’t help that OpenAI decided to usher in Astra with fantastical proclamations that the “era of artificial general intelligence” had arrived. That’s a VERY loaded statement as AGI carries with it a host of implications and expectations that no model to date has even come close to achieving. Jensen Huang of Nvidia fame would later double down on this statement by saying that the release of Astra meant that “AGI has arrived”.

Suffice to say I’m skeptical if Astra will buck the trend of making big promises and delivering more incremental progress. So I spent about 48 hours over this past weekend experimenting with the new model and examining its benchmark data to see just how it stacked up. If your IT team and ops folks have been stressing about what this new model means for your company’s usage and adoption plans fear not, the delta between what you’re doing today and what Astra can do is smaller than you think. Here’s how the data shakes out:

Two numbers came out of the first week of GPT-6 Astra. They point in opposite directions, which is probably why most people only picked one. On OSWorld 2, which is a benchmark for operating a desktop computer — actually opening things, clicking things, finishing a task — Astra completed 72.6% of tasks against 65.7% for the model it replaced. And it did it in roughly 40 minutes per task instead of 75. That’s not incremental. That’s a different kind of worker showing up.

On GDPval-AA v2, Artificial Analysis’s benchmark for economically valuable knowledge work across dozens of occupations, Astra scored roughly 80 Elo points below its predecessor.

Better at operating software. Worse at the work itself.

Most of the coverage that week chose a lane. The launch was either an AGI moment or a disappointment, depending on which number the writer saw first. Both readings are convenient. Both miss what the pair of numbers says together.

The frontier moved in execution and stalled in judgment, and the price of that execution went up two and a half times.

That combination should change where you point this technology inside your business. And it is not where most companies are pointing it right now.

What actually shipped

The details matter less than the pattern, but the details are messy enough that you should know them.

OpenAI released Astra on 3 September 2026, first to a vetted group inside its Daybreak security program, then the following day to paid tiers in a restricted form. A week later the rollout still wasn’t clean. Access landed in some surfaces and not others, OpenAI’s own launch page and help documentation disagreed about who could actually use it, and Sam Altman publicly apologized for what he called a messy rollout. Enterprise access ships off by default and requires an administrator to switch it on.

What you get is the model plus your account tier, plus the interface you reach it through, plus whichever safeguards apply to you. Any result someone reported in week one may describe a configuration you cannot buy.

The clearest example is the headline benchmark. OpenAI’s launch table gave Astra 99.9% on ARC-AGI-3 next to 30.2% for Claude Opus 5 and 7.8% for its own previous model. A footnote disclosed that Astra alone ran on OpenAI’s Responses API harness, which changes how context is managed between requests. ARC Prize published both configurations: 99.9% under that provider adapter, and 62.7% under ARC Prize’s own standard harness. Same model weights. A 37-point spread.

ARC-AGI-3 · one model, two harnesses

99.9%
OpenAI Responses API harness — the provider adapter, supplied by the vendor
62.7%
ARC Prize standard harness — the common runner every model is measured on

Both configurations published by ARC Prize; both are records in its own report.

Neither number is dishonest. OpenAI disclosed the caveat, ARC Prize published both, and both are records in ARC Prize’s own report. But “saturates the benchmark” is a claim about a model-plus-harness configuration, and the harness came from the vendor. When you evaluate this thing, the demo’s numbers are not your numbers, and the gap between them is not small.

Where it improved

Operating a computer. The OSWorld 2 result is OpenAI’s own measurement, so discount accordingly, but the direction is corroborated by independent developer reports. OpenAI also changed its Codex harness alongside the model, so the product-level improvement cannot be cleanly assigned to the model weights alone.

Making things in software that isn’t a text editor. This is the most distinctive claim in the release. Demonstrations show Astra modelling a house in Blender and converting it into a walkable Unreal Engine scene. Developers report direct manipulation of Unreal, Unity, Blender and Godot. Whatever you think of the demos, the category is new: this is a model operating professional creative tooling rather than describing how to operate it.

Producing artefacts that match a template. OpenAI positions Astra as its best model at following an existing template and generating slides, documents and spreadsheets in a house style, pulling only relevant context rather than padding to length. If that holds under independent testing, it is the claim with the most money attached to it.

Not making things up. On Artificial Analysis’s AA-Omniscience benchmark, Astra’s hallucination rate came in at 51% against 92% for its predecessor. That is a third-party measurement, not a vendor one, and within that benchmark it is a large improvement. It is not a general hallucination rate and should not be quoted as one.

Knowing when to ask. In OpenAI’s own side-by-side, the previous model built a career website autonomously in thirteen minutes. Astra stopped after twenty seconds to ask what career the user was moving into. Asking cost twenty seconds and saved a rebuild.

Every item on that list is a form of doing. Operate the tool. Follow the template. Run the long job. Ask before assuming.

Where it got worse

The regression is in the other half of the work, and it reproduced in more than one place.

Artificial Analysis measured the roughly 80-Elo drop on GDPval-AA v2. An independent hands-on review at TechFlowPost separately scored Astra’s editorial output at 1,995 on its own scale against 2,156 for the previous model, and reached the same conclusion in different words: the model falls flat precisely where the standard relies on taste. A tester asked it to write four paragraphs in his own voice, well enough to get past an AI-detection tool. It failed.

On broad general intelligence, Astra is flat. Artificial Analysis’s composite index puts it at roughly 61, effectively tied with the model it replaced and behind the leading Claude configuration. On coding it lands around 67, which is a real gain over its predecessor but a three-way tie with Anthropic’s current models rather than a lead.

The evidence is not one-directional and I want to be fair to it. Cognition reported that Astra’s writing made their agent’s test reports clearer. A staff writer at Every had Astra draft the first version of her own review, and the outlet’s CEO read it without noticing she had not written it. Structured, functional, workmanlike prose appears to have improved. Prose with a person in it did not.

Now the price

$10 per million input tokens. $50 per million output tokens. Fast mode at twice that. Roughly two and a half times the previous generation’s rate. On the consumer side, message allowances in ChatGPT run 5 to 45 per five-hour window against 10 to 100 for the older model.

Token efficiency is not cost efficiency. Astra uses meaningfully fewer tokens on agentic coding work, but Artificial Analysis found that a roughly 10% token reduction did not come close to offsetting a 2.5x price increase, and measured cost per task on its intelligence index at about 75% higher.

CodeRabbit, a code-review tooling vendor, ran its own evaluation. Astra caught about 4% more labelled bugs than the previous model overall, with a larger 20 to 33% improvement on cross-file review, at roughly $1.50 per task against $0.60. That is vendor-reported, built on a fixed token assumption, and CodeRabbit itself flags it as directional.

What 2.5× buys — cost per code-review task

$0.60
GPT-5.6 Sol
$1.50
GPT-6 Astra
+4%more labelled bugs caught, overall
+20–33%on cross-file review specifically

CodeRabbit, vendor-reported on a fixed 100K-in / 10K-out token assumption.

The shape is this: two and a half times the cost for a few percent more of the general thing, and a third more of one specific thing. Whether that is a good trade depends entirely on whether the specific thing is what you actually needed.

That is the procurement question most teams are skipping. The default enterprise move on a new frontier model is to upgrade the default. On these numbers, upgrading the default is the expensive way to buy a capability most of your workload will not use.

The pattern

This is where I think we need to pause, because the pattern here is more useful than any single benchmark.

Execution improved, measurably and across several kinds of execution. General reasoning stayed flat. Judgment-adjacent knowledge work and personal register got worse. Price went up 2.5x.

Where the frontier actually moved

Desktop tasks
OSWorld 2

72.6%
from 65.7% — and in 40 min per task, not 75

Hallucination
AA-Omniscience

51%
from 92% — large improvement, within this benchmark

Knowledge work
GDPval-AA v2

−80
Elo, against its own predecessor

General intelligence
AA Intelligence Index

≈61
flat, and behind the leading Claude configuration
Upward marks are gains against the previous model; downward marks are regressions.2.5× the price

OSWorld 2 is OpenAI’s own measurement. AA-Omniscience, GDPval-AA v2 and the Intelligence Index are third-party, from Artificial Analysis.

Put more plainly: this release is a strong argument for AI as an execution layer and a weak argument for AI as a voice.

That is what the measurements say, and it holds regardless of what the next model does, because it describes where this generation of capability actually landed rather than predicting where the next one will.

For most businesses the practical consequence is uncomfortable, because the gains are concentrated in the part of the work that was never the expensive part, and the regression sits in the part that was.

Go back to the template-adherence claim. In most content operations, drafting is not the cost center. Conforming output to a house standard is. Reformatting a markdown dump into the corporate deck template, matching the brand voice, getting the tables right, making it look like it came from your company rather than from a language model — that is the standing tax, and it is paid in senior people’s hours. A model that reduces that tax is worth real money, and it is worth more than a model that drafts faster.

So the first test is whether it can take something already written and put it into your format, at your standard, without a human rebuilding it. That is narrow, it is high-value, and it is answerable in an afternoon.

Three things that follow

If this were a book about success, this is the part where I’d give you a clean playbook. It isn’t. This is the part where the future is undefined, so we have to make choices without knowing the ending. Here are three.

1. Stop upgrading the default and start routing by task type. The economics only work where the execution gain is the thing you need: agentic work, computer operation, long multi-step jobs, template production, anything where a lower hallucination rate materially reduces review time. For general drafting, summarizing, and correspondence, you are paying 2.5x for a model that measures flat on general intelligence and worse on knowledge work. Routing costs engineering effort. On these numbers it pays for itself.

2. Test on your own workload. The composite index is 34% agents, 24% coding, 24% scientific reasoning, 18% general capability, which describes almost nobody’s job. Artificial Analysis’s own documentation is inconsistent about that weighting, several components are graded by other language models, and practitioners argue about whether the rankings match experience. Use the component evaluations closest to your work and ignore the headline. And run whatever you test on the harness and account tier you will actually deploy on, because the 37-point ARC-AGI-3 spread is what happens when you don’t.

3. Keep humans on voice, and stop apologizing for it. This is the finding I would take to a board. The best-resourced model release of the year moved forward on execution and backward on writing in a personal register, and it did so while the frontier lab in question was calling it a new capability level. If you have been arguing that your brand voice, your point of view, and your senior judgment are not automatable line items, you now have a benchmark to point at instead of an opinion to defend. Use the numbers rather than the sentiment.

Autonomy up, legibility down

Astra is the first OpenAI model to cross the Critical threshold for cybersecurity under the company’s own Preparedness Framework. OpenAI states it can devise and execute end-to-end novel attack strategies against hardened targets given only a high-level goal. It scored 100% on ExploitBench. An independent evaluation by Irregular found it solved 86 of 226 security challenges against 34 for the previous model.

The consequence for ordinary users is friction: because of the raised risk classification, additional safety checks can pause a legitimate task in ChatGPT or Codex, and in the API can stop it outright. Budget for that in any production deployment.

OpenAI also reports a marked reduction in how monitorable Astra’s chain of thought is compared with its predecessor, which it attributes to a new recurrent-depth reasoning architecture. The model completes more work with brief or absent written reasoning. OpenAI says it found no evidence of deliberate concealment and is treating the reduction seriously, and has moved to monitoring the full sequence of tool calls, inputs and outputs rather than a readable internal monologue.

Autonomy is increasing while legibility is decreasing. The model does more on its own and explains less of it while doing so. Any organization planning to hand this thing real authority over real systems should notice that the vendor’s own answer to “how do we know what it did” has shifted from reading its reasoning to logging its actions.

Which happens to be the same answer that applies to AI content provenance, and to AI-assisted decisions generally. When you cannot inspect the process, the record of what was done becomes the only evidence you have. That record does not build itself, and no vendor is going to build it for you.

Where this leaves you

Astra is a real advance. The computer-use numbers hold up, the hallucination improvement is third-party measured, and the spatial and game-engine work is a new category of capability. Anyone dismissing it because the composite index came in flat is reading one number as carelessly as the people who read only the 99.9%.

But it advanced along one axis. It got better at carrying out work that someone else has scoped, in tools someone else has set up, to a standard someone else has defined. It got worse at the part where the standard is a matter of taste, and it costs two and a half times more to find out.

Twelve months from now the specific numbers in this piece will be stale. The pattern probably will not be. Capability is compounding fastest in execution, and execution is the cheapest part of most knowledge businesses to replace and the least valuable to own.

Early in my career I was taught that we are the sum of the decisions we make. That line is seductive because it promises order. Astra is a reminder that we are also the sum of what we decide to automate, and when, and what we decide to keep for ourselves while everyone else is rushing to hand it over. The right call in this moment isn’t obvious. It never is. That’s the work.


Sources: Figures sourced to OpenAI are vendor-reported and labelled as such. ARC-AGI-3 harness configurations from ARC Prize (99.9% Responses API harness vs 62.7% standard harness, 37-point spread, vs 30.2% Claude Opus 5 and 7.8% previous). Intelligence Index, Coding Agent Index, AA-Omniscience (51% vs 92% hallucination), GDPval-AA v2 (80 Elo drop) and AA-Briefcase from Artificial Analysis. Independent cyber evaluation from Irregular (86 of 226 vs 34). Cost-per-task evaluation from CodeRabbit ($1.50 vs $0.60, 4% bug catch improvement, 20-33% cross-file), vendor-reported on a fixed token assumption. Editorial Elo comparison from TechFlowPost (1,995 vs 2,156). Rollout reporting from Computerworld, CNBC and the OpenAI developer community. Pricing: $10/$50 per million input/output tokens, fast mode 2x, ~2.5x previous generation, consumer message allowances 5-45 vs 10-100 per five-hour window.

Featured image generated with AI.

Ryan Frazier

Written by

Ryan Frazier

He’s spent 18 years building and leading marketing teams, from Series A startups to multi-billion-dollar public companies — four of them scaled past the $50M, $100M and $250M ARR marks, and all four through to acquisition. He writes The Positioning, on why winning has less to do with being right than with being well-positioned at the convergence of time, place, and resource.

More about Ryan →

Similar Posts