Every AI model release promises to change how you work. Most of them change how you write.
GPT-6 Astra, released on 4 September 2026, is aimed at something different. OpenAI built it to operate software rather than describe software, and the published evaluations back that up in a few specific areas.
This is a practical breakdown of what it finishes on its own, what it only assists with, and how to tell the difference before you hand it real work.
What It Can Do That the Previous Model Could Not
The largest single change is that Astra can drive a computer interface directly instead of telling you which buttons to press.
In practice that covers filling online forms, updating CRM records, organising a calendar, researching across websites, working inside email and document editors, installing and testing software, and running frontend quality checks on a website it has just built.
On OSWorld 2.0, a standard desktop-task benchmark, Astra scored 72.6 per cent against 65.7 per cent for GPT-5.6 Sol, and completed the runs in roughly 40 minutes instead of 75. On ScreenSpot-Pro, which measures whether a model can correctly identify the right element on a screen, it moved from 76.9 per cent to 92.7 per cent.
Alongside that, three capabilities became genuinely new. Through Sites in ChatGPT it can create, host and share a website, web app or game from a prompt, then run its own frontend checks to confirm the features work. In Codex it can now keep searchable notes across context windows, so requirements and test results from an hour ago stay retrievable during a long build. And it can produce documents, spreadsheets and presentations that follow your existing templates rather than generic output you then have to reformat.
The Success Rates Nobody Prints
OpenAI reports 41.4 per cent on AutomationBench and 40.9 per cent on its internal data science tasks. Those are large improvements over GPT-5.6 Sol at 18.1 per cent and 30.5 per cent. They are also, read plainly, failure rates of roughly 59 per cent.
Even the strongest result, 72.6 per cent on OSWorld 2.0, means the model does not complete a bit more than one desktop task in four.
This is assistance at a much higher level, not delegation. Anyone selling you a fully autonomous AI worker on the back of these numbers has not read them.
Where It Finishes the Job Reliably
The picture changes once you narrow to the tasks where the measured scores are high rather than merely improved.
1. Screen and interface grounding: At 92.7 per cent on ScreenSpot-Pro, it reliably finds the right control. Failures tend to come later in a chain, not at the click.
2. Long-document retrieval: On MRCR v2 it recalled 100 per cent of planted facts in the 256,000 to 512,000 token range and 96.3 per cent between 512,000 and one million. Contracts, tenders and multi-year records are genuinely inside its reach.
3. Infrastructure and incident work: On SRE-Bench it solved 88.0 per cent of problems on a single attempt, against 55.9 per cent for the previous model.
4. Technical drawing and structured professional output: BenchCAD came in at 95.9 per cent against 83.3 per cent.
5. Factual reliability: OpenAI's internal hallucination benchmark dropped from 12.2 per cent to 4.2 per cent.
There is one clear counterweight. Independent evaluation by Artificial Analysis found Astra scored roughly 45 Elo points below GPT-5.6 Sol on GDPval-AA v2, a benchmark of economically valuable professional work across 44 occupations, with measured declines in writing and presentation quality. For drafting, editing and client-facing prose, the newer model is a step back.
What This Changes in a Working Day
The useful mental shift is from asking for an answer to handing over an outcome.
The old pattern was: ask the model, receive instructions, perform the task yourself. The new pattern is: state the goal, let the model work through the interface steps, then review what it produced.
For a small business that maps onto a short list. Researching a lead and writing the findings into your CRM. Reading a stack of supplier PDFs and producing a comparison sheet. Building an internal tool and checking that its buttons work. Pulling a quarter of operational data into a structured analysis. None of these are experiments any more, but all of them still need a human reviewing the output before it leaves the building.
How to Start Without Wasting Money
Do not migrate anything on launch-day enthusiasm. Take three tasks your team repeats weekly, run ten real examples of each through Astra, and record two numbers: how often it finished without intervention, and how long your review took.
If the completion rate beats 70 per cent and review takes under five minutes, automate it. If completion sits near 40 per cent, keep the model as an assistant and keep the human in the loop. If the task is writing, stay on your current model until the writing regression is addressed.
One access note that catches people out. On ChatGPT Plus, Astra is currently available only inside Work and Codex, and message limits are tighter than before, with users reporting roughly 5 to 45 messages per five-hour window. In enterprise workspaces it ships switched off, so an administrator has to enable it before anyone can test anything.
Aivaura builds AI automation for small and medium businesses. We measure completion rate on your tasks before we recommend a model.
Aivaura Studio
Premium custom software and AI integration agency. Bhavnagar, Gujarat.