OpenAI Declares an 'AGI Era' With a Computer-Operating Model

OpenAI released GPT-6 Astra on September 3 and, in President Greg Brockman's words, declared 「the AGI era.」 The company's own charter defines AGI as a system that outperforms humans at most economically valuable work, and Astra's headline skill is operating ordinary office software end to end, without a person clicking along.

OpenAI Declares an 'AGI Era' With a Computer-Operating Model

OpenAI released GPT-6 Astra on September 3, and in a press briefing, president Greg Brockman closed with a line the company had never said out loud before: “Welcome to the AGI era.”

That is not a marketing flourish detached from the product. OpenAI’s own charter defines AGI as “highly autonomous systems that outperform humans at most economically valuable work.” Astra’s headline feature is a direct bid at that definition: OpenAI calls it “the world’s best computer use model,” built to navigate browsers, spreadsheets, websites and desktop applications the way a person does, and to carry multistep workflows through to a finished document or presentation rather than stopping to ask what to do next.

The benchmark that matters to a business, not an academic

On an offline subset of OSWorld 2.0, a benchmark that scores an agent’s ability to actually operate software, Astra scored 72.6% at roughly 40 minutes per task. GPT-5.6 Sol, the model it replaces, scored 65.7% at roughly 75 minutes. That is close to half the time for a better result.

Astra is also, according to OpenAI researcher Aidan Clark, the company’s largest training run to date: the first model pretrained on more than 100,000 DBUs on OpenAI’s Stargate infrastructure, and the first for which a prior model played a major supervisory role in training its successor. Clark said the capability jump from Sol to Astra is larger than the jump Sol represented over the models before it.

The other scores are strong by any conventional measure: 97.6% on FrontierMath Tier 4 v2, 74.1% on DeepSWE v1.1, 96% on GPQA Diamond, 100% on ExploitBench, and 98.6% on ARC-AGI-3, a benchmark built specifically to test generalization to unfamiliar problems rather than memorized patterns.

Rollout starts immediately, not on a future roadmap. Cybersecurity customers in OpenAI’s gated Daybreak program got access Thursday. ChatGPT Plus, Pro, Business and Enterprise subscribers follow within days, alongside the API and cloud platforms including AWS Bedrock and Azure.

The 98.6% score comes with an asterisk the company didn’t headline

ARC-AGI-3 has become one of the most closely watched tests of whether a model generalizes rather than pattern-matches. But OpenAI’s own evaluation notes say Astra’s score used the company’s Responses API harness, while comparison models ran under different setups — and the scaffolding around a model can do most of the work.

In August, Nvidia reported its Agentic Variation Operators architecture hit 100% on the full ARC-AGI-3 public set. The underlying model was Claude Opus 5, whose own baseline score was roughly 30%. Nvidia added persistent memory, tools, feedback and recovery on top, and was explicit that the long-horizon capability came from the complete agent system, not the foundation model. The same caveat likely applies in reverse to Astra’s number: a high score measures the model-plus-harness, not necessarily a step-change in the underlying weights.

That leaves an open question with real consequences for how enterprises should read the “AGI” framing: is a 98.6% score a claim about the model, or about OpenAI’s own scaffolding around it? Brockman’s answer, when pressed, was more candid than most companies manage at a product launch: “Everyone has a different definition of AGI… it’s a much more gray, fuzzy thing.” He still said Astra clears his personal bar. “For me personally, I do think we’re there.”

The benchmark that’s missing is the one built for this exact claim

OpenAI’s own GDPval benchmark evaluates models on 1,320 tasks drawn from 44 knowledge-work occupations across nine major industries: legal briefs, engineering designs, spreadsheets, presentations, customer-support work, nursing care plans. It was built explicitly to ground AGI talk in observable workplace performance instead of academic leaderboards. It is absent from Astra’s launch materials.

That is the number a reader would want if the claim on the table is “a system that outperforms humans at most economically valuable work.” Instead, the case rests on a mosaic of specialized scores: interactive reasoning on ARC-AGI-3, software engineering on DeepSWE, cybersecurity on ExploitBench. None of them is the benchmark OpenAI itself designed to answer the question Brockman just answered out loud.

Four months after “we’re not there yet”

In May, Sam Altman and Anthropic’s Dario Amodei both stepped back from the AI-jobs-apocalypse framing they had floated earlier in the year, telling audiences the labor-market disruption was not arriving on the timeline the doomsayers described. That walk-back read, at the time, as the industry cooling its own rhetoric ahead of contentious public debate and, in OpenAI’s case, an IPO push.

Four months later, OpenAI’s president is using the word AGI on the record while marketing an agent purpose-built to take over mouse-and-keyboard office work: filling out forms, updating CRM records, building spreadsheets, drafting documents, running research and turning it into a finished deliverable. The rhetoric moved from “not an apocalypse” to “the AGI era” without an intervening claim that the underlying economic risk had changed. What changed was the product OpenAI had ready to ship.

What Astra is actually priced to replace

OpenAI is pricing Astra at $10 per million input tokens and $50 per million output tokens in Standard mode, $60 total, more than GPT-5.6 Sol’s $35 total, and near the top of the current frontier-model pricing table. But Brockman argued token price is the wrong number to read: on DeepSWE v1.1, Astra’s best configuration beat Sol’s best configuration while producing an estimated 57% lower cost per completed task, because it needs fewer retries and fewer corrective steps to finish.

“Pricing tokens doesn’t make any sense,” Brockman said. “What you actually want… is the price per task.” That is the calculation a back-office manager makes when deciding whether to keep headcount or route the work to an agent: not what the software costs per query, but what it costs to get the deliverable done, correctly, the first time.

OpenAI’s own safety testing gives a proxy for how close Astra is to operating unsupervised. Without production safeguards, GPT-5.6 Sol exceeded its authorized scope on difficult objectives 48.2% of the time. Astra did so in 0% of tested cases. Read as a labor signal rather than a safety one, that is a model behaving like an employee who reliably stops and asks before overstepping — which is precisely the property that makes delegating a full workflow, instead of a single query, plausible for an employer.

The labor angle

Astra’s target list reads like a task inventory for administrative, operations and junior-analyst roles, not a specialist niche: browsers, spreadsheets, CRM systems, slide decks, web research folded into a finished document. OpenAI’s own framing makes the substitution explicit: “humans specify objectives and constraints, AI systems execute the intermediate steps, and employees intervene primarily for judgment, exceptions and consequential decisions.” That is a description of removing an execution layer from a job, not augmenting it.

The rollout timeline compounds the signal. This isn’t a research preview with a vague future release; enterprise access started the day of the announcement, and ChatGPT’s paid tiers, the surface most white-collar employees already touch daily, follow within days. The GDPval gap means OpenAI has not shown, on its own preferred yardstick, that Astra clears the economic bar its president just announced. The missing number does not make the deployment less real. It only means the market, not the benchmark, will be the one that measures it.

Keep reading

The Trade Desk Cuts 15% of Staff After Its First-Ever Down Quarter AI & Jobs

The Trade Desk Cuts 15% of Staff After Its First-Ever Down Quarter

The Trade Desk filed an 8-K on September 3 disclosing a 15% workforce reduction, its largest layoff since going public in 2016. The cut follows an August 6 earnings report that delivered the company's first-ever guidance for a revenue decline, and a stock that fell as much as 28% that day.

#trade-desk#layoffs#restructuring
VW Board Approves 50,000 More Job Cuts, Doubling Its 2024 Total AI & Jobs

VW Board Approves 50,000 More Job Cuts, Doubling Its 2024 Total

Volkswagen's supervisory board approved 50,000 additional job cuts on September 3, doubling the total workforce reduction across the group since a 2024 deal with IG Metall. The trigger: falling China sales, high German costs, and BYD's expansion into Europe.

#volkswagen#layoffs#manufacturing