GPT-6 Astra Is Out: Pricing, Benchmarks and the Migration

GPT-6 Astra Is Out: Pricing, Benchmarks and the Migration

On September 3, 2026 OpenAI shipped GPT-6 Astra — API id gpt-6-astra, the same string that had been answering 404 instead of 400 a day earlier. The rollout comes in waves: today it is organizations in the cybersecurity Trusted Access Program, over the coming days the API plus ChatGPT Plus, Pro, Business and Enterprise plans, along with Microsoft Azure and AWS Bedrock.

Below is what matters for an integration: pricing, benchmarks without the marketing, the list of request parameters you have to drop, and the three new API mechanisms this release was really built around.

Pricing and access

ParameterValue
API idgpt-6-astra (Responses API)
Input$10 per 1M tokens
Output$50 per 1M tokens
Cacheseparate read and write rates
Fast modeup to 2× the speed at 2× the price, no latency SLA
ChatGPT tiersPlus, Pro, Business, Enterprise; Pro/Business/Enterprise also get Astra Pro
Enterpriseoff by default, the workspace admin enables it
ZDRsupported for eligible API customers

The per-token price is higher than GPT-5.6 Sol's, but OpenAI consistently measures cost per task instead: Astra spends markedly fewer output tokens. On Terminal-Bench Science 0.1 it scores 64.6% against 52.6% for Claude Fable 5.1 at roughly 31% lower estimated API cost; on BenchCAD it comes out about 43% cheaper than Sol. On Agents' Last Exam it uses roughly 65% fewer output tokens than Claude Opus 5. This is the case where comparing vendors by price per million tokens tells you nothing — measure the cost of a solved task on your own request set.

Benchmarks

Numbers from OpenAI's table, maximum at any effort level.

BenchmarkAstraGPT-5.6 SolClaude Fable 5.1
Terminal-Bench 4.057.9%37.3%55.8%
Agents' Last Exam59.3%53.6%—
OSWorld 2.072.6%65.7%—
ScreenSpot-Pro92.7%76.9%—
FrontierMath Tier 4 (v2)97.6%83.0%87.8%
GPQA Diamond96.0%94.6%93.7%
ARC-AGI-399.9%7.8%—
ExploitBench100%78.5%70%
MRCR v2, 8-needle, 512K–1M96.3%73.8%—

Two results deserve a closer read. ARC-AGI-3 at 99.9% against 7.8% for Sol: per Greg Kamradt of the ARC Prize Foundation, the model beat their human action-efficiency baseline on 96% of levels. MRCR over the 512K–1M range is the long-context story: 96.3% where Sol drops to 73.8%.

Where Astra is not first

The section the announcement does not have, and the one that matters most if you are picking a model for a job rather than reading the news.

BenchmarkAstraBest in the table
Humanity's Last Exam (with tools)57.2%65.0% — Claude Fable 5.1
Artificial Analysis Intelligence Index v4.1.161.265.7 — Claude Fable 5.1
Artificial Analysis Coding Agent Index v1.467.068.1 — Claude Fable 5
FrontierCode 1.1 Main53.3%53.5% — Claude Fable 5

Astra wins where the job is to act carefully over a long horizon: terminal, browser, third-party software, long context, science. On raw reasoning in the aggregate indices the Claude line is ahead — we covered Fable 5.1 separately. The practical takeaway is dull but real: do not bet everything on one model, and run the eval across candidates with a single key.

One more honest line from the announcement: Astra's reasoning turned out harder to monitor than Sol's. OpenAI attributes this to the model solving problems in fewer written steps and calls reasoning monitorability a continuing research priority.

What breaks when you migrate from GPT-5.6

This is not "change the model name in the config". From the migration guide:

WhatWhat to do
temperature, top_p, top_logprobsremove — unsupported
logprobs (Chat Completions)remove; in Responses drop message.output_text.logprobs from include
reasoning_effort: "none"unsupported; if you were on none or minimal, start at low
Tool callingResponses API only; Chat Completions works, but not tool calling for Astra
prompt_cache_retentionreplace with prompt_cache_options.ttl: "30m" (migrating from 5.5 or earlier)
Fast mode + EU data residencyincompatible, use Standard processing

Behaviour is its own migration. Astra stops and asks a clarifying question where earlier models silently assumed. In an interactive chat that is an improvement; in an overnight pipeline it is a source of stalled jobs. OpenAI publishes counter-prompts for exactly this, and their gist is "treat the request as authorization to act and carry the task to completion". Second: the model is noticeably more sensitive to instructions in context files such as AGENTS.md and skills — reread them before you migrate, or a forgotten line in someone else's skill will start blocking work.

Three new API mechanisms

Async tool calling. Setting async: true on a function lets the model keep reasoning and calling other tools while your code runs a slow one; the result comes back later under the original call_id. For agents with slow tools this removes the main source of idle time.

Mid-turn steering. Over a WebSocket you can send a correction or a changed requirement while the model is working: the Responses API preserves the completed work and continues with your update instead of starting over.

Changing reasoning depth without losing the cache. A configuration_update input item raises or lowers reasoning_effort mid-conversation without rewriting the prompt prefix — which means without invalidating the prompt cache. The old trade-off between "think harder" and "keep the cache" is gone.

Codex also gets notes that survive across context windows instead of compaction: the model keeps its own running notes and earlier windows stay searchable. It is a flag in config.toml today and becomes the default in the coming weeks.

Diagram: config sets the model name, the request goes through one gateway to three models, the third one locked

Cybersecurity: what the model will refuse

Astra is the first OpenAI model to cross the Critical cybersecurity threshold in the Preparedness Framework. It scores 100% on ExploitBench; on an internal set built from June–August 2026 vulnerabilities it found and used two previously unknown zero-days (both disclosed to maintainers); on SRE-Bench it solved 88.0% of tasks on the first attempt against 55.9% for Sol.

The public configuration is trimmed: the model does defensive work such as secure code review and patching, but refuses to build proof-of-concept exploits. Wider access is promised through the OpenAI Daybreak program.

What matters in production: Astra-class models run asynchronous misalignment monitoring. When a check fires, ChatGPT and Codex ask the user to confirm — but in the API the task simply stops. Handle that as a valid response, not as a network error to retry.

What to do today

Astra is already in the GPTunneL catalog — open the model. What still needs preparing is the transport, because a model that thinks for tens of minutes breaks timeouts, not logic. Streaming is mandatory:

JavaScript
const BUDGET_MS = 15 * 60 * 1000;

export async function askStreaming(messages, onDelta) {
  const res = await fetch("https://gptunnel.ru/v1/chat/completions", {
    method: "POST",
    headers: {
      Authorization: process.env.GPTUNNEL_API_KEY,
      "Content-Type": "application/json",
    },
    body: JSON.stringify({
      model: process.env.MODEL_SMART ?? "gpt-5.6-sol",
      reasoning_effort: process.env.MODEL_EFFORT ?? "medium",
      stream: true,
      messages,
    }),
    signal: AbortSignal.timeout(BUDGET_MS),
  });

  const decoder = new TextDecoder();
  let buffer = "";

  for await (const chunk of res.body) {
    buffer += decoder.decode(chunk, { stream: true });
    const lines = buffer.split("\n");
    buffer = lines.pop();

    for (const line of lines) {
      const s = line.trim();
      if (!s || s.startsWith(":")) continue;    // gateway keepalive, not data
      if (s === "data: [DONE]") return;
      if (!s.startsWith("data: ")) continue;
      const delta = JSON.parse(s.slice(6)).choices[0].delta.content;
      if (delta) onDelta(delta);
    }
  }
}

Lines starting with a colon are the gateway's keepalive comments (: HELLO, : PROCESSING); a naive parser trips on them, covered separately. Keep the model name and reasoning_effort in the environment — then moving to any new model is one variable.

The working line-up today, all with a 1M context window:

No subscriptions, pay per use. Current numbers are on the pricing page; the request, streaming and limit specs are in the documentation and at docs.gptunnel.ru.

GPT-6 Astra FAQ

Is GPT-6 Astra out? Yes, announced September 3, 2026. Access today is for Trusted Access Program organizations; the API, ChatGPT Plus/Pro/Business/Enterprise, Azure and AWS Bedrock follow over the coming days.

What does GPT-6 Astra cost? $10 per million input tokens and $50 per million output tokens on standard processing. Fast mode is twice as fast at twice the price.

What is the API id? gpt-6-astra, called through the Responses API. It works in Chat Completions, but without tool calling.

Is GPT-6 Astra available in GPTunneL? Yes — open GPT-6 Astra. The rest of OpenAI's line-up sits next to it: GPT-5.6 Sol, Terra and Luna.

Is Astra the smartest model on the market? In agentic work, computer use and long context — yes, by a wide margin. On aggregate indices such as Artificial Analysis and on Humanity's Last Exam, Claude Fable 5.1 is ahead.

What are the cybersecurity restrictions? The model crossed the Critical threshold, so in its public configuration it refuses to write proof-of-concept exploits while keeping defensive work — code review, patching. Some restrictions are due to lift through OpenAI Daybreak.

The bottom line

Astra is not "Sol plus a few points" but a change of profile: a bet on long agent runs inside real software, where the win comes not from the token price but from the task being solved on the first pass and in fewer steps. The cost is a rewritten integration layer: a different API for tools, temperature and top_p gone, monitoring stops to handle, and a queue instead of a synchronous response.

Open GPT-6 Astra in GPTunneL and run your own request set with streaming and reasoning_effort — on the same balance as the rest of the line-up. If your transport is not ready for long answers yet, start with GPT-5.6 Sol and switch with one line.