The temptation is to declare the AI race over. GPT-6 Astra beats Anthropic’s flagship on many practical benchmarks in OpenAI’s launch data and is designed to work across browsers, terminals and professional software. Yet Claude Fable 5.1 still leads on two broad reasoning measures, matches Astra’s headline token price, and has a caching advantage.
The short answer is that GPT-6 Astra changes the frontier, but it does not kill the competition. It turns the contest from “who has the smartest chatbot?” into “whose agent finishes more valuable work for each dollar, minute and human intervention?” OpenAI has taken a powerful lead, not permanent ownership of the market.
First, yes: Anthropic’s strongest public model is Fable 5.1
Claude Fable 5.1 launched on September 1, 2026, and Anthropic describes it as its most capable generally available model for demanding reasoning and long-horizon agentic work. Claude Mythos 5.1 uses the same underlying model with different safeguards, but access is restricted to vetted organizations. The fair public-market comparison is therefore GPT-6 Astra versus Claude Fable 5.1.
Anthropic’s official Fable 5.1 announcement emphasizes coding, knowledge work and scientific research. OpenAI positions Astra around end-to-end computer work. These are direct competitors for valuable agentic workloads.
GPT-6 Astra vs Claude Fable 5.1: specifications and price
As of September 4, 2026, the headline specifications are unusually close:
Parameter | GPT-6 Astra | Claude Fable 5.1 |
|---|---|---|
API model ID |
|
|
Context window | 1.05M tokens | 1M tokens |
Maximum output | 128K tokens | 128K tokens |
Standard input price | $10 / MTok | $10 / MTok |
Standard output price | $50 / MTok | $50 / MTok |
Cached input/read | $1 / MTok | $0.25 / MTok |
Cache write | $12.50 / MTok | $12.50 / MTok for 5 minutes; $20 for 1 hour |
Knowledge cutoff | April 30, 2026 | June 2026 |
Inputs and outputs | Text/image in; text out | Text/image in; text out |
The GPT-6 Astra model page confirms its context, output and API pricing. Anthropic’s Fable 5.1 model documentation confirms a one-million-token window, 128K output and the same $10/$50 headline rate.
The price tie is less complete than it looks. Above 272K input tokens, Astra charges 2× input and cache rates and 1.5× output for the full request. Anthropic keeps its one-million-token context at standard rates, and Fable 5.1 charges one quarter as much for cache reads. Yet Anthropic’s newer tokenizer can produce roughly 30% more tokens for the same text than older Claude tokenizers. Nominal rates are not total-cost guarantees.
The benchmark scoreboard: a lead, not a sweep
OpenAI’s launch report publishes direct comparisons with Claude Fable 5.1. These are vendor-published results, and OpenAI says scores use the maximum result at any effort. They are useful evidence, but they are not a neutral, permanent leaderboard.
Benchmark | GPT-6 Astra | Claude Fable 5.1 | Leader |
|---|---|---|---|
AutomationBench | 41.4% | 31.4% | Astra |
BenchCAD | 95.9% | 84.3% | Astra |
Terminal-Bench 4.0 | 57.9% | 55.8% | Astra |
DeepSWE v1.1 | 74.1% | 67.4% | Astra |
Terminal-Bench Science 0.1 | 64.6% | 52.6% | Astra |
FrontierMath Tier 4 v2 | 97.6% | 87.8% | Astra |
HealthBench Professional | 63.4% | 58.1% | Astra |
Artificial Analysis Intelligence Index v4.1.1 | 61.2 | 65.7 | Fable 5.1 |
Humanity’s Last Exam, with tools | 57.2% | 65.0% | Fable 5.1 |
The pattern matters more than the win count. Astra’s largest advantage appears in tool-driven work: business automation, terminal tasks and applied science. OpenAI also reports that Astra scores 72.6% on an offline OSWorld 2.0 set in about 40 minutes per task, versus GPT-5.6 Sol’s 65.7% in roughly 75 minutes—a gain enterprises can translate into throughput.
But Fable 5.1 leads the composite Artificial Analysis Intelligence Index and Humanity’s Last Exam with tools. Anthropic’s own launch data also reports 65.0% on that latter test, along with 73.4% on CursorBench 3.2.0. In other words, Astra looks stronger across many execution-heavy tasks, while Fable remains formidable in broad reasoning, research and coding. Different safeguards and harnesses can also change outcomes; OpenAI explicitly notes several cross-vendor methodology caveats.
Does GPT-6 threaten Anthropic?
Absolutely. The threat is strongest in the category Anthropic helped define: long-running coding and knowledge-work agents. Astra is not merely answering questions better; it is competing to operate the software where work happens. OpenAI’s model guide adds asynchronous tool calls, mid-turn steering and the ability to change reasoning effort without discarding prompt-cache continuity. Those features target production agents, not demo conversations.
Price parity increases the pressure. When two frontier models both list at $10 per million input tokens and $50 per million output tokens, buyers can move the decision toward completion rate, latency and reliability. On the launch benchmarks that resemble real workflows, Astra often has the stronger story.
Still, Anthropic has defenses. Fable 5.1 supports a one-million-token context and 128K output, with distribution through Anthropic’s platform, AWS, Google Cloud and Microsoft Foundry. Its $0.25 cache-read rate can reduce the cost of agents that reuse large repositories or document collections. Anthropic estimates typical savings of about 25% and up to 45% for highly agentic work.
Anthropic can also optimize Claude Code, enterprise workflows and safety controls as one system. Benchmarks move quickly; customer trust, tooling and deployment contracts move more slowly.
Verdict: did GPT-6 Astra kill the competition?
No—but it may have killed the idea that one good coding benchmark is enough to define the frontier. GPT-6 Astra’s release raises the standard to reliable, multi-application execution with measurable speed, cost and safety. On current vendor-published evidence, it holds the more convincing overall lead in applied agentic work.
Claude Fable 5.1 is not obsolete. It wins meaningful reasoning evaluations, offers cheaper cache reads and remains a credible choice for long-context research and coding. The practical verdict is to test both models on complete workflows, measure successful outcomes per dollar, and include intervention rate and wall-clock time. GPT-6 Astra has not ended the race. It has made the next lap much more expensive for Anthropic—and much more interesting for everyone else.








