AIThis post was created with the assistance of artificial intelligence (AI).

🔍 Read the full analysis: Breaking Down The Astra Vs Fable Benchmark: The Shift From Five To Two Points on ThorstenMeyerAI.com

TL;DR

Recent revisions to the Artificial Analysis Intelligence Index have dramatically altered Astra and Fable scores, revealing that previous comparisons were based on outdated and inconsistent data. The shift highlights issues in benchmarking practices and model evaluation metrics.

Recent updates to the Artificial Analysis Intelligence Index have significantly altered the scores of GPT-6 Astra and Fable 5.1, reducing the previously reported five-point gap to just two points. This revision affects how these models are compared in terms of performance and cost-efficiency, raising questions about the reliability of benchmark-based claims.

Initially, circulating comparisons claimed that Fable 5.1 scored 66 on the AI Index while Astra scored 61, suggesting a performance lead for Fable. However, these figures were based on an earlier version of the Index. Following a recent revision—moving from version 4.1.1 to 4.2—scores for both models shifted, with Fable now scoring around 57 and Astra around 55, significantly narrowing the performance gap. The change stems from re-scoring models against a different evaluation basket, introducing variability that was not clearly communicated.

Furthermore, the benchmarking organization, Artificial Analysis, clarified that Astra’s improved economic profile—being cheaper per task—does not translate into superior general intelligence per dollar. Their own data indicates Astra is less efficient in the broader Intelligence Index but excels in coding tasks due to token reductions. The discrepancy arises because Astra’s architecture, which reasons in latent space without emitting tokens, is not accurately represented by token-based efficiency metrics. The earlier comparison, which used token counts as a proxy for compute, is now recognized as misleading.

Experts emphasize that the core issue lies in the evolving architecture of Astra, which employs a looped or recurrent transformer model that reasons internally without tokenized verbalization. As a result, traditional token-based benchmarks do not capture the true computational effort or intelligence of Astra, leading to distorted comparisons when using static index scores.

At a glance
updateWhen: ongoing; revisions occurred around Astr…
The developmentThe Artificial Analysis Intelligence Index was revised, causing a significant change in Astra and Fable scores, which impacts previous performance assessments and economic evaluations.
Five Points That Became Two — Reality Check
AI Dispatch · Reality Check · 5 September 2026

Five points that became two: what’s wrong with the Astra vs Fable benchmark

The comparison everyone is quoting — Fable 66, Astra 61, “not a rounding error” — is built on numbers that were stale when written, measuring a quantity that no longer means what it used to, aggregated in a way that hides the reversals that matter. The benchmark isn’t broken. The way it’s being read is.

Problem 1 — the numbers moved: Index v4.1.1 → v4.2 during launch week
as quotedAA today (v4.2)
Claude Fable 5.16657 max effort
GPT-6 Astra6155 max · a third source says 60
gap: 5 2 “Five points is not a rounding error.” Two points, on an aggregate of ten evals that just swapped three of them (GPQA Diamond out; AA-Briefcase + GDP.pdf in), is exactly a rounding error. Not AA’s fault — revising an index is how you keep it honest. The error is downstream: quote the version, or don’t quote the number.
◆ Problem 3 — the tell: Astra’s effort dial isn’t connected to the engine
non-reasoning554.4M tok · $990
medium52
xhigh54$2,778
max55$3,020
Non-reasoning = max. Same score, 3× the cost. Because Astra is reported to be a looped / recurrent-depth transformer — it reasons in latent space, without emitting tokens. The Index prices cost in tokens, measures verbosity in tokens, computes time in tokens. For this architecture it’s counting the receipt, not the work. “140M vs 42M tokens” compares Fable’s verbalized reasoning to Astra’s post-loop output — an artefact, not an efficiency finding. Nobody outside OpenAI knows what the loops cost in GPU-seconds.
The other three problems
02
AA’s own conclusion is the opposite of the story
AA’s benchmarking note: Astra is 75% more expensive than GPT-5.6 Sol at max effort and “largely sits behind its predecessor on the Intelligence Index vs cost frontier.” Price went 2.5× ($4/$20 → $10/$50); token savings only partly offset it. The genuine efficiency win lives in one place: the Coding Agent Index, where Astra equals Fable 5 at under half the cost. “Astra attacks the economics” stretched a true coding result over an intelligence index where AA says the reverse.
04
“Max effort” isn’t the same experiment twice
Fable at max = more tokens. Astra at max = ~nothing (see ladder). And OpenAI’s docs say Astra does not support `none` reasoning effort — yet AA lists a “non-reasoning” score. The most efficient-looking config on the leaderboard may not be one you can buy.
05
The aggregate hides the reversals
Index: Fable +2. OpenAI’s own evals (self-reported): Astra ahead 6 of 7 — AutomationBench, BenchCAD, Terminal-Bench 4.0, DeepSWE, TB-Science, FrontierMath T4; Fable takes HLE+tools. A 6–1 task split became a two-point average, and the average became the story. Ten choices deep, two points is noise wearing a number.
✓ What actually changed — and it’s not on the leaderboard
Hallucination rate 92% → 51% on AA-Omniscience — a 41-point drop; matters more than any 2 Index points “Same headline price” hides cache read $1.00 vs $0.25 (4×) + a 25% cache-write premium — the line that dominates agentic bills Coding Agent Index: Astra = Fable 5 at < half the cost — real, and narrow
The take

Three things happened at once: the Index was revised (five became two), the architecture changed (tokens stopped being compute), and the aggregate did what aggregates do (6–1 became +2). A leaderboard position now tells you less than it ever has — and the more advanced the architecture, the less it tells you. Latent reasoning is only the first architecture to break the token proxy. So with your Astra access: ignore the Index number. Take your ten real tasks. Run both models at the effort setting you’ll actually pay for. Measure the bill including the cache line. Measure the failure rate — the 41-point hallucination drop is the one number here I’d bet money on. The benchmark can’t decide for you anymore.

Sources: Artificial Analysis Intelligence Index v4.2 / v4.1.1 model pages (Fable 5.1; Astra non-reasoning/low/medium/high/xhigh/max — scores, tokens, total cost, cost-per-task method incl. cache-write & reasoning tokens) and “Benchmarking GPT-6 Astra” (75% vs Sol, cost frontier, Coding Agent Index, 92%→51% hallucination, 2.5× price, cache terms); OpenAI GPT-6 Astra developer docs (`none` unsupported, cache-write billing, logprobs removed); Alan D. Thompson, The Memo 4 Sep 2026 (looped-transformer read, unconfirmed); OpenAI’s self-reported Astra-vs-Fable table; the circulating 66/61 comparison (pre-v4.2). Scores are version-dependent and were changing at time of writing. Not investment advice.
thorstenmeyerai.com

Implications for AI Performance and Benchmark Reliability

The recent revision of the AI Index and the shifting scores highlight the challenges in benchmarking AI models accurately. As architectures evolve—particularly with models reasoning in latent space—traditional token-based metrics become less reliable indicators of true performance or efficiency. This development underscores the importance of transparent, architecture-aware evaluation methods and cautions against overreliance on static benchmark figures for assessing model progress or economic value.

For users, investors, and developers, this means that claims of superiority based solely on previous benchmark scores may be outdated or misleading. The focus should shift toward understanding the underlying architecture and the specific metrics that best reflect a model’s capabilities in real-world tasks. The broader lesson is that AI benchmarks must adapt to architectural innovations to remain meaningful and trustworthy.

Evals for AI Engineers: Systematically Measuring and Improving AI Applications

Evals for AI Engineers: Systematically Measuring and Improving AI Applications

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Background on Astra, Fable, and Benchmark Changes

The Artificial Analysis Intelligence Index has been a key reference point for comparing AI models, with scores influencing perceptions of model performance and cost-effectiveness. Originally, the index measured verbalized reasoning tokens and output tokens, serving as a proxy for compute and intelligence. Astra, a new GPT-6 model, was introduced with architectural innovations—looped transformers that reason in latent space—aimed at improving efficiency.

Initial reports suggested Astra outperformed Fable 5.1 in both performance and economics, based on token counts and index scores. However, these comparisons relied on an earlier version of the index, which was later revised. The recent update, reflecting architectural changes and new evaluation baskets, caused scores to shift, revealing that earlier claims were based on outdated data. Experts note that the index’s methodology has not kept pace with architectural advances, leading to potential misinterpretations.

Additionally, the debate over Astra’s true efficiency has intensified, as its architecture reduces token use during reasoning, but this is not captured by the index’s token-based metrics. This discrepancy underscores the evolving complexity of AI benchmarking and the need for more architecture-aware evaluation methods.

“The benchmark scores we see are often outdated the moment they’re published, especially when models and evaluation methods are rapidly evolving.”

— Thorsten Meyer, AI researcher

Amazon

AI model performance evaluation software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Remaining Questions About Astra’s True Efficiency

It remains unclear how Astra’s architectural innovations—particularly its latent-space reasoning—translate into real-world compute costs and performance. OpenAI has not publicly disclosed detailed metrics on GPU-seconds or energy consumption for Astra’s latent loops, making it difficult to quantify true efficiency gains. The impact of these architectural features on general intelligence metrics versus coding-specific tasks is also still under debate.

Further, the extent to which the revised index accurately captures Astra’s capabilities, given its architectural differences, is uncertain. As benchmarking methods evolve, the community questions whether current metrics can keep pace with AI model innovation.

Amazon

AI model efficiency testing kits

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Future Steps for Benchmarking and Model Evaluation

Moving forward, researchers and evaluators are expected to develop new benchmarks that account for architectural differences, especially for models reasoning in latent space. OpenAI and other organizations may publish more transparent metrics, including GPU-hours and energy consumption, to better reflect true compute costs.

Additionally, the AI community is likely to revisit existing benchmarks, updating evaluation baskets and scoring methods to ensure they remain relevant. Stakeholders should approach current scores with caution and prioritize understanding the underlying architecture and evaluation methodology.

Finally, as models continue to evolve, transparency around architecture-specific efficiencies and capabilities will become increasingly important for accurate assessment and comparison.

Amazon

AI coding task performance tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Key Questions

Why did the Astra and Fable benchmark scores change?

The scores shifted due to a revision of the Artificial Analysis Intelligence Index, which updated its evaluation baskets and scoring methodology, reflecting Astra’s architectural innovations and leading to different scores than earlier versions.

Does Astra outperform Fable in general intelligence?

Based on the latest AI Index scores, Astra does not outperform Fable in general intelligence per dollar. Its architectural design favors coding efficiency, but its broader intelligence metrics are now seen as less favorable.

Can token counts reliably measure Astra’s efficiency?

No. Astra’s architecture reasons in latent space without emitting tokens, so token-based metrics no longer accurately reflect its true compute costs or efficiency.

What does this mean for AI benchmarking?

It indicates the need for new evaluation methods that consider architectural differences, moving beyond token counts to more comprehensive metrics like GPU-hours, energy use, and task-specific performance.

Will Astra’s performance improve with future updates?

Potentially, but current data suggests architectural innovations primarily improve cost-efficiency in specific tasks, not necessarily overall intelligence. Future updates may refine its capabilities further.

Source: ThorstenMeyerAI.com

You May Also Like

Novak Djokovic has a new job — advisor to private equity firm General Atlantic

Tennis star Novak Djokovic becomes a global strategic advisor for private equity firm General Atlantic, focusing on health, wellness, and sports investments.

AWS Raises Prices for Nvidia Compute by 20%

Amazon Web Services has raised prices for Nvidia-based compute instances by 20%, impacting cloud users reliant on Nvidia GPU services.

The Slate Auto pickup truck starts at $24,950

The Slate Auto electric pickup truck begins at $24,950, making it the most affordable EV and truck in the US market, with preorders opening today.

Japan’s LNG carrier revival pins hopes on help from South Korea

Japan seeks to restart LNG carrier construction with help from South Korea, but legal and technological hurdles remain.