---
title: "Qwen 3.8 Benchmarks: What's Actually Verified So Far (2026)"
slug: qwen-3-8-benchmarks-what-is-verified-2026
description: "Alibaba says Qwen 3.8 is second only to Claude Fable 5, but hasn't published a single score. Here's every real data point so far, and what to watch next."
author: "Yotta Labs"
date: 2026-07-24
categories: ["Inference"]
canonical: https://www.yottalabs.ai/post/qwen-3-8-benchmarks-what-is-verified-2026
---

# Qwen 3.8 Benchmarks: What's Actually Verified So Far (2026)

![](https://cdn.sanity.io/images/wy75wyma/production/375252261c2d38696944197465a90b1567db8952-1200x627.png)

Alibaba claims Qwen 3.8 is the second-best model in the world. Five days after the WAIC preview, there is still no benchmark table, no model card, and no score behind that claim. Meanwhile, "qwen 3.8 benchmark" is one of the most-searched questions about the model, and most of what's circulating is either speculation or repackaged vendor marketing.

This post is the other thing: every verifiable data point that exists on Qwen 3.8-Max as of July 24, clearly sourced, and updated as real numbers land.

## TL;DR

- Official benchmarks: none. Alibaba's "second only to Fable 5" ranking rests on internal evaluations it has not published
- Independent data point #1: in a real-world architecture evaluation, Qwen 3.8-Max preview scored 80/100, just behind Kimi K3 at 83
- Independent data point #2: community tracking of Code Arena placed a stealth Qwen 3.8 preview in frontier range on coding preference, a fragile early signal, not a score sheet
- The verified family baseline: Qwen 3.7-Max's published numbers (GPQA Diamond 92.4, SWE-bench Verified 80.4, Terminal-Bench 2.0 69.7) are the floor 3.8 should clear
- Pricing: no standalone API pricing exists. Preview access runs at a reported 10% of standard rates through Alibaba's Token Plan and Qoder platforms
- Contrast: [Kimi K3](https://www.yottalabs.ai/post/kimi-k3-specs-benchmarks-how-to-access-2026) already has third-party rankings and published pricing. Qwen 3.8 has hype and a promise
- What flips this post from "wait" to "evaluate": a published benchmark table, a model card, or standalone API pricing

## The Official Claim, and What's Behind It

Alibaba's exact words: Qwen 3.8 is "one of the most powerful models available today, comparable to leading frontier AI models, second only to Fable 5," referring to Anthropic's Claude Fable 5.

What's behind it, as of today: internal evaluations. No benchmark names, no scores, no prompts, no harness, no methodology. The claim is not implausible, the [Qwen 3.7-Max generation](https://www.yottalabs.ai/post/qwen-3-7-max-release-date-features-open-source-status-and-how-to-access-2026) posted genuinely competitive published numbers, but a ranking with no scores attached is marketing until the table ships.

For production teams the distinction matters: you can't size a deployment, estimate cost per task, or compare against your current model using a press quote.

## What's Actually Been Tested

Two independent data points exist so far. Both are early, both are limited, and both are more informative than anything official.

**A real-world architecture evaluation.** An independent evaluator ran Qwen 3.8-Max preview and Kimi K3 through a production-style software architecture task: analyzing 269 files across two real projects and proposing system designs, blind-reviewed and fact-checked. Result: Kimi K3 83/100, Qwen 3.8-Max preview 80/100. Notable detail: none of Qwen's 44 tool calls failed, a good sign for agent reliability. One test, one workload type, but a real one.

**Code Arena community tracking.** Before the WAIC announcement, a stealth preview model widely believed to be Qwen 3.8 ran on Code Arena, where community analysis placed its coding-preference Elo in frontier range. This is the most fragile kind of evidence, unofficial identification, preference voting rather than task completion, but it is consistent with the architecture eval: 3.8 appears to sit near, not above, the current frontier.

The honest synthesis of both: Qwen 3.8-Max preview looks like a legitimate frontier-class model that trades blows with Kimi K3 rather than dominating it. "Second only to Fable 5" remains unproven.

## The Verified Baseline: What Qwen 3.7-Max Actually Scored

The best calibration for 3.8 claims is the family's last published score sheet. Qwen 3.7-Max, released May 19 with a full benchmark table, posted: 92.4 on GPQA Diamond, 80.4 on SWE-bench Verified, 69.7 on Terminal-Bench 2.0, 91.6 on LiveCodeBench, and 90.4 on MRCR-v2 128k, competitive with Claude Opus 4.6 across most of the agentic suite. Full breakdown with pricing context in our [Qwen 3.7 Max vs Claude Opus 4.6 comparison](https://www.yottalabs.ai/post/qwen-3-7-max-vs-claude-opus-4-6-pricing-benchmarks-2026).

Those are the numbers 3.8-Max needs to clear for the generation jump to be real. When the official table drops, that's the first check worth doing: 3.8 vs 3.7's verified floor, not 3.8 vs the marketing.

## Qwen 3.8 vs Kimi K3: The Evidence Gap

The sharpest way to see 3.8's evidence problem is next to its direct rival, launched the same week:

<!-- unsupported block: table -->

K3 has receipts; 3.8 has a claim. That can flip in a single announcement, and this post will be updated the day it does. Full K3 breakdown in our [Kimi K3 guide](https://www.yottalabs.ai/post/kimi-k3-specs-benchmarks-how-to-access-2026).

## Qwen 3.8 Pricing: What We Know

Short version: there is no official standalone API pricing yet.

What exists today is preview access through Alibaba's Token Plan subscription and the Qoder and QoderWork platforms, at a reported 10% of standard rates during the preview window. No per-token price list has been published, which means no cost-per-task math is possible yet, the number that actually decides deployments.

For family context, Qwen 3.7-Max runs $1.25 per million input tokens and $3.75 per million output on [Yotta AI Gateway](https://www.yottalabs.ai/ai-gateway) today. Where 3.8 lands relative to that, and to Kimi K3's aggressive $0.30 cached-input rate, is the pricing question of the quarter. When Alibaba publishes real numbers, we'll break them down the same way we did for [3.7-Max vs Claude Opus 4.6](https://www.yottalabs.ai/post/qwen-3-7-max-vs-claude-opus-4-6-pricing-benchmarks-2026).

## How to Evaluate Qwen 3.8 Without Official Benchmarks

If you can't wait for the table, the playbook is the same one we recommend for every frontier release:

1. **Build your baseline now.** Run your actual workload, your prompts, your tools, your evaluation criteria, against the models you can access today. One OpenAI-compatible API in front of [Qwen 3.7-Max, GLM 5.2's family, DeepSeek, and Claude](https://www.yottalabs.ai/ai-gateway) makes side-by-side runs a config change instead of four integrations.
1. **Get preview access if the workload justifies it.** The Token Plan route works today for hands-on testing at discounted preview rates.
1. **When 3.8's standard API opens, slot it into the same harness.** Your own eval beats any leaderboard: vendor benchmarks, including the eventual official table, are directional at best. The [defensive integration pattern](https://www.yottalabs.ai/post/how-to-run-qwen-3-7-in-production) means adding it is a routing rule, not a rewrite.

## Frequently Asked Questions

**Are there official Qwen 3.8 benchmarks?** No. As of July 24, 2026, Alibaba has published no benchmark table, no model card, and no methodology behind its "second only to Fable 5" claim.

**Is Qwen 3.8 really the second-best model in the world?** Unverified. The only independent data so far suggests it's frontier-class but trades blows with Kimi K3 rather than clearly beating it. The ranking rests on Alibaba's internal evaluations.

**How does Qwen 3.8 compare to Kimi K3 on benchmarks?** In the one independent head-to-head published so far, K3 edged the 3.8 preview 83 to 80 on a real-world architecture task. K3 also has third-party index rankings and published pricing; 3.8 has neither yet.

**How much does Qwen 3.8 cost?** No official API pricing exists. Preview access via Alibaba's Token Plan runs at a reported 10% of standard rates. For family reference, Qwen 3.7-Max is $1.25 in / $3.75 out per million tokens.

**When will official Qwen 3.8 benchmarks be released?** Alibaba hasn't said. The signals to watch: a model card, a benchmark table, standalone API pricing, or the promised open-weight release, any of which would enable independent verification.

**Should I wait for benchmarks before evaluating Qwen 3.8?** Build your evaluation harness now against models you can already access, then slot 3.8 in when its standard API opens. Your own workload results beat any leaderboard.

## Bottom Line

Everything verified about Qwen 3.8-Max fits in one paragraph: it's a 2.4 trillion parameter multimodal MoE, it's frontier-class in the two independent tests that exist, it trailed Kimi K3 narrowly in the only head-to-head, and every stronger claim than that currently traces back to Alibaba's own press materials.

That's not a knock on the model. It's the state of the evidence, and evidence is what deployment decisions run on. This post gets updated the day the official table, model card, or API pricing lands. Until then: the [full 3.8-Max preview breakdown](https://www.yottalabs.ai/post/qwen-3-8-max-release-date-specs-how-to-access-2026) covers access and specs, and the [Kimi K3 guide](https://www.yottalabs.ai/post/kimi-k3-specs-benchmarks-how-to-access-2026) covers the rival that already showed its numbers.
