---
title: "DeepSeek V4.1 Pro: Release Date, What’s Confirmed, and How to Prepare (2026)"
slug: deepseek-v4-1-pro-release-date-what-is-known-how-to-prepare-2026
description: "DeepSeek has confirmed V4.1 Pro is coming but hasn’t dated it. What DeepSeek actually said, what the V4.1 Flash architecture tells you about Pro, the V4 timeline as a guide, and what V4 Pro users should do now."
author: "Yotta Labs"
date: 2026-09-21
categories: ["Inference"]
canonical: https://www.yottalabs.ai/post/deepseek-v4-1-pro-release-date-what-is-known-how-to-prepare-2026
---

# DeepSeek V4.1 Pro: Release Date, What’s Confirmed, and How to Prepare (2026)

![](https://cdn.sanity.io/images/wy75wyma/production/42261ebc8630b54f1be57edb478ab67c5aa2f80f-2240x1260.png)

*DeepSeek V4.1 Pro isn’t out. DeepSeek has said it’s coming, twice, in the same announcement. Here’s what’s confirmed, what’s inference, and how to be ready.*

Search for “DeepSeek V4.1 Pro release date” and you get pages about V4 Pro, V4.1 Flash, or both, none of which answer the question. The short answer is that there is no date. The longer answer is more useful: DeepSeek named the model, tied V4 Pro’s future to it, and shipped the small version of its architecture on September 10. That’s enough to plan around, as long as the confirmed parts stay separate from the guesses.

This post keeps them separate. It will be updated the day V4.1 Pro ships.

## TL;DR

- DeepSeek V4.1 Pro has no announced release date, size, price, or benchmarks
- What DeepSeek has confirmed: V4.1 Flash is “the smallest model in our new architecture family,” and the V4 Pro transition plan runs “until V4.1-Pro launches”
- DeepSeek first said V4 Pro requests would route to V4.1 Flash from September 14, then reversed that before it took effect; V4 Pro stays on the API at unchanged prices until V4.1 Pro arrives
- The V4 generation went Flash on July 31 to Pro GA on August 13, thirteen days apart; V4.1 Flash shipped September 10
- The architecture is public in the Flash: causal encoder-decoder, 8B active parameters on input and 16B on output, native image input, a large Engram memory component, 1M context
- Nothing called V4.1 Pro exists on any API, on Hugging Face, or on [Yotta AI Gateway](https://www.yottalabs.ai/ai-gateway) yet

## DeepSeek V4.1 Pro status at a glance

<!-- unsupported block: table -->

## What DeepSeek has actually confirmed

Two sentences, both from the [V4.1 Flash announcement](https://www.yottalabs.ai/post/deepseek-v4-1-flash-pricing-specs-v4-pro-routing-2026) on September 10.

The first is the family framing: V4.1 Flash is “the smallest model in our new architecture family, with native visual understanding.” Smallest implies larger, and a family implies more than one member.

The second is the one that names the model. Explaining the V4 Pro phase-out, DeepSeek wrote that V4 Pro requests would route to V4.1 Flash “until V4.1-Pro launches.” That’s a product name and a commitment that it ships, from the vendor, in writing. It is not a date.

Then the plan changed. Before the routing took effect, DeepSeek’s API changelog added: “In response to user demand, we have decided to continue providing API services for DeepSeek V4 Pro after September 14, 2026, with the billing method remaining unchanged.” So V4 Pro stays up at its existing prices. The reversal says nothing about V4.1 Pro’s timing either way; it says enough customers wanted the 1.6T model kept alive that DeepSeek chose not to force them onto a Flash-class model in the gap.

Everything below is inference from those facts, from the V4 rollout, or from the Flash architecture, and it’s labeled as such.

## The release date: what the evidence says

There is no official date, and unlike Qwen 4, there’s no leak worth weighing either. What exists is cadence.

**The V4 timeline.** DeepSeek previewed the V4 family on April 24, released V4 Flash with open weights on July 31, and took V4 Pro to GA on August 13. Flash to Pro was thirteen days at the GA stage, though Pro had existed in preview since April. The [full V4 timeline](https://www.yottalabs.ai/post/deepseek-v4-release-date-specs-how-to-access-2026) is the closest precedent DeepSeek has set.

**The V4.1 timeline so far.** V4.1 Flash shipped September 10 with no preview period and no Pro preview alongside it. That’s a different shape from V4, where both tiers were previewed together. Read one way, it means Pro is further behind Flash this time. Read another, it means DeepSeek shipped Flash the moment it was ready rather than holding it for a joint launch, which says nothing about Pro’s state.

**The routing plan.** DeepSeek was prepared to send every V4 Pro customer to a Flash-class model for the duration of the gap. A company expecting a two-week gap doesn’t usually bother with a routing plan and a public phase-out notice. That reads as a gap measured in more than days, but it’s a reading, not a statement.

The practical version: plan for V4.1 Pro in the coming weeks to months, don’t plan around a specific date, and expect the pattern V4 set: API first, weights on Hugging Face shortly after, third-party gateways within days.

## What V4.1 Pro will probably look like

If V4.1 Flash is the small version of the architecture, and DeepSeek says it is, the [V4.1 Flash specs](https://www.yottalabs.ai/post/deepseek-v4-1-flash-pricing-specs-v4-pro-routing-2026) are the blueprint.

The Flash is a 552B-parameter mixture of experts with a causal encoder-decoder layout: 8B parameters active on input, 16B on output, 20 encoder and 20 decoder layers, 384 routed experts plus one shared. It carries a 196.6B-parameter Engram memory component alongside the transformer, uses a compressed attention scheme that brings the KV cache to 890 bytes per token, takes images natively, and runs a 1M-token context with up to 384K output. DeepSeek’s headline for the efficiency work is a quarter of the HBM and an eighth of the SSD storage of the previous generation.

Scale that up and the expected shape of V4.1 Pro is: total parameters well above the Flash’s 552B, an active count still small relative to total, the same encoder-decoder and Engram design, native vision as standard rather than an “-Exp” variant, and a KV cache that stays tiny even at 1M context. Whether DeepSeek returns to a V4 Pro-scale 1.6T total, or lands somewhere below it because the new architecture gets more out of each parameter, is unknown. DeepSeek’s own comparison puts V4.1 Flash ahead of V4 Pro on its agentic benchmarks (74.2 vs 62.7 on DeepSWE, 31.2 vs 12.4 on Terminal-Bench 4.0) while V4 Pro keeps the edge on knowledge tests (92.4 vs 90.9 on GPQA, 42.7 vs 36.8 on Humanity’s Last Exam). A Pro-scale model in the new architecture is presumably meant to close that second gap.

Pricing is the other open question. V4 Pro is $1.32 in and $3.96 out per million tokens at peak on DeepSeek’s API, half that off-peak. V4.1 Flash came in at $0.30 and $1.20 peak, under the V4 Flash it replaced. A V4.1 Pro priced under V4 Pro would fit the pattern; anything specific is a guess.

## What V4 Pro users should do now

The reversal bought time, not certainty. DeepSeek said it is phasing out V4 Pro before it said it would keep serving it, and the second statement came from customer pressure. Treat V4 Pro as a model with a known successor and an unknown end date.

Three things pay off whichever way the timing goes.

**Run V4.1 Flash against your V4 Pro workload now.** DeepSeek’s claim that Flash beats Pro is on agentic benchmarks and on cost, speed, and total runtime. If your work is agentic or code-heavy, you may find the downgrade isn’t one and the bill drops by roughly 70 percent. If it’s knowledge-heavy, you’ll find out where Flash falls short, which tells you exactly what to test the day V4.1 Pro appears. The [V4.1 Flash vs GLM 5.3 Flash comparison](https://www.yottalabs.ai/post/deepseek-v4-1-flash-vs-glm-5-3-flash-2026) has the independent speed and quality numbers for the Flash tier.

**Keep V4 Pro behind an OpenAI-compatible interface.** Every DeepSeek generation change has been a model-string swap for teams set up this way. DeepSeek V4 Pro and V4 Flash are live on [Yotta AI Gateway](https://www.yottalabs.ai/ai-gateway) at flat rates with no peak window ($0.99 / $2.97 for Pro), alongside GLM 5.3, Qwen 3.8, Kimi K3, and Grok 4.6, which makes the V4.1 Pro A/B a routing rule when it lands rather than a migration.

**Plan hardware for the total, not the active count.** If you self-host and V4.1 Pro ships open weights, the active-parameter number will make the headlines and the checkpoint size will decide the bill. V4.1 Flash already moved the floor: 510 GB on disk and about 614 GB of GPU memory, a full node where V4 Flash ran on two H200s. The [V4.1 Flash hardware breakdown](https://www.yottalabs.ai/post/deepseek-v4-1-flash-hardware-requirements-gpu-memory-2026) is the reference; a Pro-scale checkpoint in the same architecture is multi-node territory, and the only question is how many.

## Frequently asked questions

**When is DeepSeek V4.1 Pro coming out?** DeepSeek hasn’t said. The model is confirmed by name in the September 10 announcement, but no date, window, or preview has been given. The V4 generation went from Flash to Pro GA in thirteen days; V4.1 has been Flash-only since September 10.

**Is DeepSeek V4.1 Pro confirmed?** Yes, as a product that will exist. DeepSeek wrote that its V4 Pro routing plan would run “until V4.1-Pro launches” and called V4.1 Flash “the smallest model in our new architecture family.” Nothing beyond the name is confirmed.

**What happened to DeepSeek V4 Pro?** It’s still on the API. DeepSeek announced on September 10 that V4 Pro requests would route to V4.1 Flash from September 14, then reversed that before the change took effect “in response to user demand,” keeping V4 Pro available with unchanged billing.

**Will V4.1 Pro have open weights?** Unknown. V4 Flash, V4 Pro, and V4.1 Flash all shipped under MIT on Hugging Face, so the precedent is strong, but DeepSeek hasn’t said.

**How big will DeepSeek V4.1 Pro be?** Unknown. V4 Pro was 1.6 trillion total parameters with 49B active. V4.1 Flash is 552B total with 8B active on input and 16B on output. Any V4.1 Pro figure you see is speculation until DeepSeek publishes one.

**What will V4.1 Pro cost?** Unknown. V4 Pro is $1.32 / $3.96 per million tokens at peak on DeepSeek’s API and $0.99 / $2.97 flat on Yotta AI Gateway. V4.1 Flash launched cheaper than the Flash it replaced, so a price at or under V4 Pro’s would fit the pattern.

**Should I switch from V4 Pro to V4.1 Flash now?** Test it. DeepSeek’s own numbers have Flash ahead of V4 Pro on agentic and coding benchmarks and behind it on knowledge benchmarks, at less than a third of the price. Whether that trade works depends on your workload, and running the test now means you know your answer before V4.1 Pro forces the question.

**Is DeepSeek V4.1 Pro on Yotta?** It doesn’t exist yet. DeepSeek V4 Pro, V4 Flash, V3.2, and R1 are live on Yotta AI Gateway, and this post will be updated when V4.1 Pro ships and again if it joins the catalog.

## Bottom line

DeepSeek V4.1 Pro is confirmed by name, undated, and described only by the small model that shares its architecture. The reversal on V4 Pro routing means nobody is being forced off the current flagship before the new one exists, and that’s the whole official picture.

The teams that handled the V4 wave well weren’t the ones who guessed the date. They were the ones whose serving was a model string away from the new model and who had already run the previous generation on their own workload. V4 Pro and V4.1 Flash on [Yotta AI Gateway](https://www.yottalabs.ai/ai-gateway) are one API key away, and running Flash against your Pro workload this week is the preparation.
