---
title: "Mistral Large 4 vs DeepSeek V4 Pro, GLM 5.3, and Kimi K3 (2026)"
slug: mistral-large-4-vs-deepseek-v4-pro-vs-glm-5-3-vs-kimi-k3-2026
description: "Mistral Large 4 lands in the same price band as GLM 5.3 and DeepSeek V4 Pro, a few points behind them on the coding benchmark they all report. Specs, price, benchmarks, hardware, and license, side by side."
author: "Yotta Labs"
date: 2026-10-09
categories: ["Inference"]
canonical: https://www.yottalabs.ai/post/mistral-large-4-vs-deepseek-v4-pro-vs-glm-5-3-vs-kimi-k3-2026
---

# Mistral Large 4 vs DeepSeek V4 Pro, GLM 5.3, and Kimi K3 (2026)

![](https://cdn.sanity.io/images/wy75wyma/production/d5ccaeb420cda467443f88cd6ab85de99aea13c1-1200x627.png)

*Mistral Large 4 puts a European lab in the trillion-parameter open-weight class. Here's how it lines up against the three it has to beat: DeepSeek V4 Pro, GLM 5.3, and Kimi K3.*

The biggest open-weight models of 2026 have come from Chinese labs: DeepSeek V4 Pro, Z.ai's GLM 5.3, and Moonshot's Kimi K3. Mistral joined them on October 6, 2026 with the public preview of Mistral Large 4, a 1.05 trillion-parameter model with weights promised by the end of the month. That makes it a field of four.

They're closer to each other than the headlines suggest. All four are sparse Mixture-of-Experts models with a 1M-token context window, all four publish or have promised weights, and three of the four are priced within a few cents of each other. The differences that decide which one you use are in the benchmark column, the license, and how much hardware the weights need.

## The short version

- **Price:** Mistral Large 4, GLM 5.3, and DeepSeek V4 Pro all sit between about $1 and $1.40 per million input tokens. Kimi K3 costs more than twice that on input and three to five times on output.
- **Coding benchmark:** on DeepSWE v1.1, each vendor's own run puts Kimi K3 at 67.5, GLM 5.3 at 66.9, DeepSeek V4 Pro at 62.7, and Mistral Large 4 at 61.7.
- **Self-hosting:** GLM 5.3 is the easiest of the released three, on one 8x H200 node. Mistral Large 4 could land in the same class if Mistral ships a 4-bit checkpoint, but its weights aren't out yet.
- **License:** DeepSeek V4 Pro is MIT. GLM 5.3 and Kimi K3 use custom licenses. Mistral hasn't announced one.
- **Where it's built:** Mistral trained and serves Large 4 in its own European datacenters, which is the reason some teams will pick it regardless of the other columns.

## Specs side by side

<!-- unsupported block: table -->

Mistral's own pages give two figures for active parameters, 49 billion in the announcement and 52B in the model docs, which is why the table says about 50B. Either way it matches DeepSeek V4 Pro almost exactly on compute per token while carrying about a third fewer total parameters.

On inputs: Mistral Large 4 and Kimi K3 both take images as well as text, and GLM 5.3 is text-only. If your workload includes screenshots, documents, or charts, that narrows the field before price does.

## Price

<!-- unsupported block: table -->

Mistral's list price is $1.36 in and $4.18 out. Its docs page currently shows $0.68 and $2.09 with the list prices struck through, so check which one is in effect before you budget.

What that looks like on a real workload. Take 10 million input tokens and 2 million output tokens a day, no caching:

<!-- unsupported block: table -->

Three of the four are in the same band, and Mistral Large 4 at list price sits between DeepSeek V4 Pro and GLM 5.3. Kimi K3 is the premium option, at close to three times the monthly bill of Mistral Large 4 or GLM 5.3. At the lower rate Mistral's docs page shows today, Large 4 would be the cheapest of the four on this workload at $329.40 a month.

## Benchmarks

DeepSWE v1.1 is the one benchmark all four labs report, so it's the only fair column.

<!-- unsupported block: table -->

Every figure is the vendor's own run at its own settings, so read it as a rough ordering and not a ranking. For reference, OpenAI reports 68.8 for [GPT-6 Sol](https://www.yottalabs.ai/post/gpt-6-sol-luna-pricing-vs-astra-open-models-2026) on the same test.

On that reading Mistral Large 4 is level with DeepSeek V4 Pro and about five points behind GLM 5.3 and Kimi K3 on coding. Mistral isn't pitching it as a coding leader. The numbers it leads with are in security: 93% on Cybench, and 82% on a vulnerability reproduction and patching test that Mistral says is the highest of any model. It also reports 28.3% on Terminal-Bench 4.0 and training across more than 160 languages, including every official EU language.

None of Mistral's figures have independent replication yet, and the model is still in preview. Mistral says it is continuing to refine it before the weights ship.

## Hardware, if you self-host

<!-- unsupported block: table -->

The Mistral row is arithmetic from the parameter count, not a tested configuration. The full working is in the [Mistral Large 4 hardware requirements](https://www.yottalabs.ai/post/mistral-large-4-hardware-requirements-gpu-memory-2026) post, and the per-model breakdowns are in the [GLM 5.3](https://www.yottalabs.ai/post/glm-5-3-hardware-requirements-gpu-memory-2026) and [Kimi K3](https://www.yottalabs.ai/post/kimi-k3-hardware-requirements-gpu-memory-2026) hardware posts.

The pattern across all four: active parameters set the cost per token, total parameters set the memory. Kimi K3 is the only one that needs more than a single node today. On Yotta, [8-GPU H200 and B300 nodes by the hour](https://www.yottalabs.ai/pricing) cover the single-node rows.

## License and where it runs

This is where the four separate most.

**DeepSeek V4 Pro** is MIT, the most permissive of the group. You can host it, fine-tune it, and ship it commercially without reading a custom license.

**GLM 5.3 and Kimi K3** each ship under their own license. Both are open-weight, and both need a read from your legal team before a commercial deployment.

**Mistral Large 4** has no license yet. Mistral Large 3 shipped under Apache 2.0, and Mistral describes Large 4 as open-weight and built for self-deployment on private cloud and on-premise, but the terms aren't published. Don't plan a commercial deployment around it until they are.

Then there's jurisdiction. Mistral trained Large 4 on its own infrastructure in Europe and serves the preview from the same datacenters. For teams with EU data-residency rules, or procurement policies that restrict Chinese-lab models, that can decide the question before price or benchmarks come into it.

## Which one to use

**Pick DeepSeek V4 Pro** if you want the lowest price and the cleanest license. It's the cheapest on a flat rate, MIT-licensed, and generally available since August.

**Pick GLM 5.3** if coding performance per dollar is the goal and your inputs are text. It's within a point of Kimi K3 on DeepSWE at less than half the input price, and it's the easiest of the released three to self-host.

**Pick Kimi K3** if you need the top of the benchmark column with image input and the price is justified by the work.

**Pick Mistral Large 4** if you need a European model, image input at mid-tier pricing, or its security strengths, and you can live with preview status until the weights and license land.

In practice most teams won't pick one. DeepSeek V4 Pro, GLM 5.3, and Kimi K3 are all on [Yotta AI Gateway](https://www.yottalabs.ai/ai-gateway) behind one OpenAI-compatible key, so testing them against each other is a model-string change, and Mistral Large 4 can sit in the same routing table through Mistral's API. Run your own traffic through each and let cost per result decide. The [Best Chinese LLMs roundup](https://www.yottalabs.ai/post/best-chinese-llm-models-2026-deepseek-qwen-glm-kimi-compared) and the [best open-source LLMs](https://www.yottalabs.ai/post/best-open-source-llms-2026) post cover the wider field.

## Frequently asked questions

**Is Mistral Large 4 better than DeepSeek V4 Pro?** On DeepSWE v1.1 they're about level: 61.7 for Mistral Large 4 and 62.7 for DeepSeek V4 Pro, each on the vendor's own run. Mistral Large 4 takes image input and is built and served in Europe. DeepSeek V4 Pro is cheaper on a flat rate and MIT-licensed.

**Is Mistral Large 4 cheaper than GLM 5.3?** Slightly, at list price: $1.36 in and $4.18 out against $1.40 and $4.40 for GLM 5.3 on Yotta AI Gateway. Mistral's docs page currently shows a lower rate of $0.68 and $2.09.

**How does Mistral Large 4 compare to Kimi K3?** Kimi K3 scores higher on DeepSWE v1.1, 67.5 against 61.7, and costs more than twice as much on input and more than three times on output. Kimi K3 is also much larger, 2.8T parameters against 1.05T, and needs a multi-node cluster to self-host.

**Which is the largest open-weight model?** Of these four, Kimi K3 at 2.8 trillion parameters, followed by DeepSeek V4 Pro at 1.6T, Mistral Large 4 at 1.05T, and GLM 5.3 at about 753B.

**Which of these models can I self-host today?** DeepSeek V4 Pro, GLM 5.3, and Kimi K3 all have public weights. Mistral Large 4's weights are due by the end of October 2026.

**Which has the most permissive license?** DeepSeek V4 Pro, under MIT. GLM 5.3 and Kimi K3 use custom licenses, and Mistral hasn't announced terms for Large 4.

**Is Mistral Large 4 open source?** Mistral calls it open-weight, with weights due by the end of October 2026. The license hasn't been announced.

## Bottom line

Mistral Large 4 doesn't beat the Chinese flagships on the coding benchmark they share, and it doesn't need to. It matches DeepSeek V4 Pro there, prices in the same band as GLM 5.3, takes image input, and is the only one of the four trained and served in Europe. For teams that couldn't use a Chinese-lab model, that makes it the first option of its size they can use.

For everyone else the choice is the usual one. DeepSeek V4 Pro on price and license, GLM 5.3 on coding per dollar, Kimi K3 at the top end, and Mistral Large 4 once the weights and license are out. The three that are released are on [Yotta AI Gateway](https://www.yottalabs.ai/ai-gateway), one API key away.
