Skip to content
  • Models
  • Rankings
  • Ori
Sign Up
Sign Up
OpenRouterOpenRouter
© 2026 OpenRouter, Inc

Product

  • Chat
  • Rankings
  • Benchmarks
  • Apps
  • Discover
  • Models
  • Collections
  • Providers
  • Pricing
  • Business
  • Enterprise
  • Labs

Company

  • About
  • Blog
  • Careers
    Hiring
  • Privacy
  • Terms of Service
  • Trust Center
  • Support
  • Works With OR
  • Data
  • Brand

Developer

  • Documentation
  • API Reference
  • Developer Platform
  • Status

Connect

  • Discord
  • GitHub
  • LinkedIn
  • X
  • YouTube
Favicon for deepseek

DeepSeek: DeepSeek V3.2

deepseek/deepseek-v3.2

Model weights
Compare

DeepSeek-V3.2 is a large language model designed to harmonize high computational efficiency with strong reasoning and agentic tool-use performance. It introduces DeepSeek Sparse Attention (DSA), a fine-grained sparse attention mechanism that reduces training and inference cost while preserving quality in long-context scenarios. A scalable reinforcement learning post-training framework further improves reasoning, with reported performance in the GPT-5 class, and the model has demonstrated gold-medal results on the 2025 IMO and IOI. V3.2 also uses a large-scale agentic task synthesis pipeline to better integrate reasoning into tool-use settings, boosting compliance and generalization in interactive environments.

Users can control the reasoning behaviour with the reasoning enabled boolean. Learn more in our docsOpens in new tab

Modalities

In / Out Price

28% off

$0.2088 / $0.3096per 1M

Context

164K

Released

Dec 1, 2025

Compare
ProvidersPricingPerformanceUptimeBenchmarksAppsActivityFAQExplore

Providers

Different companies host the same model. OpenRouter routes your request to one of them based on the routing mode you pick — Balanced (price + speed), Nitro (fastest), or Exacto (highest tool-calling accuracy).

Pricing

The average price customers actually pay for this model, next to the prices providers post. Caching and discounts mean the price actually paid is often well below the listed one.

Performance

Throughput is how fast the model writes (tokens per second — higher is better). Latency is total round-trip time (lower is better). TTFT is time-to-first-token — how long before you see anything appear (lower is better).

Uptime

Uptime is the percentage of the past 3 days that at least one provider was responding to requests. Availability is the percentage of time that inference was successfully served. OpenRouter continuously monitors and uses the next-best provider when one returns an error.

Benchmarks

Scores on standardized evaluations. Higher percentages are better — and rank percentile shows where this model lands among all models on OpenRouter.

Benchmark score summary for DeepSeek: DeepSeek V3.2 (Artificial Analysis and Design Arena)
SourceBenchmarkScore
Artificial AnalysisDeepSeek V3.2 (Reasoning) Coding Index44.2
Artificial AnalysisDeepSeek V3.2 (Reasoning) GPQA Diamond84.0%
Artificial AnalysisDeepSeek V3.2 (Reasoning) HLE24.6%
Artificial AnalysisDeepSeek V3.2 (Reasoning) IFBench60.7%
Artificial AnalysisDeepSeek V3.2 (Reasoning) τ²-Bench Telecom90.6%
Artificial AnalysisDeepSeek V3.2 (Reasoning) AA-LCR73.3%
Artificial AnalysisDeepSeek V3.2 (Reasoning) GDPval-AA15.5%
Artificial AnalysisDeepSeek V3.2 (Reasoning) CritPt2.9%
Artificial AnalysisDeepSeek V3.2 (Reasoning) Terminal-Bench Hard35.6%
Artificial AnalysisDeepSeek V3.2 (Reasoning) AA-Omniscience Accuracy33.0%
Artificial AnalysisDeepSeek V3.2 (Reasoning) AA-Omniscience Non-Hallucination Rate17.3%
Design ArenaDeepSeek-V3.2 Models Arena 3D Elo1161
Design ArenaDeepSeek-V3.2 Models Arena Asciiart Elo1100
Design ArenaDeepSeek-V3.2 Models Arena Code Categories Elo1179
Design ArenaDeepSeek-V3.2 Models Arena Data Visualization Elo1175
Design ArenaDeepSeek-V3.2 Models Arena Game Development Elo1154
Design ArenaDeepSeek-V3.2 Models Arena SVG Elo1057
Design ArenaDeepSeek-V3.2 Models Arena UI Component Elo1164
Design ArenaDeepSeek-V3.2 Models Arena Website Elo1187

Apps

Public apps that send the most traffic to this model. Good signal for what real production workloads look like — and a hint at which use cases this model is best suited for.

Activity

Token volume and request traffic to this model over time.

Quick Start

Drop-in code to call this model. OpenRouter's API is OpenAI-compatible — most SDKs work by just swapping the base URL. The only thing that changes between models is the model slug below.

Explore more models

AI Model RankingsRanking
28% off
$0.29$0.2088$0.43$0.3096$0.03$0.0216----2.10s27 tps
99.95%
25% off
$0.286$0.2145$0.429$0.3218$0.0286$0.02145----1.43s25 tps
99.39%
$0.259$0.42$0.135----3.22s12 tps
96.42%
$0.26$0.38$0.13----1.63s19 tps
99.40%
$0.26$0.38$0.13----1.34s28 tps
99.74%
19% off
$0.33$0.2683$0.48$0.3902$0.16$0.1301----1.55s19 tps
99.19%
$0.269$0.40$0.1345----1.20s27 tps
99.84%
$0.28$0.42$0.028----1.09s44 tps
100.00%
$0.30$0.96$0.09----1.50s26 tps
96.09%
$0.3705$1.112$0.0741$0.4635$0.037050.96s40 tps
99.26%
$0.50$1.50$0.25----0.47s45 tps
99.98%
$0.56$1.68------2.44s28 tps
98.23%
$1.00$1.00$0.50----4.37s7 tps
99.53%
$3.00$4.50------27.53s0 tps
55.50%
$3.00$4.50------1.83s33 tps
89.15%

Throughput

45tok/s

P50, best across providers

Latency

0.47s

P50, best provider

AutoExacto Benchmarks
GPQA DiamondTAU-BenchBaidu Qianfan82.9%--DigitalOcean81.6%--StreamLake81.6%74.3%AtlasCloud81.9%73.2%SiliconFlow81.6%73.1%
+10 more providers
Uptime (3d)The model was reachable. Request routed to a provider.

100.00%

Availability (3d)The model returned inference from any provider. Errors and empty responses count against it.

99.75%

Availability over the last 3 days

Last 72 hours
Availability 99.75%
3 Days Ago2 Days AgoYesterdayNow

Availability over the last 24 hours

OpenRouter Availability
99.73%
Without Routing
96.28%

When an error occurs in an upstream provider, we can recover by routing to another healthy provider, if your request filters allow it. You can access per-provider uptime data programmatically through the Endpoints API. Learn more about our load balancing and customization options.

1.
Favicon for https://janitorai.com/
Janitor AI
Janitor AI is a chatbot platform where users create and chat with custom AI characters for interactive roleplay, storytelling, and immersive fiction.
261Btokens
2.
Favicon for https://shapes.inc/
Shapes Sandbox Gateway
General purpose social agents
46.3Btokens
3.
Favicon for https://www.mydreamcompanion.com/
My Dream Companion
new
46.3Btokens
4.
Favicon for https://swipelink.fr/
swipelink-smartlink
new
41.7Btokens
5.
Favicon for https://lorebary.com/
Sophia's LoreBary
Organize and enhance your roleplay creations.
30.5Btokens

Frequently asked questions

DeepSeek-V3.2 is a large language model designed to harmonize high computational efficiency with strong reasoning and agentic tool-use performance. It introduces DeepSeek Sparse Attention (DSA), a fine-grained sparse attention mechanism that reduces training and inference cost while preserving quality in long-context scenarios.

DeepSeek V3.2 costs $0.2088/M input tokens and $0.3096/M output tokens, with separate rates for Cache Read at $0.0216/M tokens.

DeepSeek V3.2 has a 163,840 token context window.

Yes. DeepSeek V3.2 accepts tools and tool_choice for function calling. It does not support response_format, so JSON output is not enforced.

DeepSeek V3.2 is served by 15 providers on OpenRouter: GMICloud, StreamLake, SiliconFlow, DeepInfra, AtlasCloud, Venice, NovitaAI, Baidu Qianfan and 7 more. Requests are routed to the best available provider, with automatic failover to the others, and you can pin or exclude providers with provider routing.

DeepSeek V3.2 was released on December 1, 2025.

More models from DeepSeek

DeepSeek Pro Latest

This model always redirects to the latest model in the DeepSeek Pro family.

Text1.0M context
DeepSeek Flash Latest

This model always redirects to the latest model in the DeepSeek Flash family.

Text1.0M context
DeepSeek V4.1 Flash

DeepSeek V4.1 Flash is a sparse mixture-of-experts model from DeepSeek, and the first built on the company's Causal Encoder-Decoder (CED) architecture. It activates 8B parameters on input and 16B on output from a 552B-parameter backbone, an asymmetric split that keeps per-token compute low relative to the model's total size. Image understanding is native to the architecture, with visual and text embeddings trained jointly from the start of pre-training rather than added afterward as in the earlier experimental V4 Flash Vision Exp.

It is suited for coding, terminal, and computer-use agents, along with long-horizon tasks that must run to completion across many steps and long-context analysis. Compressed KV caching cuts cache memory to roughly a quarter of the previous Flash generation, significantly reducing costs on agentic workloads. DeepSeek positions it as the cost-efficient tier of the V4.1 family and reports that it exceeds V4 Pro on performance, speed, and task completion time.

Text1.0M context$0.15 / $0.60
DeepSeek V4 Flash Vision Exp

DeepSeek V4 Flash Vision Exp is an experimental vision-enabled version of DeepSeek V4 Flash 0731 from DeepSeek, adding image understanding while matching the base model on text capabilities including agents, reasoning, and world knowledge. It is a sparse mixture-of-experts model with 13B active parameters out of 284B total.

It is suited for document and chart understanding, visual question answering, and multimodal agent workflows that interleave text and images.

Text1.0M context$0.2156 / $0.6468
DeepSeek V4 Flash Vision Exp

DeepSeek V4 Flash Vision Exp is an experimental vision-enabled version of DeepSeek V4 Flash 0731 from DeepSeek, adding image understanding while matching the base model on text capabilities including agents, reasoning, and world knowledge. It is a sparse mixture-of-experts model with 13B active parameters out of 284B total.

It is suited for document and chart understanding, visual question answering, and multimodal agent workflows that interleave text and images.

Text1.0M context$0.11 / $0.33
DeepSeek V4 Pro 0813

DeepSeek V4 Pro 0813 is a large-scale mixture-of-experts model from DeepSeek. This is the GA release of DeepSeek V4 Pro.

Text1.0M context$0.5795 / $1.738
DeepSeek V4 Pro 0813

DeepSeek V4 Pro 0813 is a large-scale mixture-of-experts model from DeepSeek. This is the GA release of DeepSeek V4 Pro.

Text1.0M context$0.66 / $1.98
DeepSeek V4 Flash Latest

This model always redirects to the latest model in the DeepSeek V4 Flash family.

Text1.0M context
DeepSeek V4 Flash 0731

DeepSeek V4 Flash 0731 is a sparse mixture-of-experts model from DeepSeek, with 13B active parameters out of 284B total. This re-post-trained revision is suited for coding, reasoning, and agent workflows. This is the GA release of DeepSeek V4 Flash.

Text1.3M context$0.04 / $0.10
DeepSeek V4 Flash 0731

DeepSeek V4 Flash 0731 is a sparse mixture-of-experts model from DeepSeek, with 13B active parameters out of 284B total. This re-post-trained revision is suited for coding, reasoning, and agent workflows. This is the GA release of DeepSeek V4 Flash.

Text1.0M context$0.11 / $0.33
DeepSeek V4 Pro 0423

DeepSeek V4 Pro is a large-scale Mixture-of-Experts model from DeepSeek with 1.6T total parameters and 49B activated parameters, supporting a 1M-token context window. It is designed for advanced reasoning, coding, and long-horizon agent workflows, with strong performance across knowledge, math, and software engineering benchmarks.

Built on the same architecture as DeepSeek V4 Flash, it introduces a hybrid attention system for efficient long-context processing. Reasoning efforts high and xhigh are supported; xhigh maps to max reasoning. It is well suited for complex workloads such as full-codebase analysis, multi-step automation, and large-scale information synthesis, where both capability and efficiency are critical

Text1.0M context$0.9481 / $1.896
DeepSeek V4 Flash 0423

DeepSeek V4 Flash is an efficiency-optimized Mixture-of-Experts model from DeepSeek with 284B total parameters and 13B activated parameters, supporting a 1M-token context window. It is designed for fast inference and high-throughput workloads, while maintaining strong reasoning and coding performance.

The model includes hybrid attention for efficient long-context processing. Reasoning efforts high and xhigh are supported; xhigh maps to max reasoning. It is well suited for applications such as coding assistants, chat systems, and agent workflows where responsiveness and cost efficiency are important.

Text1.0M context$0.05 / $0.14
DeepSeek V3.2 Speciale

DeepSeek-V3.2-Speciale is a high-compute variant of DeepSeek-V3.2 optimized for maximum reasoning and agentic performance. It builds on DeepSeek Sparse Attention (DSA) for efficient long-context processing, then scales post-training reinforcement learning to push capability beyond the base model. Reported evaluations place Speciale ahead of GPT-5 on difficult reasoning workloads, with proficiency comparable to Gemini-3.0-Pro, while retaining strong coding and tool-use reliability. Like V3.2, it benefits from a large-scale agentic task synthesis pipeline that improves compliance and generalization in interactive environments.

Text131K context
DeepSeek V3.2 Exp

DeepSeek-V3.2-Exp is an experimental large language model released by DeepSeek as an intermediate step between V3.1 and future architectures. It introduces DeepSeek Sparse Attention (DSA), a fine-grained sparse attention mechanism designed to improve training and inference efficiency in long-context scenarios while maintaining output quality. Users can control the reasoning behaviour with the reasoning enabled boolean. Learn more in our docs

The model was trained under conditions aligned with V3.1-Terminus to enable direct comparison. Benchmarking shows performance roughly on par with V3.1 across reasoning, coding, and agentic tool-use tasks, with minor tradeoffs and gains depending on the domain. This release focuses on validating architectural optimizations for extended context lengths rather than advancing raw task accuracy, making it primarily a research-oriented model for exploring efficient transformer designs.

Text164K context$0.27 / $0.41
DeepSeek V3.1 Terminus

DeepSeek-V3.1 Terminus is an update to DeepSeek V3.1 that maintains the model's original capabilities while addressing issues reported by users, including language consistency and agent capabilities, further optimizing the model's performance in coding and search agents. It is a large hybrid reasoning model (671B parameters, 37B active) that supports both thinking and non-thinking modes. It extends the DeepSeek-V3 base with a two-phase long-context training process, reaching up to 128K tokens, and uses FP8 microscaling for efficient inference. Users can control the reasoning behaviour with the reasoning enabled boolean. Learn more in our docs

The model improves tool use, code generation, and reasoning efficiency, achieving performance comparable to DeepSeek-R1 on difficult benchmarks while responding more quickly. It supports structured tool calling, code agents, and search agents, making it suitable for research, coding, and agentic workflows.

Text164K context$0.27 / $1
DeepSeek V3.1

DeepSeek-V3.1 is a large hybrid reasoning model (671B parameters, 37B active) that supports both thinking and non-thinking modes via prompt templates. It extends the DeepSeek-V3 base with a two-phase long-context training process, reaching up to 128K tokens, and uses FP8 microscaling for efficient inference. Users can control the reasoning behaviour with the reasoning enabled boolean. Learn more in our docs

The model improves tool use, code generation, and reasoning efficiency, achieving performance comparable to DeepSeek-R1 on difficult benchmarks while responding more quickly. It supports structured tool calling, code agents, and search agents, making it suitable for research, coding, and agentic workflows.

It succeeds the DeepSeek V3-0324 model and performs well on a variety of tasks.

Text164K context$0.25 / $0.95
DeepSeek V3.1 Base

This is a base model, trained only for raw next-token prediction. Unlike instruct/chat models, it has not been fine-tuned to follow user instructions. Prompts need to be written more like training text or examples rather than simple requests (e.g., “Translate the following sentence…” instead of just “Translate this”).

DeepSeek-V3.1 Base is a 671B parameter open Mixture-of-Experts (MoE) language model with 37B active parameters per forward pass and a context length of 128K tokens. Trained on 14.8T tokens using FP8 mixed precision, it achieves high training efficiency and stability, with strong performance across language, reasoning, math, and coding tasks.

Text164K context
R1 Distill Qwen 7B

DeepSeek-R1-Distill-Qwen-7B is a 7 billion parameter dense language model distilled from DeepSeek-R1, leveraging reinforcement learning-enhanced reasoning data generated by DeepSeek's larger models. The distillation process transfers advanced reasoning, math, and code capabilities into a smaller, more efficient model architecture based on Qwen2.5-Math-7B. This model demonstrates strong performance across mathematical benchmarks (92.8% pass@1 on MATH-500), coding tasks (Codeforces rating 1189), and general reasoning (49.1% pass@1 on GPQA Diamond), achieving competitive accuracy relative to larger models while maintaining smaller inference costs.

Text131K context
DeepSeek R1 0528 Qwen3 8B

DeepSeek-R1-0528 is a lightly upgraded release of DeepSeek R1 that taps more compute and smarter post-training tricks, pushing its reasoning and inference to the brink of flagship models like O3 and Gemini 2.5 Pro. It now tops math, programming, and logic leaderboards, showcasing a step-change in depth-of-thought. The distilled variant, DeepSeek-R1-0528-Qwen3-8B, transfers this chain-of-thought into an 8 B-parameter form, beating standard Qwen3 8B by +10 pp and tying the 235 B “thinking” giant on AIME 2024.

Text131K context
R1 0528

May 28th update to the original DeepSeek R1 Performance on par with OpenAI o1, but open-sourced and with fully open reasoning tokens. It's 671B parameters in size, with 37B active in an inference pass.

Fully open-source model.

Text164K context$0.50 / $2.15
DeepSeek Prover V2

DeepSeek Prover V2 is a 671B parameter model, speculated to be geared towards logic and mathematics. Likely an upgrade from DeepSeek-Prover-V1.5 Not much is known about the model yet, as DeepSeek released it on Hugging Face without an announcement or description.

Text164K context
DeepSeek V3 Base

Note that this is a base model mostly meant for testing, you need to provide detailed prompts for the model to return useful responses.

DeepSeek-V3 Base is a 671B parameter open Mixture-of-Experts (MoE) language model with 37B active parameters per forward pass and a context length of 128K tokens. Trained on 14.8T tokens using FP8 mixed precision, it achieves high training efficiency and stability, with strong performance across language, reasoning, math, and coding tasks.

DeepSeek-V3 Base is the pre-trained model behind DeepSeek V3

Text131K context
DeepSeek V3 0324

DeepSeek V3, a 685B-parameter, mixture-of-experts model, is the latest iteration of the flagship chat model family from the DeepSeek team.

It succeeds the DeepSeek V3 model and performs really well on a variety of tasks.

Text164K context$0.24 / $0.90
DeepSeek R1 Zero

DeepSeek-R1-Zero is a model trained via large-scale reinforcement learning (RL) without supervised fine-tuning (SFT) as a preliminary step. It's 671B parameters in size, with 37B active in an inference pass.

It demonstrates remarkable performance on reasoning. With RL, DeepSeek-R1-Zero naturally emerged with numerous powerful and interesting reasoning behaviors.

DeepSeek-R1-Zero encounters challenges such as endless repetition, poor readability, and language mixing. See DeepSeek R1 for the SFT model.

Text164K context
R1 Distill Llama 8B

DeepSeek R1 Distill Llama 8B is a distilled large language model based on Llama-3.1-8B-Instruct, using outputs from DeepSeek R1. The model combines advanced distillation techniques to achieve high performance across multiple benchmarks, including:

The model leverages fine-tuning from DeepSeek R1's outputs, enabling competitive performance comparable to larger frontier models.

Hugging Face:

Text
R1 Distill Qwen 1.5B

DeepSeek R1 Distill Qwen 1.5B is a distilled large language model based on Qwen 2.5 Math 1.5B, using outputs from DeepSeek R1. It's a very small and efficient model which outperforms GPT 4o 0513 on Math Benchmarks.

Other benchmark results include:

The model leverages fine-tuning from DeepSeek R1's outputs, enabling competitive performance comparable to larger frontier models.

Text131K context
R1 Distill Qwen 32B

DeepSeek R1 Distill Qwen 32B is a distilled large language model based on Qwen 2.5 32B, using outputs from DeepSeek R1. It outperforms OpenAI's o1-mini across various benchmarks, achieving new state-of-the-art results for dense models.\n\nOther benchmark results include:\n\n- AIME 2024 pass@1: 72.6\n- MATH-500 pass@1: 94.3\n- CodeForces Rating: 1691\n\nThe model leverages fine-tuning from DeepSeek R1's outputs, enabling competitive performance comparable to larger frontier models.

Text128K context
R1 Distill Qwen 14B

DeepSeek R1 Distill Qwen 14B is a distilled large language model based on Qwen 2.5 14B, using outputs from DeepSeek R1. It outperforms OpenAI's o1-mini across various benchmarks, achieving new state-of-the-art results for dense models.

Other benchmark results include:

The model leverages fine-tuning from DeepSeek R1's outputs, enabling competitive performance comparable to larger frontier models.

Text131K context
R1 Distill Llama 70B

DeepSeek R1 Distill Llama 70B is a distilled large language model based on Llama-3.3-70B-Instruct, using outputs from DeepSeek R1. The model combines advanced distillation techniques to achieve high performance across multiple benchmarks, including:

The model leverages fine-tuning from DeepSeek R1's outputs, enabling competitive performance comparable to larger frontier models.

Text8K context$0.80 / $0.80
R1

DeepSeek R1 is here: Performance on par with OpenAI o1, but open-sourced and with fully open reasoning tokens. It's 671B parameters in size, with 37B active in an inference pass.

Fully open-source model & technical report.

MIT licensed: Distill & commercialize freely!

Text64K context$0.70 / $2.50
DeepSeek V2.5

DeepSeek-V2.5 is an upgraded version that combines DeepSeek-V2-Chat and DeepSeek-Coder-V2-Instruct. The new model integrates the general and coding abilities of the two previous versions. For model details, please visit DeepSeek-V2 page for more information.

Text128K context