Skip to main content
Browse Models

Deepseek AI

DeepSeek R1 Distill Qwen 32B

Released

2025-01-20

Family

Qwen 2

Type

Fine-Tuned Model

Downloads

External Download

You are about to open a link to an external source. Verify the URL before continuing.

Download

4-bit GGUF (Q4_K_M)

GGUF · DeepSeek-R1-Distill-Qwen-32B-Q4_K_M.gguf

5-bit GGUF (Q5_K_M)

GGUF · DeepSeek-R1-Distill-Qwen-32B-Q5_K_M.gguf

6-bit GGUF (Q6_K)

GGUF · DeepSeek-R1-Distill-Qwen-32B-Q6_K.gguf

8-bit GGUF (Q8_0)

GGUF · DeepSeek-R1-Distill-Qwen-32B-Q8_0.gguf

Ollama Model (q4_K_M)

Ollama

Model Report

Overview

DeepSeek R1 Distill Qwen 32B is a generative artificial intelligence model developed by DeepSeek-AI. Positioned within the DeepSeek-R1 model series, this model emphasizes enhanced reasoning capabilities, harnessing advanced reinforcement learning (RL) and knowledge distillation strategies. Built upon the Qwen2.5-32B base model, DeepSeek R1 Distill Qwen 32B is designed to excel at mathematical reasoning, code generation, and general cognitive tasks. The distillation process transfers reasoning patterns from the larger DeepSeek-R1 teacher model, yielding a dense, efficient system suitable for broad research and application contexts.

Bar chart comparing performance of DeepSeek models and competitors on reasoning, code, and knowledge benchmarks

Figure 1. Benchmark comparison between DeepSeek-R1 models and peer systems across tasks such as AIME 2024, Codeforces, GPQA Diamond, and MATH-500, highlighting the model's relative accuracy and percentile performance.

Model Architecture and Training Process

DeepSeek R1 Distill Qwen 32B inherits its dense architecture from Qwen2.5-32B, further refined through a targeted distillation strategy. The teacher model, DeepSeek-R1, is a Mixture-of-Experts (MoE) language model based on the DeepSeek-V3-Base architecture. DeepSeek-R1 features 671 billion total parameters, with 37 billion activated per forward pass.

The development of DeepSeek-R1 and its distilled derivatives leverages a two-tiered RL methodology: an initial RL stage to establish complex reasoning behaviors, and a subsequent RL phase to align outputs with human preferences of helpfulness and harmlessness. Prior to these RL stages, supervised fine-tuning (SFT) is employed to "seed" basic reasoning and general language skills. The distillation step for models such as DeepSeek R1 Distill Qwen 32B uses approximately 800,000 curated samples generated by the DeepSeek-R1 teacher, fine-tuning the target model for advanced reasoning performance.

Data Sources and Training Methodology

Training data for DeepSeek R1 Distill Qwen 32B is generated via the DeepSeek-R1 model and comprises two principal categories. The first and largest portion consists of around 600,000 samples oriented toward reasoning-intensive challenges—such as mathematics, coding, science, and logic—curated using rule-based and generative reward mechanisms. The second portion includes approximately 200,000 general-purpose samples, covering writing, factual question answering, translation, and self-cognition, with a subset explicitly leveraging Chain-of-Thought (CoT) reasoning.

The upstream DeepSeek-R1 teacher model is trained via a staged process. It begins with DeepSeek-R1-Zero, which employs RL using Group Relative Policy Optimization (GRPO) without any initial SFT. This yields emergent reasoning capabilities but introduces challenges, such as repetition and mixed-language responses. The subsequent DeepSeek-R1 stage improves on this by pre-conditioning with a small volume of carefully-selected CoT data ("cold start"), enhancing readability and language consistency. The final distilled datasets, produced after RL convergence, are curated with rejection sampling and SFT to exclude undesirable attributes (e.g., mixed language or excessive length), ensuring high-quality training inputs for downstream distillation, as detailed in the DeepSeek-R1 research paper.

Performance and Benchmark Evaluation

DeepSeek R1 Distill Qwen 32B demonstrates robust performance on a diverse set of evaluation benchmarks in mathematical reasoning, code generation, and general knowledge tasks. On tasks like AIME 2024, MATH-500, and LiveCodeBench, it outperforms several competing dense models and closely follows or, in some cases, surpasses models that are considerably larger or trained with different methodologies.

Across key benchmarks, DeepSeek R1 Distill Qwen 32B has been shown to:

  • Achieve a Pass@1 score of 72.6 on AIME 2024, outperforming both QwQ-32B-Preview and o1-mini.
  • Reach a Pass@1 of 94.3 on MATH-500 and 57.2 on LiveCodeBench, reflecting strong mathematical and coding abilities.
  • Excel on GPQA Diamond and Codeforces, ranking among the top performers for general knowledge and programming competition proficiency.

The performance gains observed with DeepSeek R1 Distill Qwen 32B are attributed to its distillation from the large-scale RL-trained DeepSeek-R1 model, rather than relying solely on direct RL training at the smaller scale. This approach enables efficient transfer of the teacher model's reasoning strategies into a more compact and versatile system, as highlighted across multiple technical reports and benchmarks.

Applications and Use Cases

DeepSeek R1 Distill Qwen 32B is particularly well-suited for scenarios demanding intricate reasoning, robust problem-solving, and high-level cognitive tasks. Typical applications include advanced mathematics problem solving (as evaluated by AIME and MATH-500 benchmarks), code synthesis and debugging (as reflected in LiveCodeBench and Codeforces performance), and general knowledge reasoning (as assessed by GPQA Diamond).

Given its proficiency in Chain-of-Thought style tasks, the model finds further application in question answering, technical and scientific writing, summarization, and scenarios where explanations and stepwise logic are critical. The broader DeepSeek-R1 family, from which this model is derived, is also used for creative writing, editing, and tasks requiring extended context handling, as described in the DeepSeek-R1 research documentation.

Limitations

While DeepSeek R1 Distill Qwen 32B demonstrates substantial capabilities, certain limitations are documented. The model is optimized for Chinese and English, which can result in language mixing or reduced performance when queried in other languages. Its outputs are prompt-sensitive, and few-shot prompts may lead to diminished reasoning ability; zero-shot querying is generally recommended. The model performs best when instructions are provided directly in the user prompt, without using a separate system prompt.

In comparison to earlier models such as DeepSeek-V3, DeepSeek-R1 derivatives may be less capable in some complex structured tasks, including function calling or multi-turn dialogues. Additionally, in domain-specific software engineering benchmarks, improvements over previous generations are limited by the current scope of RL training data.

Licensing

DeepSeek R1 Distill Qwen 32B is released under the MIT License, which grants broad rights for use, modification, and distribution, including in commercial settings. The underlying Qwen2.5-32B base model is distributed under the Apache 2.0 License, and users should be mindful of corresponding terms and attribution requirements. Models distilled from Llama series inherit their respective Llama-3.1 or Llama-3.3 licenses.

Helpful Links

About Qwen 2: Qwen 2 (and 2.5) is a family of advanced AI models developed by Alibaba, designed to excel in various tasks including general language understanding, coding, and mathematics.

More in the Qwen 2 Family

Alibaba Cloud /

Qwen2.5 VL 3B

A 3B-parameter multimodal language model that processes images, videos, and text with capabilities for visual reasoning, document analysis, and computer interface interaction.
Alibaba Cloud /

Qwen2.5 VL 7B

A 7-billion parameter multimodal model capable of processing images, documents, and videos with native dynamic resolution support and structured output capabilities.
Alibaba Cloud /

Qwen2.5 VL 72B

A 72-billion parameter multimodal model that processes images, videos, and text for document analysis, object detection, and visual agent tasks.
Alibaba Cloud /

QwQ 32B Preview

Experimental research model focused on advancing AI reasoning capabilities by thinking through problems step by step, continuously questioning assumptions, and exploring different paths of thought. Impressive benchmark scores in math and coding tasks.
Alibaba Cloud /

QwQ 32B

A 32.5-billion parameter transformer model optimized for mathematical reasoning, coding, and complex problem-solving through reinforcement learning training.
Alibaba Cloud /

Qwen 2.5 Math 1.5B

A 1.5 billion parameter specialized model focused on mathematical reasoning and problem-solving capabilities in English and Chinese.
Deepseek AI /

DeepSeek R1 Distill Qwen 1.5B

A 1.5 billion parameter language model created through distillation techniques, focusing on mathematical reasoning and chain-of-thought problem-solving capabilities.
Agentica /

DeepCoder 1.5B Preview

A 1.5B parameter code generation model fine-tuned with reinforcement learning and iterative context lengthening for long-context reasoning tasks.
Alibaba Cloud /

Qwen 2.5 Math 7B

A 7.62-billion parameter mathematical reasoning model supporting English and Chinese with chain-of-thought and tool-integrated reasoning capabilities.
Alibaba Cloud /

Qwen 2.5 Math PRM 7B

A 7B parameter process reward model that evaluates mathematical reasoning steps individually rather than just final answers, achieving 67.6% accuracy on mathematical benchmarks.
Deepseek AI /

DeepSeek R1 Distill Qwen 7B

A 7.62B parameter distilled language model based on Qwen2.5-Math-7B, trained via knowledge distillation for mathematical and logical reasoning tasks.
Alibaba Cloud /

Qwen 2.5 Math 72B

A 72.7 billion parameter language model specialized for solving mathematical problems in English and Chinese using chain-of-thought and tool-integrated reasoning.
Alibaba Cloud /

Qwen 2.5 Math PRM 72B

A 72.8 billion parameter process reward model that evaluates intermediate mathematical reasoning steps using consensus filtering and achieves 78.3% F1 on ProcessBench.
Alibaba Cloud /

Qwen 2.5 Coder 7B

A 7.61 billion parameter transformer model designed for code generation and reasoning across 92 programming languages with a 128,000-token context window.
Alibaba Cloud /

Qwen 2.5 Coder 32B

Qwen2.5-Coder builds on Qwen2.5 by training on 5.5 trillion additional tokens of code data including source code, text-code grounding data, and synthetic data. This leads to significant improvements in code-related tasks.
Alibaba Cloud /

Qwen 2.5 7B

A 7.61-billion parameter transformer-based language model with 128K token context length, trained on 18 trillion tokens supporting 29+ languages.
Alibaba Cloud /

Qwen2.5 7B 1M

A 7-billion parameter transformer model with 1-million token context capacity utilizing Dual Chunk Attention for long-range text processing.
Alibaba Cloud /

Qwen 2.5 14B

A 14.7 billion parameter transformer model supporting 29 languages with 128K context window, designed as foundation for fine-tuning applications.
Alibaba Cloud /

Qwen2.5 14B 1M

A 14.7B parameter transformer model with 1-million token context capacity, utilizing dual chunk attention and progressive training for extended sequence processing.
Deepseek AI /

DeepSeek R1 Distill Qwen 14B

A 14B dense language model distilled from a mixture-of-experts architecture, optimized for mathematical reasoning and code generation tasks.
Agentica /

DeepCoder 14B Preview

A 14-billion parameter code reasoning model fine-tuned using distributed reinforcement learning with long-context capabilities up to 64,000 tokens.
Deep Cogito /

Cogito V1 Preview 14B

A 14.8 billion parameter instruction-tuned language model trained using Iterated Distillation and Amplification with hybrid reasoning capabilities across 30+ languages.
Alibaba Cloud /

Qwen 2.5 32B

A 32.5 billion parameter multilingual transformer model trained on 18 trillion tokens with 128K context length supporting text generation, coding, and mathematical reasoning.
Deep Cogito /

Cogito V1 Preview 32B

A 32-billion parameter instruction-tuned model based on Qwen2.5 architecture featuring dual operational modes and iterated distillation alignment methodology.
Alibaba Cloud /

Qwen 2.5 72B

This update to Qwen 2, pretrained on 18 trillion tokens, shows impressive performance in benchmarks (MMLU 85+, HumanEval 85+, MATH 80+), outperforming models of similar size. Qwen 2.5 has multilingual support for over 29 languages and can understand and generate structured data.
Alibaba Cloud /

Qwen 2 7B

A 7.6 billion parameter multilingual decoder-only Transformer model designed for further post-training with 32,000-token context and strong coding capabilities.
Alibaba Cloud /

Qwen 2 72B

A 72-billion parameter transformer model supporting extended context lengths up to 128K tokens and demonstrating strong multilingual capabilities across diverse benchmarks.

More from Deepseek AI

Deepseek AI /

DeepSeek R1 Distill Llama 8B

Distilled 8B-parameter model optimized for mathematical reasoning and code generation through knowledge transfer from larger reinforcement learning-trained teacher models.
Deepseek AI /

DeepSeek R1 Distill Llama 70B

A 70B parameter dense language model distilled from DeepSeek-R1 using Llama 3.3 architecture, optimized for mathematical and coding reasoning tasks.
Deepseek AI /

DeepSeek R1 (0528)

A 671B-parameter MoE model with 37B active parameters featuring enhanced reasoning capabilities through reinforcement learning and chain-of-thought training methodologies.
Deepseek AI /

DeepSeek R1

A 671B parameter Mixture-of-Experts model trained with reinforcement learning to enhance reasoning capabilities in mathematics, coding, and logical tasks.
Deepseek AI /

DeepSeek V3 (0324)

Large-scale MoE language model utilizing 671B parameters with 37B activated per token, featuring enhanced reasoning and multilingual capabilities.
Deepseek AI /

DeepSeek V3

A 671-billion parameter Mixture-of-Experts language model with 37 billion active parameters per token, featuring auxiliary-loss-free load balancing and FP8 mixed-precision training.
Deepseek AI /

DeepSeek VL2

A Mixture-of-Experts vision-language model series featuring dynamic image tiling, visual grounding capabilities, and efficient sparse computation across three parameter variants.
Deepseek AI /

DeepSeek VL2 Small

A 2.8B parameter mixture-of-experts vision-language model with dynamic tiling for multimodal understanding, OCR, and visual grounding tasks.
Deepseek AI /

DeepSeek VL2 Tiny

A compact Mixture-of-Experts vision-language model with 1.0B activated parameters supporting multimodal tasks including OCR, document analysis, and visual grounding.
Deepseek AI /

DeepSeek V2.5

A 236-billion parameter mixture-of-experts language model with multi-head latent attention, activating 21 billion parameters per token for bilingual text generation.
Deepseek AI /

DeepSeek V2

A 236-billion parameter Mixture-of-Experts language model that activates only 21 billion parameters per token for efficient multilingual text generation.
Deepseek AI /

DeepSeek Coder V2

Open-source Mixture-of-Experts model with 236B total parameters specialized for code generation, mathematical reasoning, and programming across 338 languages.
Deepseek AI /

DeepSeek Coder V2 Lite

A 16B parameter Mixture-of-Experts model designed for code generation, completion, and reasoning across 338 programming languages with 128K token context length.