Skip to main content
Browse Models

Alibaba Cloud

Qwen 2.5 32B

Released

2024-09-19

Family

Qwen 2

Type

Foundation Model

Downloads

External Download

You are about to open a link to an external source. Verify the URL before continuing.

Download

Instruct model, 4-bit GGUF (Q4_K_M)

GGUF · Qwen2.5-32B-Instruct-Q4_K_M.gguf

Instruct model, 5-bit GGUF (Q5_K_M)

GGUF · Qwen2.5-32B-Instruct-Q5_K_M.gguf

Instruct model, 6-bit GGUF (Q6_K)

GGUF · Qwen2.5-32B-Instruct-Q6_K.gguf

Instruct model, 8-bit GGUF (Q8_0)

GGUF · Qwen2.5-32B-Instruct-Q8_0.gguf

Instruct model, 16-bit GGUF (F16)

GGUF · Qwen2.5-32B-Instruct-f16.gguf

Ollama Model (q4_K_M)

Ollama

Model Report

Overview

Qwen2.5-32B is a large-scale, 32.5 billion parameter generative language model developed by the Qwen Team at Alibaba Group as part of the broader Qwen2.5 family. The model exemplifies advances in natural language processing, supporting broad multilingual capabilities, robust instruction following, and specialized performance in coding and mathematics. Qwen2.5-32B is positioned as a base model within its series and is primarily intended for continued development through post-training techniques. The Qwen2.5 series, released in September 2024, expanded upon earlier generations of Qwen models to support a variety of research and industrial use cases, maintaining an open-source ethos under the Apache 2.0 license for most variants.

Table of Qwen2.5 model specifications

Figure 1. Qwen2.5 model family specifications, including general and specialized versions such as Qwen2.5-Coder and Qwen2.5-Math, as detailed in the official documentation.

An introductory video overview of the Qwen2.5 project, explaining its main features and capabilities. · Source

Model Architecture and Technical Specifications

Qwen2.5-32B is built as a dense, decoder-only transformer, employing architectural innovations such as Rotary Position Embeddings (RoPE), SwiGLU (Swish Gated Linear Unit), and RMSNorm (Root Mean Square Normalization), along with a bias in the attention query-key-value computation. The model contains 64 layers, with 40 query heads and 8 key-value heads in a Grouped Query Attention setup, supporting efficient and scalable context processing. The total parameter count reaches 32.5 billion, with 31.0 billion non-embedding parameters.

The model supports a context window of up to 128,000 tokens, with internal specification supporting up to 131,072 tokens, and can produce outputs of up to 8,000 tokens per generation. This extensive context length significantly enhances the model’s capacity for document-level understanding and advanced reasoning tasks. Multilingual by design, Qwen2.5-32B recognizes and generates text in over 29 languages, including but not limited to Chinese, English, French, Spanish, Russian, Japanese, and Arabic, accompanied by strong instruction-following and translation abilities as outlined in the Qwen2.5 technical summary.

Further, the model introduces enhancements over previous Qwen2 series models, including improved knowledge coverage, better handling of structured data and long-form text, and improved resilience to diverse prompting scenarios. Its design accounts for downstream adaptation, enabling continued pretraining, supervised fine-tuning (SFT), reinforcement learning from human feedback (RLHF), and other post-training methodologies.

Training Regimen

The training of Qwen2.5-32B was conducted on a large-scale multilingual and multi-domain corpus comprising up to 18 trillion tokens, incorporating diverse sources to maximize knowledge acquisition and adaptability. The training process introduced refinements over prior series, emphasizing improvements in long context modeling, comprehension of structured outputs such as tables and JSON data, and stability across a variety of prompt types. These advances are documented in the Qwen2.5 blog.

Following pretraining, post-training algorithms were systematically integrated to further enhance performance in instruction adherence, complex task reasoning, and cross-lingual transfer. Although a comprehensive technical report is pending release as of September 2024, summary methodologies and preliminary findings are provided through project communications and the public model card.

Performance and Benchmarking

Qwen2.5-32B demonstrates competitive performance across a spectrum of established natural language understanding and generation benchmarks. Evaluations summarized in the official documentation indicate particularly strong results in multi-task learning (MMLU), coding (HumanEval), and mathematics (MATH) assessments, outpacing comparable or larger models such as Phi-3.5-MoE-Instruct and Gemma2-27B-IT on several key indicators.

Qwen2.5-32B Instruct Performance Table

Figure 2. Performance comparison of Qwen2.5-32B and Qwen2.5-14B against baseline models on benchmarks including MMLU, GPQA, HumanEval, and MATH, highlighting strengths in knowledge, code generation, and mathematical reasoning.

Published figures show that Qwen2.5-32B achieves MMLU scores surpassing 85, HumanEval coding assessments above 85, and mathematics test results exceeding 80, representing significant improvements over the previous Qwen2 generation, as detailed in the Qwen2.5 evaluation report.

Use Cases and Applications

Qwen2.5-32B is designed primarily for scientific research, custom model development, and integration as a large language foundation for a variety of downstream tasks. The model's capabilities extend across natural language understanding, code synthesis, mathematical reasoning, and handling of structurally rich outputs. Qwen models are utilized in natural language processing, multilingual translation, agent-based reasoning, tool integration, and as building blocks for domain-specific assistant systems.

The Qwen2.5-32B base model is intended for further post-training activities—such as SFT, RLHF, or continued pretraining—rather than as a direct conversational endpoint. Matched with robust tool use scaffolding, the model can be integrated into complex agent frameworks. Official resources and implementation guides are available via the Qwen documentation portal.

Release and Licensing

The Qwen2.5-32B model, released in September 2024, contributes to a lineage of models whose development includes the earlier Qwen1.5 and Qwen2 series. Most models within the Qwen2.5 family, including Qwen2.5-32B, are licensed under the permissive Apache 2.0 license, with licensing details and exceptions, such as for some 3B and 72B parameter variants, made explicit in respective repositories. This licensing approach supports open research and community-driven advancements while upholding transparent governance around use and distribution.

Limitations

Qwen2.5-32B, as a base language model, is not recommended for direct conversational deployment without further post-training. The technical report offering comprehensive methodological detail remained unreleased as of September 2024, with the best available insights provided through current blog updates and preliminary documentation. As a research-oriented model, its optimal performance and safety in open-ended conversational systems requires additional, task-specific fine-tuning and evaluation.

Helpful Links

About Qwen 2: Qwen 2 (and 2.5) is a family of advanced AI models developed by Alibaba, designed to excel in various tasks including general language understanding, coding, and mathematics.

More in the Qwen 2 Family

Alibaba Cloud /

Qwen2.5 VL 3B

A 3B-parameter multimodal language model that processes images, videos, and text with capabilities for visual reasoning, document analysis, and computer interface interaction.
Alibaba Cloud /

Qwen2.5 VL 7B

A 7-billion parameter multimodal model capable of processing images, documents, and videos with native dynamic resolution support and structured output capabilities.
Alibaba Cloud /

Qwen2.5 VL 72B

A 72-billion parameter multimodal model that processes images, videos, and text for document analysis, object detection, and visual agent tasks.
Alibaba Cloud /

QwQ 32B Preview

Experimental research model focused on advancing AI reasoning capabilities by thinking through problems step by step, continuously questioning assumptions, and exploring different paths of thought. Impressive benchmark scores in math and coding tasks.
Alibaba Cloud /

QwQ 32B

A 32.5-billion parameter transformer model optimized for mathematical reasoning, coding, and complex problem-solving through reinforcement learning training.
Alibaba Cloud /

Qwen 2.5 Math 1.5B

A 1.5 billion parameter specialized model focused on mathematical reasoning and problem-solving capabilities in English and Chinese.
Deepseek AI /

DeepSeek R1 Distill Qwen 1.5B

A 1.5 billion parameter language model created through distillation techniques, focusing on mathematical reasoning and chain-of-thought problem-solving capabilities.
Agentica /

DeepCoder 1.5B Preview

A 1.5B parameter code generation model fine-tuned with reinforcement learning and iterative context lengthening for long-context reasoning tasks.
Alibaba Cloud /

Qwen 2.5 Math 7B

A 7.62-billion parameter mathematical reasoning model supporting English and Chinese with chain-of-thought and tool-integrated reasoning capabilities.
Alibaba Cloud /

Qwen 2.5 Math PRM 7B

A 7B parameter process reward model that evaluates mathematical reasoning steps individually rather than just final answers, achieving 67.6% accuracy on mathematical benchmarks.
Deepseek AI /

DeepSeek R1 Distill Qwen 7B

A 7.62B parameter distilled language model based on Qwen2.5-Math-7B, trained via knowledge distillation for mathematical and logical reasoning tasks.
Alibaba Cloud /

Qwen 2.5 Math 72B

A 72.7 billion parameter language model specialized for solving mathematical problems in English and Chinese using chain-of-thought and tool-integrated reasoning.
Alibaba Cloud /

Qwen 2.5 Math PRM 72B

A 72.8 billion parameter process reward model that evaluates intermediate mathematical reasoning steps using consensus filtering and achieves 78.3% F1 on ProcessBench.
Alibaba Cloud /

Qwen 2.5 Coder 7B

A 7.61 billion parameter transformer model designed for code generation and reasoning across 92 programming languages with a 128,000-token context window.
Alibaba Cloud /

Qwen 2.5 Coder 32B

Qwen2.5-Coder builds on Qwen2.5 by training on 5.5 trillion additional tokens of code data including source code, text-code grounding data, and synthetic data. This leads to significant improvements in code-related tasks.
Alibaba Cloud /

Qwen 2.5 7B

A 7.61-billion parameter transformer-based language model with 128K token context length, trained on 18 trillion tokens supporting 29+ languages.
Alibaba Cloud /

Qwen2.5 7B 1M

A 7-billion parameter transformer model with 1-million token context capacity utilizing Dual Chunk Attention for long-range text processing.
Alibaba Cloud /

Qwen 2.5 14B

A 14.7 billion parameter transformer model supporting 29 languages with 128K context window, designed as foundation for fine-tuning applications.
Alibaba Cloud /

Qwen2.5 14B 1M

A 14.7B parameter transformer model with 1-million token context capacity, utilizing dual chunk attention and progressive training for extended sequence processing.
Deepseek AI /

DeepSeek R1 Distill Qwen 14B

A 14B dense language model distilled from a mixture-of-experts architecture, optimized for mathematical reasoning and code generation tasks.
Agentica /

DeepCoder 14B Preview

A 14-billion parameter code reasoning model fine-tuned using distributed reinforcement learning with long-context capabilities up to 64,000 tokens.
Deep Cogito /

Cogito V1 Preview 14B

A 14.8 billion parameter instruction-tuned language model trained using Iterated Distillation and Amplification with hybrid reasoning capabilities across 30+ languages.
Deepseek AI /

DeepSeek R1 Distill Qwen 32B

A 32B parameter language model created through knowledge distillation, optimized for mathematical reasoning, code generation, and complex problem-solving tasks.
Deep Cogito /

Cogito V1 Preview 32B

A 32-billion parameter instruction-tuned model based on Qwen2.5 architecture featuring dual operational modes and iterated distillation alignment methodology.
Alibaba Cloud /

Qwen 2.5 72B

This update to Qwen 2, pretrained on 18 trillion tokens, shows impressive performance in benchmarks (MMLU 85+, HumanEval 85+, MATH 80+), outperforming models of similar size. Qwen 2.5 has multilingual support for over 29 languages and can understand and generate structured data.
Alibaba Cloud /

Qwen 2 7B

A 7.6 billion parameter multilingual decoder-only Transformer model designed for further post-training with 32,000-token context and strong coding capabilities.
Alibaba Cloud /

Qwen 2 72B

A 72-billion parameter transformer model supporting extended context lengths up to 128K tokens and demonstrating strong multilingual capabilities across diverse benchmarks.

More from Alibaba Cloud

Alibaba Cloud /

Qwen3 0.6B

A 0.6B parameter language model featuring dual thinking modes, multilingual capabilities, and 32K context length through strong-to-weak distillation training.
Alibaba Cloud /

Qwen3 1.7B

A 1.7 billion parameter multilingual transformer supporting dual-mode reasoning with step-by-step "thinking" and rapid "non-thinking" response capabilities.
Alibaba Cloud /

Qwen3 4B

A 4-billion parameter transformer model featuring dual reasoning modes, extensive multilingual training, and competitive performance across mathematical, coding, and logical reasoning benchmarks.
Alibaba Cloud /

Qwen3 8B

Dense 8.2 billion parameter transformer model featuring hybrid thinking capabilities with 32K token context and multilingual support across 119 languages.
Alibaba Cloud /

Qwen3 14B

A 14.8 billion parameter transformer model featuring hybrid thinking/non-thinking reasoning modes, 32K context length, and multilingual capabilities across 119 languages.
Alibaba Cloud /

Qwen3 32B

A 32.8 billion parameter language model featuring hybrid thinking modes for both rapid responses and step-by-step reasoning across multilingual tasks.
Alibaba Cloud /

Qwen3 30B A3B

Mixture-of-experts model with 30.5B total parameters, 3.3B activated per token, featuring hybrid reasoning modes and multilingual support across 119 languages.
Alibaba Cloud /

Qwen3 235B A22B

Large language model with Mixture-of-Experts architecture featuring dual operational modes for both rapid inference and complex reasoning tasks.
Alibaba Cloud /

Qwen 1.5 32B

Foundation 32B Qwen 1.5 model from Alibaba Cloud.
Alibaba Cloud /

Qwen 1.5 72B

Foundation 72B Qwen 1.5 model from Alibaba Cloud.