Skip to main content
Browse Models

Alibaba Cloud

Qwen 1.5 32B

Released

2024-01-22

Family

Qwen 1

Type

Foundation Model

Downloads

External Download

You are about to open a link to an external source. Verify the URL before continuing.

Download

Chat model, 4-bit GGUF (Q4_K_M)

GGUF · Qwen1.5-32B-Chat-Q4_K_M.gguf

Chat model, 5-bit GGUF (Q5_K_M)

GGUF · Qwen1.5-32B-Chat-Q5_K_M.gguf

Chat model, 6-bit GGUF (Q6_K)

GGUF · Qwen1.5-32B-Chat-Q6_K.gguf

Chat model, 8-bit GGUF (Q8_0)

GGUF · Qwen1.5-32B-Chat-Q8_0.gguf

Ollama Model (q4_K_M)

Ollama

Model Report

Overview

Qwen1.5-32B is a large-scale generative language model in the Qwen1.5 series, developed by the Qwen Team and released in February 2024. This model, part of the Qwen lineup, incorporates designs for enhanced alignment, multilingual capabilities, long-context processing, and aims for improved performance across a range of natural language processing tasks.

Qwen1.5 timeline and branding

Figure 1. Timeline graphic highlighting releases in the Qwen series, with a focus on the launch of Qwen1.5.

Model Architecture and Training

The Qwen1.5-32B model is built as part of a modular series covering a spectrum of model sizes, from 0.5B to 110B parameters, alongside a mixture-of-experts (MoE) variant. All models in this series are unified by support for an extended context window of up to 32,768 tokens, which enhances their ability to process and retain information over lengthy documents or complex, multi-turn conversations. The architecture is accessible via the Hugging Face Transformers library (version 4.37.0 and above), allowing integration without modification or reliance on custom execution environments, as described in the official Qwen1.5 release documentation.

Alignment to human preferences is achieved using reinforcement learning techniques such as Direct Policy Optimization (DPO) and Proximal Policy Optimization (PPO), which are incorporated during the fine-tuning stages of development. These methods refine the model's tendency to follow human instructions and produce responses that are helpful, accurate, and contextually appropriate, as detailed in the Qwen team's technical blog post. While the specific pretraining datasets employed remain undisclosed, the focus on diverse data and multilingual resources is stated as a priority for broad generalization.

The Qwen1.5-32B model is published in multiple quantized formats, including Int4/Int8 GPTQ, AWQ, and GGUF, enabling efficient inference and adaptation to a range of deployment scenarios.

Performance and Benchmark Evaluation

Qwen1.5-32B demonstrates competitive results across numerous language understanding, reasoning, and coding benchmarks. The base model’s performance includes an MMLU score of 73.4, with scores indicating knowledge and reasoning competencies; a C-Eval score of 83.5; GSM8K at 77.4 for mathematical reasoning; MATH at 36.1; HumanEval at 37.2 for programming tasks; MBPP at 49.4; BBH at 66.8; and CMMLU at 82.3. These results, published in the Qwen1.5 performance overview, indicate its performance relative to peer language models of similar scale.

Although the Qwen1.5-32B model's multilingual and tool-use specific results are less extensively documented than those of larger siblings such as Qwen-1_5-72B, the series as a whole exhibits capabilities in language understanding, translation, and code-related tasks. The model family has been evaluated using benchmarks such as MT-Bench and AlpacaEval v2 for instruction following, L-Eval for long context handling, RGB for retrieval-augmented generation, and additional assessments for code interpretation and tool use.

Scatterplot of model performance on MT-Bench and AlpacaEval benchmarks

Figure 2. Scatter plot comparing alignment with human preferences across Qwen1.5 and other prominent language models on MT-Bench and AlpacaEval 2.0.

Multilinguality and Context Handling

Qwen1.5-32B includes support for multiple languages spanning Europe, East Asia, and Southeast Asia, with demonstrable performance across exams, translation, comprehension, and mathematical reasoning tasks. The design emphasizes uniform multilingual evaluation and large context capacity, allowing applications in translation, cross-lingual information retrieval, and multilingual dialogue.

Illustrative results in the broader Qwen1.5 family indicate score parity or, in some cases, an advantage compared to other leading models in multilingual benchmarks, as highlighted in bar charts comparing Qwen-1_5-72B-Chat with models like GPT-3.5 across languages such as Arabic, French, Vietnamese, Korean, Japanese, and others in a performance comparison.

Multilingual performance chart for Qwen1.5 and GPT-3.5

Figure 3. Bar chart comparing multilingual evaluation scores between Qwen1.5-72B-Chat and GPT-3.5 across various languages.

The consistent support for context lengths up to 32,768 tokens increases usability in scenarios requiring rich document understanding and handling of complex multi-step instructions without premature truncation or loss of coherence.

Integration, Alignment, and System Connectivity

Beyond core language modeling, Qwen1.5-32B is developed with capabilities for integration with external systems. Features include support for retrieval-augmented generation (RAG), enabling the use of up-to-date or private knowledge via connected external databases or APIs. The model series is also designed for API-based interaction, function/tool calling, and interaction with agent-oriented frameworks.

Alignment with human intent and safety guidelines is addressed via the aforementioned DPO and PPO training regimes. These steps contribute to instruction-following reliability and human-aligned outputs, validated through comprehensive crowd-sourcing and evaluation results in alignment-oriented benchmarks, as described in the Qwen1.5 technical overview.

Comparative Position and Limitations

Within the Qwen1.5 series, the 32B model occupies a middle-size tier, offering a balance between scale and resource efficiency. Larger siblings, such as Qwen-1_5-72B, are reported to outperform models like Llama-2-70B across diverse metrics and demonstrate high scores in alignment and multilingual evaluations. However, on code interpretation and visual reasoning tasks, the series is noted to trail behind frontier models such as GPT-4, particularly in tool-use and math-centric reasoning.

While the models achieve parity in many multi-turn, chat, and reasoning benchmarks, ongoing work is planned by the Qwen Team to bolster code reasoning and interpreter accuracy, as noted in the model’s official blog announcement.

Applications and Use Cases

Qwen1.5-32B finds application across a wide range of domains, including language comprehension and generation, coding assistance, complex reasoning, multilingual dialogue, retrieval-augmented workflows, and integration with agent frameworks capable of tool invocation. Fine-tuning support and configurability allow adaptation to specialized domains and research. Local inference in quantized formats and integration with open-source tools further extend practical accessibility.

Helpful Links

About Qwen 1: The Qwen 1 series (going up to Qwen 1.5), developed by Alibaba Cloud, comprises large language and multimodal models ranging from 0.5 to 72 billion parameters, designed for tasks such as text generation, image understanding, and conversation

More from Alibaba Cloud

Alibaba Cloud /

Qwen3 0.6B

A 0.6B parameter language model featuring dual thinking modes, multilingual capabilities, and 32K context length through strong-to-weak distillation training.
Alibaba Cloud /

Qwen3 1.7B

A 1.7 billion parameter multilingual transformer supporting dual-mode reasoning with step-by-step "thinking" and rapid "non-thinking" response capabilities.
Alibaba Cloud /

Qwen3 4B

A 4-billion parameter transformer model featuring dual reasoning modes, extensive multilingual training, and competitive performance across mathematical, coding, and logical reasoning benchmarks.
Alibaba Cloud /

Qwen3 8B

Dense 8.2 billion parameter transformer model featuring hybrid thinking capabilities with 32K token context and multilingual support across 119 languages.
Alibaba Cloud /

Qwen3 14B

A 14.8 billion parameter transformer model featuring hybrid thinking/non-thinking reasoning modes, 32K context length, and multilingual capabilities across 119 languages.
Alibaba Cloud /

Qwen3 32B

A 32.8 billion parameter language model featuring hybrid thinking modes for both rapid responses and step-by-step reasoning across multilingual tasks.
Alibaba Cloud /

Qwen3 30B A3B

Mixture-of-experts model with 30.5B total parameters, 3.3B activated per token, featuring hybrid reasoning modes and multilingual support across 119 languages.
Alibaba Cloud /

Qwen3 235B A22B

Large language model with Mixture-of-Experts architecture featuring dual operational modes for both rapid inference and complex reasoning tasks.
Alibaba Cloud /

Qwen2.5 VL 3B

A 3B-parameter multimodal language model that processes images, videos, and text with capabilities for visual reasoning, document analysis, and computer interface interaction.
Alibaba Cloud /

Qwen2.5 VL 7B

A 7-billion parameter multimodal model capable of processing images, documents, and videos with native dynamic resolution support and structured output capabilities.
Alibaba Cloud /

Qwen2.5 VL 72B

A 72-billion parameter multimodal model that processes images, videos, and text for document analysis, object detection, and visual agent tasks.
Alibaba Cloud /

QwQ 32B Preview

Experimental research model focused on advancing AI reasoning capabilities by thinking through problems step by step, continuously questioning assumptions, and exploring different paths of thought. Impressive benchmark scores in math and coding tasks.
Alibaba Cloud /

QwQ 32B

A 32.5-billion parameter transformer model optimized for mathematical reasoning, coding, and complex problem-solving through reinforcement learning training.
Alibaba Cloud /

Qwen 2.5 Math 1.5B

A 1.5 billion parameter specialized model focused on mathematical reasoning and problem-solving capabilities in English and Chinese.
Alibaba Cloud /

Qwen 2.5 Math 7B

A 7.62-billion parameter mathematical reasoning model supporting English and Chinese with chain-of-thought and tool-integrated reasoning capabilities.
Alibaba Cloud /

Qwen 2.5 Math PRM 7B

A 7B parameter process reward model that evaluates mathematical reasoning steps individually rather than just final answers, achieving 67.6% accuracy on mathematical benchmarks.
Alibaba Cloud /

Qwen 2.5 Math 72B

A 72.7 billion parameter language model specialized for solving mathematical problems in English and Chinese using chain-of-thought and tool-integrated reasoning.
Alibaba Cloud /

Qwen 2.5 Math PRM 72B

A 72.8 billion parameter process reward model that evaluates intermediate mathematical reasoning steps using consensus filtering and achieves 78.3% F1 on ProcessBench.
Alibaba Cloud /

Qwen 2.5 Coder 7B

A 7.61 billion parameter transformer model designed for code generation and reasoning across 92 programming languages with a 128,000-token context window.
Alibaba Cloud /

Qwen 2.5 Coder 32B

Qwen2.5-Coder builds on Qwen2.5 by training on 5.5 trillion additional tokens of code data including source code, text-code grounding data, and synthetic data. This leads to significant improvements in code-related tasks.
Alibaba Cloud /

Qwen 2.5 7B

A 7.61-billion parameter transformer-based language model with 128K token context length, trained on 18 trillion tokens supporting 29+ languages.
Alibaba Cloud /

Qwen2.5 7B 1M

A 7-billion parameter transformer model with 1-million token context capacity utilizing Dual Chunk Attention for long-range text processing.
Alibaba Cloud /

Qwen 2.5 14B

A 14.7 billion parameter transformer model supporting 29 languages with 128K context window, designed as foundation for fine-tuning applications.
Alibaba Cloud /

Qwen2.5 14B 1M

A 14.7B parameter transformer model with 1-million token context capacity, utilizing dual chunk attention and progressive training for extended sequence processing.
Alibaba Cloud /

Qwen 2.5 32B

A 32.5 billion parameter multilingual transformer model trained on 18 trillion tokens with 128K context length supporting text generation, coding, and mathematical reasoning.
Alibaba Cloud /

Qwen 2.5 72B

This update to Qwen 2, pretrained on 18 trillion tokens, shows impressive performance in benchmarks (MMLU 85+, HumanEval 85+, MATH 80+), outperforming models of similar size. Qwen 2.5 has multilingual support for over 29 languages and can understand and generate structured data.
Alibaba Cloud /

Qwen 2 7B

A 7.6 billion parameter multilingual decoder-only Transformer model designed for further post-training with 32,000-token context and strong coding capabilities.
Alibaba Cloud /

Qwen 2 72B

A 72-billion parameter transformer model supporting extended context lengths up to 128K tokens and demonstrating strong multilingual capabilities across diverse benchmarks.