Skip to main content
Browse Models

Alibaba Cloud

QwQ 32B Preview

Released

2024-11-27

Family

Qwen 2

Type

Foundation Model

Downloads

External Download

You are about to open a link to an external source. Verify the URL before continuing.

Download

4-bit GGUF (Q4_K_M)

GGUF · QwQ-32B-Preview-Q4_K_M.gguf

5-bit GGUF (Q5_K_M)

GGUF · QwQ-32B-Preview-Q5_K_M.gguf

6-bit GGUF (Q6_K)

GGUF · QwQ-32B-Preview-Q6_K.gguf

8-bit GGUF (Q8_0)

GGUF · QwQ-32B-Preview-Q8_0.gguf

Ollama Model (q4_K_M)

Ollama

Model Report

Overview

QwQ 32B Preview is an experimental large language model developed by the Qwen Team, with the goal of advancing reasoning capabilities in artificial intelligence. Pronounced similarly to "quill" (/kwju:/), the QwQ 32B Preview model is designed to approach challenging problems through a process of curiosity-driven, reflective analysis, particularly excelling in mathematical and coding tasks. Its development aims to enhance AI's ability to engage in deep, self-questioning thought, especially on problems at the frontier of knowledge. The primary technical advancements and evaluation results associated with QwQ 32B Preview are publicly documented on the QwQ 32B Preview blog post and the corresponding Qwen2 Technical Report.

Performance table for QwQ-32B-Preview and other models across major AI benchmarks

Figure 1. Benchmark table comparing QwQ 32B Preview with other large language models on GPQA, AIME, MATH-500, and LiveCodeBench. The table enables direct comparison of model performance in technical and reasoning domains.

Model Architecture

QwQ 32B Preview is classified as a causal language model, built on a transformer-based architecture designed for next-token prediction. The construction of the model features notable components widely adopted in natural language processing systems: Rotary Position Embedding (RoPE) for positional encoding, the SwiGLU activation function for increased representation power, RMSNorm for normalization, and attention QKV bias for improved attention mechanisms. The model was developed through an initial pretraining stage followed by post-training to further refine its capabilities, aligning with modern practices in large-scale language model training as described in the Qwen2 Technical Report.

QwQ 32B Preview shares its foundational architecture with Qwen2.5-32B, inheriting design choices and implementation structures suited for high-capacity reasoning and knowledge integration.

Parameters and Specifications

The QwQ 32B Preview model contains approximately 32.5 billion parameters in total, of which 31.0 billion are non-embedding parameters, underscoring its scale for general-purpose reasoning. The transformer backbone comprises 64 layers, with a grouped-query attention (GQA) head structure utilizing 40 query attention heads and 8 key/value heads. Its context window spans a maximum of 32,768 tokens, permitting the consideration of extended textual inputs and outputs within a single sequence. The model is stored and distributed in BF16 tensor formats to optimize memory efficiency. The total model size, as distributed via safetensors, is 32.8 billion parameters, according to the official documentation.

Performance and Benchmarking

QwQ 32B Preview has been evaluated on several specialized benchmarks measuring analytical, mathematical, and programming competencies. According to the QwQ 32B Preview blog, it attains a score of 65.2% on GPQA, a benchmark testing graduate-level scientific problem-solving. On the AIME benchmark, which measures mathematical reasoning across diverse topics such as algebra, counting, and geometry, QwQ 32B Preview scores 50.0%. On MATH-500, a challenging arithmetic and algebraic reasoning test, the model achieves 90.6%. In coding and real-world logic tests as found in LiveCodeBench, QwQ 32B Preview scores 50.0%. These results distinguish the model as proficient across multiple technical reasoning domains.

The benchmark table found in the overview summarizes and directly compares QwQ 32B Preview’s results with other prominent large language models, clarifying its strengths in comparison to contemporary systems. The table includes models such as OpenAI's o1-series, GPT-4o, Claude 3.5 Sonnet, and Qwen2.5-72B Instruct, illustrating the landscape of AI reasoning capabilities.

Limitations

As an early preview release, QwQ 32B Preview exhibits several notable limitations. The model may occasionally mix or switch languages unexpectedly within the same response, an artifact of its multilingual training data and generation process as noted in official Qwen documentation. Users have also observed that, in certain circumstances, the model may fall into recursive, circular reasoning patterns, resulting in unnecessarily lengthy explanations that fail to reach a definitive conclusion. In addition, while optimized for technical reasoning, the model currently shows limited capabilities in broader domains such as nuanced language understanding or strong common-sense reasoning. Safety features are considered preliminary; as such, robust usage guidelines and caution are recommended when deploying the model for general or public-facing tasks, as described in the release notes.

Applications and Research Context

QwQ 32B Preview is intended primarily for research use in domains where deep, step-by-step reasoning is required, such as advanced mathematics, algorithmic programming, and logical problem solving. Its design is informed by ongoing work in reflective reasoning, where the model seeks to analyze a problem, consider alternative approaches, and systematically check its own logic before providing a solution. This capacity is particularly visible when the model tackles tasks involving multi-step computation or equation manipulation, such as decomposing a number into its prime factors or systematically diagnosing errors in mathematical derivations. The Qwen Team situates this work within broader efforts in large language model research, including reflective process supervision and reinforcement learning enhanced by system feedback, as described in the QwQ 32B Preview blog post.

Release Information and Model Family

QwQ 32B Preview was officially announced on November 28, 2024, as documented in the official blog post. It is directly based on the Qwen2.5-32B base model, with further post-training designed to enhance its reasoning performance. Related models in the Qwen family include Qwen2.5-32B-Instruct, a finetuned variant focused on instruction following. The ongoing development of the Qwen model family aims to incrementally improve advanced reasoning, critique, and multi-step logic in language models, supporting an open research ecosystem. For sustained discussions and collaboration, the Qwen community is accessible via a Discord channel and the Qwen organization page on ModelScope.

Further Reading and Resources

For those interested in exploring QwQ 32B Preview and related Qwen models further, the following resources offer comprehensive technical details, usage instructions, and research background:

These resources provide additional context, guides for research applications, ongoing development updates, and opportunities for community interaction.

About Qwen 2: Qwen 2 (and 2.5) is a family of advanced AI models developed by Alibaba, designed to excel in various tasks including general language understanding, coding, and mathematics.

More in the Qwen 2 Family

Alibaba Cloud /

Qwen2.5 VL 3B

A 3B-parameter multimodal language model that processes images, videos, and text with capabilities for visual reasoning, document analysis, and computer interface interaction.
Alibaba Cloud /

Qwen2.5 VL 7B

A 7-billion parameter multimodal model capable of processing images, documents, and videos with native dynamic resolution support and structured output capabilities.
Alibaba Cloud /

Qwen2.5 VL 72B

A 72-billion parameter multimodal model that processes images, videos, and text for document analysis, object detection, and visual agent tasks.
Alibaba Cloud /

QwQ 32B

A 32.5-billion parameter transformer model optimized for mathematical reasoning, coding, and complex problem-solving through reinforcement learning training.
Alibaba Cloud /

Qwen 2.5 Math 1.5B

A 1.5 billion parameter specialized model focused on mathematical reasoning and problem-solving capabilities in English and Chinese.
Deepseek AI /

DeepSeek R1 Distill Qwen 1.5B

A 1.5 billion parameter language model created through distillation techniques, focusing on mathematical reasoning and chain-of-thought problem-solving capabilities.
Agentica /

DeepCoder 1.5B Preview

A 1.5B parameter code generation model fine-tuned with reinforcement learning and iterative context lengthening for long-context reasoning tasks.
Alibaba Cloud /

Qwen 2.5 Math 7B

A 7.62-billion parameter mathematical reasoning model supporting English and Chinese with chain-of-thought and tool-integrated reasoning capabilities.
Alibaba Cloud /

Qwen 2.5 Math PRM 7B

A 7B parameter process reward model that evaluates mathematical reasoning steps individually rather than just final answers, achieving 67.6% accuracy on mathematical benchmarks.
Deepseek AI /

DeepSeek R1 Distill Qwen 7B

A 7.62B parameter distilled language model based on Qwen2.5-Math-7B, trained via knowledge distillation for mathematical and logical reasoning tasks.
Alibaba Cloud /

Qwen 2.5 Math 72B

A 72.7 billion parameter language model specialized for solving mathematical problems in English and Chinese using chain-of-thought and tool-integrated reasoning.
Alibaba Cloud /

Qwen 2.5 Math PRM 72B

A 72.8 billion parameter process reward model that evaluates intermediate mathematical reasoning steps using consensus filtering and achieves 78.3% F1 on ProcessBench.
Alibaba Cloud /

Qwen 2.5 Coder 7B

A 7.61 billion parameter transformer model designed for code generation and reasoning across 92 programming languages with a 128,000-token context window.
Alibaba Cloud /

Qwen 2.5 Coder 32B

Qwen2.5-Coder builds on Qwen2.5 by training on 5.5 trillion additional tokens of code data including source code, text-code grounding data, and synthetic data. This leads to significant improvements in code-related tasks.
Alibaba Cloud /

Qwen 2.5 7B

A 7.61-billion parameter transformer-based language model with 128K token context length, trained on 18 trillion tokens supporting 29+ languages.
Alibaba Cloud /

Qwen2.5 7B 1M

A 7-billion parameter transformer model with 1-million token context capacity utilizing Dual Chunk Attention for long-range text processing.
Alibaba Cloud /

Qwen 2.5 14B

A 14.7 billion parameter transformer model supporting 29 languages with 128K context window, designed as foundation for fine-tuning applications.
Alibaba Cloud /

Qwen2.5 14B 1M

A 14.7B parameter transformer model with 1-million token context capacity, utilizing dual chunk attention and progressive training for extended sequence processing.
Deepseek AI /

DeepSeek R1 Distill Qwen 14B

A 14B dense language model distilled from a mixture-of-experts architecture, optimized for mathematical reasoning and code generation tasks.
Agentica /

DeepCoder 14B Preview

A 14-billion parameter code reasoning model fine-tuned using distributed reinforcement learning with long-context capabilities up to 64,000 tokens.
Deep Cogito /

Cogito V1 Preview 14B

A 14.8 billion parameter instruction-tuned language model trained using Iterated Distillation and Amplification with hybrid reasoning capabilities across 30+ languages.
Alibaba Cloud /

Qwen 2.5 32B

A 32.5 billion parameter multilingual transformer model trained on 18 trillion tokens with 128K context length supporting text generation, coding, and mathematical reasoning.
Deepseek AI /

DeepSeek R1 Distill Qwen 32B

A 32B parameter language model created through knowledge distillation, optimized for mathematical reasoning, code generation, and complex problem-solving tasks.
Deep Cogito /

Cogito V1 Preview 32B

A 32-billion parameter instruction-tuned model based on Qwen2.5 architecture featuring dual operational modes and iterated distillation alignment methodology.
Alibaba Cloud /

Qwen 2.5 72B

This update to Qwen 2, pretrained on 18 trillion tokens, shows impressive performance in benchmarks (MMLU 85+, HumanEval 85+, MATH 80+), outperforming models of similar size. Qwen 2.5 has multilingual support for over 29 languages and can understand and generate structured data.
Alibaba Cloud /

Qwen 2 7B

A 7.6 billion parameter multilingual decoder-only Transformer model designed for further post-training with 32,000-token context and strong coding capabilities.
Alibaba Cloud /

Qwen 2 72B

A 72-billion parameter transformer model supporting extended context lengths up to 128K tokens and demonstrating strong multilingual capabilities across diverse benchmarks.

More from Alibaba Cloud

Alibaba Cloud /

Qwen3 0.6B

A 0.6B parameter language model featuring dual thinking modes, multilingual capabilities, and 32K context length through strong-to-weak distillation training.
Alibaba Cloud /

Qwen3 1.7B

A 1.7 billion parameter multilingual transformer supporting dual-mode reasoning with step-by-step "thinking" and rapid "non-thinking" response capabilities.
Alibaba Cloud /

Qwen3 4B

A 4-billion parameter transformer model featuring dual reasoning modes, extensive multilingual training, and competitive performance across mathematical, coding, and logical reasoning benchmarks.
Alibaba Cloud /

Qwen3 8B

Dense 8.2 billion parameter transformer model featuring hybrid thinking capabilities with 32K token context and multilingual support across 119 languages.
Alibaba Cloud /

Qwen3 14B

A 14.8 billion parameter transformer model featuring hybrid thinking/non-thinking reasoning modes, 32K context length, and multilingual capabilities across 119 languages.
Alibaba Cloud /

Qwen3 32B

A 32.8 billion parameter language model featuring hybrid thinking modes for both rapid responses and step-by-step reasoning across multilingual tasks.
Alibaba Cloud /

Qwen3 30B A3B

Mixture-of-experts model with 30.5B total parameters, 3.3B activated per token, featuring hybrid reasoning modes and multilingual support across 119 languages.
Alibaba Cloud /

Qwen3 235B A22B

Large language model with Mixture-of-Experts architecture featuring dual operational modes for both rapid inference and complex reasoning tasks.
Alibaba Cloud /

Qwen 1.5 32B

Foundation 32B Qwen 1.5 model from Alibaba Cloud.
Alibaba Cloud /

Qwen 1.5 72B

Foundation 72B Qwen 1.5 model from Alibaba Cloud.