Skip to main content
Browse Models

Alibaba Cloud

Qwen 2.5 72B

Released

2024-09-19

Family

Qwen 2

Type

Foundation Model

Downloads

External Download

You are about to open a link to an external source. Verify the URL before continuing.

Download

Instruct model, 4-bit GGUF (Q4_K_M)

GGUF · Qwen2.5-72B-Instruct-Q4_K_M.gguf

Ollama Model (q4_K_M)

Ollama

Model Report

Overview

Qwen2.5-72B is a large language model in the Qwen series developed by Alibaba Group’s Qwen Team. Released in September 2024, it represents an extensively scaled, dense, decoder-only transformer architecture designed to facilitate a broad range of natural language processing tasks. Qwen2.5-72B supports a significant context window, robust multilingual capabilities, and improved handling of structured data. Its architecture and training regime allow it to deliver strong performance across established linguistic, coding, and mathematical benchmarks, positioning it as a versatile component within the open-source AI ecosystem.

This official video presents an overview of the Qwen2.5 series, highlighting its principal capabilities and improvements within the Qwen model family. · Source

Model Architecture and Technical Specifications

Qwen2.5-72B is built on a causal transformer framework, integrating technologies such as Rotary Position Embeddings (RoPE), SwiGLU activation functions, RMSNorm normalization, and attention mechanism enhancements including QKV bias. The model features 72.7 billion parameters, of which 70 billion are allocated exclusively to non-embedding roles, reflecting a focus on scalable inference and generation performance. This architecture supports a maximum context length of 128,000 tokens and provides output sequences up to 8,000 tokens in length, accommodating both brief and extended conversational or document-based tasks.

Qwen2.5 Specification Table

Figure 1. Specification table for Qwen2.5, listing model sizes, parameters, layer counts, head configuration, context lengths, generation lengths, and available specialized variants, including Qwen2.5-Coder and Qwen2.5-Math.

Qwen2.5-72B is licensed under the Apache 2.0 license, encouraging open research and development initiatives. The model is provided as a base version, meaning it is optimally used as a foundation for further post-training processes such as supervised fine-tuning or reinforcement learning from human feedback.

Training Regimen and Model Capabilities

Qwen2.5 models are pretrained on a massive corpus comprising up to 18 trillion tokens, encompassing diverse textual domains and over 29 languages. This extensive pretraining confers a pronounced aptitude for multilingual tasks, allowing the model to process and generate text in languages including Chinese, English, French, Spanish, German, Russian, Japanese, Korean, Arabic, Vietnamese, and Thai. The model architecture exhibits significant improvements in structured data comprehension, particularly for extracting and generating table-based and JSON-formatted content, facilitating robust interactions with structured enterprise and scientific datasets.

A principal focus of Qwen2.5-72B is enhanced instruction following and robustness to system prompt diversity, which supports sophisticated role-play scenarios and advanced chatbot conditioning. The model also demonstrates strong utility as an AI agent, effectively integrating with external tools via programmatic interfaces. Although certain “thinking mode” and “general-purpose mode” distinctions are introduced in successor models (such as Qwen3), Qwen2.5-72B supports highly capable logical reasoning, coding, and mathematical operations through its foundation training and exposure to expert model insights, including Qwen2.5-Coder and Qwen2.5-Math.

Performance and Benchmarking

Qwen2.5-72B achieves high levels of benchmark performance relative to contemporary large language models. On knowledge-intensive benchmarks such as MMLU, it surpasses a score of 85, reflecting improvements in factual accuracy and general knowledge retrieval over its predecessor, Qwen2. For coding tasks, the model exceeds a HumanEval score of 85, evidencing enhanced performance in software engineering and code synthesis domains. Mathematical reasoning is similarly improved, with MATH benchmark scores above 80, demonstrating capacity for quantitative and logical inference.

Qwen2.5 Math Benchmark Performance

Figure 2. Scatter plot of Qwen2.5-Math model performance against parameter size, highlighting competitive accuracy on the MATH benchmark across all sizes, with Qwen2.5-72B-Instruct achieving top accuracy.

Qwen2.5-72B Instruct Model Benchmark Table

Figure 3. Benchmark table comparing Qwen2.5-72B Instruct with other leading language models, including Mistral-Large V2 and Llama-3.1-70B, across diverse evaluation tasks.

The base version of Qwen2.5-72B also compares closely to models with larger parameter counts, such as Llama-3-405B, in academic and general reasoning benchmarks. In instruction-tuned configurations, Qwen2.5-72B demonstrates competitive results with established open-source models such as Mistral-Large V2 and the Llama-3.1-70B series. These evaluations are publicly documented and cross-referenced in the Qwen2.5 blog release.

Qwen2.5-72B Base Model Performance Benchmarks

Figure 4. Data table showing Qwen2.5-72B’s base model performance on academic and reasoning benchmarks, directly compared to Qwen2-72B, Mixtral, Llama-3-70B, and Llama-3-405B.

Applications and Use Cases

Qwen2.5-72B’s capabilities make it well-suited for a spectrum of applications requiring deep natural language understanding, text generation, tool use, and AI agent behaviors. Its improvements in handling extensive context windows benefit document analysis, legal text processing, scientific literature review, and other large-scale text operations. Enhanced structured data handling supports business intelligence, data annotation, and interactive analytics through direct JSON or table output.

The model’s strong multilingual support extends its applicability to translation, cross-lingual summarization, and global conversational interfaces. Additionally, its specialized expert training enables robust performance in software development (via Qwen2.5-Coder) and academic mathematics (via Qwen2.5-Math), with specialized versions targeting each domain for improved performance.

Qwen2.5-72B is released as a base model, optimized for further post-training rather than direct deployment in conversational scenarios. Typical workflows involve supervised fine-tuning (SFT), reinforcement learning from human feedback (RLHF), or continued domain-specific pretraining to achieve optimal performance for target applications.

Model Family, Availability, and Evolution

Qwen2.5-72B is the largest member of the Qwen2.5 dense, decoder-only model family, which spans a parameter spectrum from 0.5B through 72B. Specialized derivatives within the series include Qwen2.5-Coder (optimized for programming tasks) and Qwen2.5-Math (optimized for mathematical and reasoning challenges), both available at different parameter scales. The Qwen2.5 series builds on earlier Qwen releases, and has since been superseded by the Qwen3 series, which introduces further architectural enhancements and a new naming scheme. All open-source Qwen models are published under the Apache 2.0 license.

Development milestones include the release of Qwen2.5 models in September 2024, with previous iterations and MoE-based variants tracing back through the Qwen1.5 and Qwen2 timelines.

Limitations

Although Qwen2.5-72B demonstrates substantial performance across benchmark suites, its base version is not recommended for direct conversational or end-user deployment without additional post-training. For optimal usage, post-training strategies such as instruction-tuning should be employed. While API-based models derived from Qwen (such as Qwen-Plus) demonstrate strong comparative results to leading large models, they may underperform relative to select proprietary models on certain benchmarks.

Helpful Links

About Qwen 2: Qwen 2 (and 2.5) is a family of advanced AI models developed by Alibaba, designed to excel in various tasks including general language understanding, coding, and mathematics.

More in the Qwen 2 Family

Alibaba Cloud /

Qwen2.5 VL 3B

A 3B-parameter multimodal language model that processes images, videos, and text with capabilities for visual reasoning, document analysis, and computer interface interaction.
Alibaba Cloud /

Qwen2.5 VL 7B

A 7-billion parameter multimodal model capable of processing images, documents, and videos with native dynamic resolution support and structured output capabilities.
Alibaba Cloud /

Qwen2.5 VL 72B

A 72-billion parameter multimodal model that processes images, videos, and text for document analysis, object detection, and visual agent tasks.
Alibaba Cloud /

QwQ 32B Preview

Experimental research model focused on advancing AI reasoning capabilities by thinking through problems step by step, continuously questioning assumptions, and exploring different paths of thought. Impressive benchmark scores in math and coding tasks.
Alibaba Cloud /

QwQ 32B

A 32.5-billion parameter transformer model optimized for mathematical reasoning, coding, and complex problem-solving through reinforcement learning training.
Alibaba Cloud /

Qwen 2.5 Math 1.5B

A 1.5 billion parameter specialized model focused on mathematical reasoning and problem-solving capabilities in English and Chinese.
Deepseek AI /

DeepSeek R1 Distill Qwen 1.5B

A 1.5 billion parameter language model created through distillation techniques, focusing on mathematical reasoning and chain-of-thought problem-solving capabilities.
Agentica /

DeepCoder 1.5B Preview

A 1.5B parameter code generation model fine-tuned with reinforcement learning and iterative context lengthening for long-context reasoning tasks.
Alibaba Cloud /

Qwen 2.5 Math 7B

A 7.62-billion parameter mathematical reasoning model supporting English and Chinese with chain-of-thought and tool-integrated reasoning capabilities.
Alibaba Cloud /

Qwen 2.5 Math PRM 7B

A 7B parameter process reward model that evaluates mathematical reasoning steps individually rather than just final answers, achieving 67.6% accuracy on mathematical benchmarks.
Deepseek AI /

DeepSeek R1 Distill Qwen 7B

A 7.62B parameter distilled language model based on Qwen2.5-Math-7B, trained via knowledge distillation for mathematical and logical reasoning tasks.
Alibaba Cloud /

Qwen 2.5 Math 72B

A 72.7 billion parameter language model specialized for solving mathematical problems in English and Chinese using chain-of-thought and tool-integrated reasoning.
Alibaba Cloud /

Qwen 2.5 Math PRM 72B

A 72.8 billion parameter process reward model that evaluates intermediate mathematical reasoning steps using consensus filtering and achieves 78.3% F1 on ProcessBench.
Alibaba Cloud /

Qwen 2.5 Coder 7B

A 7.61 billion parameter transformer model designed for code generation and reasoning across 92 programming languages with a 128,000-token context window.
Alibaba Cloud /

Qwen 2.5 Coder 32B

Qwen2.5-Coder builds on Qwen2.5 by training on 5.5 trillion additional tokens of code data including source code, text-code grounding data, and synthetic data. This leads to significant improvements in code-related tasks.
Alibaba Cloud /

Qwen 2.5 7B

A 7.61-billion parameter transformer-based language model with 128K token context length, trained on 18 trillion tokens supporting 29+ languages.
Alibaba Cloud /

Qwen2.5 7B 1M

A 7-billion parameter transformer model with 1-million token context capacity utilizing Dual Chunk Attention for long-range text processing.
Alibaba Cloud /

Qwen 2.5 14B

A 14.7 billion parameter transformer model supporting 29 languages with 128K context window, designed as foundation for fine-tuning applications.
Alibaba Cloud /

Qwen2.5 14B 1M

A 14.7B parameter transformer model with 1-million token context capacity, utilizing dual chunk attention and progressive training for extended sequence processing.
Deepseek AI /

DeepSeek R1 Distill Qwen 14B

A 14B dense language model distilled from a mixture-of-experts architecture, optimized for mathematical reasoning and code generation tasks.
Agentica /

DeepCoder 14B Preview

A 14-billion parameter code reasoning model fine-tuned using distributed reinforcement learning with long-context capabilities up to 64,000 tokens.
Deep Cogito /

Cogito V1 Preview 14B

A 14.8 billion parameter instruction-tuned language model trained using Iterated Distillation and Amplification with hybrid reasoning capabilities across 30+ languages.
Alibaba Cloud /

Qwen 2.5 32B

A 32.5 billion parameter multilingual transformer model trained on 18 trillion tokens with 128K context length supporting text generation, coding, and mathematical reasoning.
Deepseek AI /

DeepSeek R1 Distill Qwen 32B

A 32B parameter language model created through knowledge distillation, optimized for mathematical reasoning, code generation, and complex problem-solving tasks.
Deep Cogito /

Cogito V1 Preview 32B

A 32-billion parameter instruction-tuned model based on Qwen2.5 architecture featuring dual operational modes and iterated distillation alignment methodology.
Alibaba Cloud /

Qwen 2 7B

A 7.6 billion parameter multilingual decoder-only Transformer model designed for further post-training with 32,000-token context and strong coding capabilities.
Alibaba Cloud /

Qwen 2 72B

A 72-billion parameter transformer model supporting extended context lengths up to 128K tokens and demonstrating strong multilingual capabilities across diverse benchmarks.

More from Alibaba Cloud

Alibaba Cloud /

Qwen3 0.6B

A 0.6B parameter language model featuring dual thinking modes, multilingual capabilities, and 32K context length through strong-to-weak distillation training.
Alibaba Cloud /

Qwen3 1.7B

A 1.7 billion parameter multilingual transformer supporting dual-mode reasoning with step-by-step "thinking" and rapid "non-thinking" response capabilities.
Alibaba Cloud /

Qwen3 4B

A 4-billion parameter transformer model featuring dual reasoning modes, extensive multilingual training, and competitive performance across mathematical, coding, and logical reasoning benchmarks.
Alibaba Cloud /

Qwen3 8B

Dense 8.2 billion parameter transformer model featuring hybrid thinking capabilities with 32K token context and multilingual support across 119 languages.
Alibaba Cloud /

Qwen3 14B

A 14.8 billion parameter transformer model featuring hybrid thinking/non-thinking reasoning modes, 32K context length, and multilingual capabilities across 119 languages.
Alibaba Cloud /

Qwen3 32B

A 32.8 billion parameter language model featuring hybrid thinking modes for both rapid responses and step-by-step reasoning across multilingual tasks.
Alibaba Cloud /

Qwen3 30B A3B

Mixture-of-experts model with 30.5B total parameters, 3.3B activated per token, featuring hybrid reasoning modes and multilingual support across 119 languages.
Alibaba Cloud /

Qwen3 235B A22B

Large language model with Mixture-of-Experts architecture featuring dual operational modes for both rapid inference and complex reasoning tasks.
Alibaba Cloud /

Qwen 1.5 32B

Foundation 32B Qwen 1.5 model from Alibaba Cloud.
Alibaba Cloud /

Qwen 1.5 72B

Foundation 72B Qwen 1.5 model from Alibaba Cloud.