Skip to main content
Browse Models

Alibaba Cloud

Qwen 1.5 72B

Released

2024-01-23

Family

Qwen 1

Type

Foundation Model

Downloads

External Download

You are about to open a link to an external source. Verify the URL before continuing.

Download

Chat model, 4-bit GGUF (Q4_K_M)

GGUF · Qwen1.5-72B-Chat-Q4_K_M.gguf

Ollama Model (q4_K_M)

Ollama

Model Report

Overview

Qwen1.5-72B is a large-scale generative AI model and a central member of the Qwen1.5 series, developed by the Qwen Team and released in February 2024. This iteration emphasizes advanced alignment with human preferences, robust multilingual proficiency, and enhanced developer accessibility. Qwen1.5-72B has been benchmarked extensively against leading models, displaying consistent performance gains in language understanding, code generation, and external system integration.

Qwen1.5 release timeline and branding

Figure 1. Qwen1.5 in context: a timeline of the Qwen model development, highlighting the February 2024 release of Qwen1.5 alongside earlier iterations.

Model Architecture and Training Approaches

Qwen1.5-72B follows the transformer-based architecture typical of recent large language models, offering both base and chat-optimized variants across multiple parameter scales. The 72B-parameter version supports a context window of up to 32,768 tokens, with settings optionally extendable in configuration files for longer sequences, enabling flexible adaptation for tasks requiring extended memory.

Alignment was a central concern in model design; techniques such as Direct Policy Optimization (DPO) and Proximal Policy Optimization (PPO) were utilized to fine-tune instruction-following ability and to closely match model outputs to human preference. The open-source documentation does not detail the proprietary datasets used for pre-training, but references the inclusion of broad multilingual materials and the use of community open-source repositories for evaluation benchmarks. Quantized model formats, such as GPTQ (Int4/Int8), AWQ, and GGUF, are officially provided, facilitating efficient deployment and lower resource utilization for inference.

Multilingual and General Language Capabilities

Qwen1.5-72B demonstrates considerable strength in multilingual understanding, covering at least 12 languages from diverse regions including Europe, East Asia, and Southeast Asia. Performance analysis on representative tasks shows substantial gains over earlier models—and in several cases, over peers like Llama 2 70B and GPT-3.5—in translation, language reasoning, and code execution benchmarks.

Multilingual performance comparison

Figure 2. Qwen1.5-72B-Chat outperforms GPT-3.5 in multilingual language tasks across several languages including Arabic, French, Vietnamese, Korean, Japanese, Spanish, Indonesian, Portuguese, Russian, and Thai.

In tasks such as the MMLU benchmark (77.5), C-Eval (84.1), GSM8K (79.5), MATH (34.1), and HumanEval (41.5), Qwen1.5-72B displays broad competence across knowledge, reasoning, and coding domains. These benchmarks indicate not only linguistic fluency but also advanced capability in mathematics and programming, although there remains a performance gap in code interpretation when compared to models such as GPT-4.

Alignment with Human Preferences and Benchmark Performance

The Qwen1.5-72B-Chat model demonstrates high scores on benchmarks designed to measure alignment with human judgment and instruction-following quality. On MT-Bench the model achieves an average score of 8.61, with a 27.18 win rate on AlpacaEval 2.0.

Scatter plot of human preference alignment benchmarks

Figure 3. Qwen1.5-72B-Chat achieves competitive results on MT-Bench and AlpacaEval 2.0, as shown in this benchmark comparison plot—prompting evaluation against models like GPT-4-Turbo and Mistral Medium.

Performance is especially notable when compared to contemporaries such as Claude-2.1, GPT-3.5-Turbo, and Mixtral 8x7B-Instruct, and approaches the scores of Mistral Medium. However, the model continues to trail GPT-4-Turbo on several of these gold-standard alignment benchmarks.

Integration, External System Connectivity, and Application Domains

One of the distinctive features of Qwen1.5-72B is its capacity for external system integration, including Retrieval-Augmented Generation (RAG), API-based tool and function calling for AI agent use cases, and Python code interpretation. The model performs competitively on RAG benchmarks, particularly in tasks requiring information retrieval under noise or counterfactual scenarios, and shows effective agent execution in both English and Chinese as measured by T-Eval.

Extensive evaluations demonstrate high tool-use accuracy (above 90% for both selection and input categories) and strong, though not leading, performance in mathematical and visualization-based code execution. The LEval long-context benchmark further highlights the model's ability to maintain performance in tasks involving complex, extended documents or dialogues.

Qwen1.5-72B has practical applications in conversational systems, code assistants, translation engines, mathematical problem solving, and customizable AI agents—thanks to its multilingual fluency, contextual memory, and open-ended instruction following.

Limitations and Future Directions

Despite robust scores in multilingual and alignment assessments, Qwen1.5-72B exhibits certain limitations. On advanced code interpretation and mathematical reasoning tasks, the model’s performance does not yet match that of GPT-4, especially for complex programming or difficult visualization problems. The efficacy of context windows significantly longer than 32,768 tokens, enabled by configuration changes, may vary and is not uniformly validated.

Continued improvements are anticipated in subsequent iterations to address these areas, with particular focus on enhancing coding capabilities and ensuring stable performance in longer-context applications.

Licensing and Access

Qwen1.5 models, including the 72B parameter variant, are open-source and are distributed via platforms like Hugging Face and ModelScope. While model licensing is not explicitly detailed in all official documentation, their open-source status facilitates broad research, adaptation, and integration into downstream AI pipelines.

External Resources

About Qwen 1: The Qwen 1 series (going up to Qwen 1.5), developed by Alibaba Cloud, comprises large language and multimodal models ranging from 0.5 to 72 billion parameters, designed for tasks such as text generation, image understanding, and conversation

More from Alibaba Cloud

Alibaba Cloud /

Qwen3 0.6B

A 0.6B parameter language model featuring dual thinking modes, multilingual capabilities, and 32K context length through strong-to-weak distillation training.
Alibaba Cloud /

Qwen3 1.7B

A 1.7 billion parameter multilingual transformer supporting dual-mode reasoning with step-by-step "thinking" and rapid "non-thinking" response capabilities.
Alibaba Cloud /

Qwen3 4B

A 4-billion parameter transformer model featuring dual reasoning modes, extensive multilingual training, and competitive performance across mathematical, coding, and logical reasoning benchmarks.
Alibaba Cloud /

Qwen3 8B

Dense 8.2 billion parameter transformer model featuring hybrid thinking capabilities with 32K token context and multilingual support across 119 languages.
Alibaba Cloud /

Qwen3 14B

A 14.8 billion parameter transformer model featuring hybrid thinking/non-thinking reasoning modes, 32K context length, and multilingual capabilities across 119 languages.
Alibaba Cloud /

Qwen3 32B

A 32.8 billion parameter language model featuring hybrid thinking modes for both rapid responses and step-by-step reasoning across multilingual tasks.
Alibaba Cloud /

Qwen3 30B A3B

Mixture-of-experts model with 30.5B total parameters, 3.3B activated per token, featuring hybrid reasoning modes and multilingual support across 119 languages.
Alibaba Cloud /

Qwen3 235B A22B

Large language model with Mixture-of-Experts architecture featuring dual operational modes for both rapid inference and complex reasoning tasks.
Alibaba Cloud /

Qwen2.5 VL 3B

A 3B-parameter multimodal language model that processes images, videos, and text with capabilities for visual reasoning, document analysis, and computer interface interaction.
Alibaba Cloud /

Qwen2.5 VL 7B

A 7-billion parameter multimodal model capable of processing images, documents, and videos with native dynamic resolution support and structured output capabilities.
Alibaba Cloud /

Qwen2.5 VL 72B

A 72-billion parameter multimodal model that processes images, videos, and text for document analysis, object detection, and visual agent tasks.
Alibaba Cloud /

QwQ 32B Preview

Experimental research model focused on advancing AI reasoning capabilities by thinking through problems step by step, continuously questioning assumptions, and exploring different paths of thought. Impressive benchmark scores in math and coding tasks.
Alibaba Cloud /

QwQ 32B

A 32.5-billion parameter transformer model optimized for mathematical reasoning, coding, and complex problem-solving through reinforcement learning training.
Alibaba Cloud /

Qwen 2.5 Math 1.5B

A 1.5 billion parameter specialized model focused on mathematical reasoning and problem-solving capabilities in English and Chinese.
Alibaba Cloud /

Qwen 2.5 Math 7B

A 7.62-billion parameter mathematical reasoning model supporting English and Chinese with chain-of-thought and tool-integrated reasoning capabilities.
Alibaba Cloud /

Qwen 2.5 Math PRM 7B

A 7B parameter process reward model that evaluates mathematical reasoning steps individually rather than just final answers, achieving 67.6% accuracy on mathematical benchmarks.
Alibaba Cloud /

Qwen 2.5 Math 72B

A 72.7 billion parameter language model specialized for solving mathematical problems in English and Chinese using chain-of-thought and tool-integrated reasoning.
Alibaba Cloud /

Qwen 2.5 Math PRM 72B

A 72.8 billion parameter process reward model that evaluates intermediate mathematical reasoning steps using consensus filtering and achieves 78.3% F1 on ProcessBench.
Alibaba Cloud /

Qwen 2.5 Coder 7B

A 7.61 billion parameter transformer model designed for code generation and reasoning across 92 programming languages with a 128,000-token context window.
Alibaba Cloud /

Qwen 2.5 Coder 32B

Qwen2.5-Coder builds on Qwen2.5 by training on 5.5 trillion additional tokens of code data including source code, text-code grounding data, and synthetic data. This leads to significant improvements in code-related tasks.
Alibaba Cloud /

Qwen 2.5 7B

A 7.61-billion parameter transformer-based language model with 128K token context length, trained on 18 trillion tokens supporting 29+ languages.
Alibaba Cloud /

Qwen2.5 7B 1M

A 7-billion parameter transformer model with 1-million token context capacity utilizing Dual Chunk Attention for long-range text processing.
Alibaba Cloud /

Qwen 2.5 14B

A 14.7 billion parameter transformer model supporting 29 languages with 128K context window, designed as foundation for fine-tuning applications.
Alibaba Cloud /

Qwen2.5 14B 1M

A 14.7B parameter transformer model with 1-million token context capacity, utilizing dual chunk attention and progressive training for extended sequence processing.
Alibaba Cloud /

Qwen 2.5 32B

A 32.5 billion parameter multilingual transformer model trained on 18 trillion tokens with 128K context length supporting text generation, coding, and mathematical reasoning.
Alibaba Cloud /

Qwen 2.5 72B

This update to Qwen 2, pretrained on 18 trillion tokens, shows impressive performance in benchmarks (MMLU 85+, HumanEval 85+, MATH 80+), outperforming models of similar size. Qwen 2.5 has multilingual support for over 29 languages and can understand and generate structured data.
Alibaba Cloud /

Qwen 2 7B

A 7.6 billion parameter multilingual decoder-only Transformer model designed for further post-training with 32,000-token context and strong coding capabilities.
Alibaba Cloud /

Qwen 2 72B

A 72-billion parameter transformer model supporting extended context lengths up to 128K tokens and demonstrating strong multilingual capabilities across diverse benchmarks.