Alibaba Cloud
Qwen 1.5 72B
Downloads
Model Report
Overview
Qwen1.5-72B is a large-scale generative AI model and a central member of the Qwen1.5 series, developed by the Qwen Team and released in February 2024. This iteration emphasizes advanced alignment with human preferences, robust multilingual proficiency, and enhanced developer accessibility. Qwen1.5-72B has been benchmarked extensively against leading models, displaying consistent performance gains in language understanding, code generation, and external system integration.

Figure 1. Qwen1.5 in context: a timeline of the Qwen model development, highlighting the February 2024 release of Qwen1.5 alongside earlier iterations.
Model Architecture and Training Approaches
Qwen1.5-72B follows the transformer-based architecture typical of recent large language models, offering both base and chat-optimized variants across multiple parameter scales. The 72B-parameter version supports a context window of up to 32,768 tokens, with settings optionally extendable in configuration files for longer sequences, enabling flexible adaptation for tasks requiring extended memory.
Alignment was a central concern in model design; techniques such as Direct Policy Optimization (DPO) and Proximal Policy Optimization (PPO) were utilized to fine-tune instruction-following ability and to closely match model outputs to human preference. The open-source documentation does not detail the proprietary datasets used for pre-training, but references the inclusion of broad multilingual materials and the use of community open-source repositories for evaluation benchmarks. Quantized model formats, such as GPTQ (Int4/Int8), AWQ, and GGUF, are officially provided, facilitating efficient deployment and lower resource utilization for inference.
Multilingual and General Language Capabilities
Qwen1.5-72B demonstrates considerable strength in multilingual understanding, covering at least 12 languages from diverse regions including Europe, East Asia, and Southeast Asia. Performance analysis on representative tasks shows substantial gains over earlier models—and in several cases, over peers like Llama 2 70B and GPT-3.5—in translation, language reasoning, and code execution benchmarks.

Figure 2. Qwen1.5-72B-Chat outperforms GPT-3.5 in multilingual language tasks across several languages including Arabic, French, Vietnamese, Korean, Japanese, Spanish, Indonesian, Portuguese, Russian, and Thai.
In tasks such as the MMLU benchmark (77.5), C-Eval (84.1), GSM8K (79.5), MATH (34.1), and HumanEval (41.5), Qwen1.5-72B displays broad competence across knowledge, reasoning, and coding domains. These benchmarks indicate not only linguistic fluency but also advanced capability in mathematics and programming, although there remains a performance gap in code interpretation when compared to models such as GPT-4.
Alignment with Human Preferences and Benchmark Performance
The Qwen1.5-72B-Chat model demonstrates high scores on benchmarks designed to measure alignment with human judgment and instruction-following quality. On MT-Bench the model achieves an average score of 8.61, with a 27.18 win rate on AlpacaEval 2.0.

Figure 3. Qwen1.5-72B-Chat achieves competitive results on MT-Bench and AlpacaEval 2.0, as shown in this benchmark comparison plot—prompting evaluation against models like GPT-4-Turbo and Mistral Medium.
Performance is especially notable when compared to contemporaries such as Claude-2.1, GPT-3.5-Turbo, and Mixtral 8x7B-Instruct, and approaches the scores of Mistral Medium. However, the model continues to trail GPT-4-Turbo on several of these gold-standard alignment benchmarks.
Integration, External System Connectivity, and Application Domains
One of the distinctive features of Qwen1.5-72B is its capacity for external system integration, including Retrieval-Augmented Generation (RAG), API-based tool and function calling for AI agent use cases, and Python code interpretation. The model performs competitively on RAG benchmarks, particularly in tasks requiring information retrieval under noise or counterfactual scenarios, and shows effective agent execution in both English and Chinese as measured by T-Eval.
Extensive evaluations demonstrate high tool-use accuracy (above 90% for both selection and input categories) and strong, though not leading, performance in mathematical and visualization-based code execution. The LEval long-context benchmark further highlights the model's ability to maintain performance in tasks involving complex, extended documents or dialogues.
Qwen1.5-72B has practical applications in conversational systems, code assistants, translation engines, mathematical problem solving, and customizable AI agents—thanks to its multilingual fluency, contextual memory, and open-ended instruction following.
Limitations and Future Directions
Despite robust scores in multilingual and alignment assessments, Qwen1.5-72B exhibits certain limitations. On advanced code interpretation and mathematical reasoning tasks, the model’s performance does not yet match that of GPT-4, especially for complex programming or difficult visualization problems. The efficacy of context windows significantly longer than 32,768 tokens, enabled by configuration changes, may vary and is not uniformly validated.
Continued improvements are anticipated in subsequent iterations to address these areas, with particular focus on enhancing coding capabilities and ensuring stable performance in longer-context applications.
Licensing and Access
Qwen1.5 models, including the 72B parameter variant, are open-source and are distributed via platforms like Hugging Face and ModelScope. While model licensing is not explicitly detailed in all official documentation, their open-source status facilitates broad research, adaptation, and integration into downstream AI pipelines.
External Resources
More in the Qwen 1 Family
More from Alibaba Cloud
Qwen3 0.6B
Qwen3 1.7B
Qwen3 4B
Qwen3 8B
Qwen3 14B
Qwen3 32B
Qwen3 30B A3B
Qwen3 235B A22B
Qwen2.5 VL 3B
Qwen2.5 VL 7B
Qwen2.5 VL 72B
QwQ 32B Preview
QwQ 32B
Qwen 2.5 Math 1.5B
Qwen 2.5 Math 7B
Qwen 2.5 Math PRM 7B
Qwen 2.5 Math 72B
Qwen 2.5 Math PRM 72B
Qwen 2.5 Coder 7B
Qwen 2.5 Coder 32B
Qwen 2.5 7B
Qwen2.5 7B 1M
Qwen 2.5 14B
Qwen2.5 14B 1M
Qwen 2.5 32B
Qwen 2.5 72B
Qwen 2 7B
Qwen 2 72B
Compatible Apps

Open WebUI
A polished, self-hosted chat interface for LLMs with Ollama integration, multimodal prompts, and extensive workspace customization.
Web UI
Chat UIs · Beginner Friendly
llama.cpp
GGUF model inference with a polished web UI and OpenAI-format API. CPU-only build — high hardware compatibility, works on any machine without a GPU.
Web UI · API · CLI
LLM Inference · Chat UIs
llama.cpp (CUDA)
GGUF model inference with a polished web UI and OpenAI-format API. CUDA build — GPU-accelerated for NVIDIA GPUs.
Web UI · API · CLI
LLM Inference · Chat UIs

Text Generation Web UI
A feature-rich interface for running and experimenting with open-weight LLMs, including multiple inference backends, plugins, and tuning controls.
Web UI · API
Chat UIs · LLM Inference