Skip to main content
Browse Models

Alibaba Cloud

Qwen 2.5 7B

Released

2024-09-19

Family

Qwen 2

Type

Foundation Model

Downloads

External Download

You are about to open a link to an external source. Verify the URL before continuing.

Download

Instruct model, 4-bit GGUF (Q4_K_M)

GGUF · Qwen2.5-7B-Instruct-Q4_K_M.gguf

Instruct model, 5-bit GGUF (Q5_K_M)

GGUF · Qwen2.5-7B-Instruct-Q5_K_M.gguf

Instruct model, 6-bit GGUF (Q6_K)

GGUF · Qwen2.5-7B-Instruct-Q6_K.gguf

Instruct model, 8-bit GGUF (Q8_0)

GGUF · Qwen2.5-7B-Instruct-Q8_0.gguf

Instruct model, 16-bit GGUF (F16)

GGUF · Qwen2.5-7B-Instruct-f16.gguf

Ollama Model (q4_K_M)

Ollama

Model Report

Overview

Qwen2.5-7B is a foundational large language model (LLM) developed by the Qwen Team at Alibaba Group as part of the Qwen2.5 series. Released on September 19, 2024, Qwen2.5-7B occupies a key position among a suite of multilingual and multimodal models designed for a broad range of natural language processing tasks. Distinguished by its transformer architecture and extensive pretraining, Qwen2.5-7B is intended as a base model for further post-training or adaptation. The Qwen2.5 series also features specialized models targeting domains like code generation (Qwen2.5-Coder) and mathematics (Qwen2.5-Math).

Qwen2.5 Model Family Specification Table

Figure 1. Specifications for the Qwen2.5 model family, including base, Coder, and Math variants, such as model size, architecture, context length, and license.

Overview video introducing the features and scope of the Qwen2.5 project. · Source

Architecture and Training

At its core, Qwen2.5-7B is a transformer-based causal language model, comprising approximately 7.61 billion total parameters, with 6.53 billion allocated to non-embedding roles. The model utilizes 28 layers, a grouped query attention (GQA) mechanism—featuring 28 attention heads for queries and 4 for key/value—and incorporates advanced architectural features such as rotary position embeddings (RoPE), SwiGLU activation functions, root mean square normalization (RMSNorm), and attention QKV bias to enhance performance and efficiency. These enhancements support robust multilingual capabilities and improved generalization.

Qwen2.5-7B was trained on a dataset encompassing up to 18 trillion tokens, sourced from a diverse collection of large-scale multilingual and multimodal data, as described in the Qwen2.5 technical report. Post-training refinement is conducted through alignment with human preferences to enhance the model’s applicability in real-world tasks.

Technical Features and Capabilities

Qwen2.5-7B is designed as a base model, not instruction-tuned, making it suitable for fine-tuning and further adaptation across specialized applications. With support for context lengths up to 128,000 tokens and output sequences up to 8,000 tokens, the model is suitable for long-form text generation and contextual understanding. It provides broad multilingual support across more than 29 languages, including but not limited to Chinese, English, French, Spanish, Russian, Japanese, Arabic, Vietnamese, and Thai.

The Qwen2.5 series introduces several enhancements over its predecessor, Qwen2, as documented in the release blog. These enhancements include an increased scope of world knowledge, advancements in coding and mathematical problem solving, improved robustness to various system prompts, refined structured data understanding, and the generation of structured outputs such as JSON. The model design also supports prompt engineering scenarios, contributing to role-play and dialog management.

Math benchmark scatter plot

Figure 2. Math accuracy of Qwen2.5-Math models—including 7B—in comparison with other models, showing performance relative to parameter size. (Prompt: MATH zero-shot accuracy evaluation)

Model Performance and Evaluation

Benchmarking results for the Qwen2.5 series, including the 7B variant, are detailed in the Qwen2.5 release post. The family's performance on standard academic and industry evaluations is documented:

  • On MMLU (Massive Multitask Language Understanding), Qwen2.5 series models achieve scores above 85.
  • On HumanEval (code generation tasks), scores also surpass 85 for the series as a whole.
  • For mathematical reasoning, as evaluated on MATH, accuracies above 80 are reported.
  • Comparative analysis against other open-source models—such as Llama-3.1-70B and Mistral-Large-V2—indicates that larger Qwen2.5 variants exhibit performance characteristics reported as comparable to or surpassing these models.

Specialized expert models within the Qwen2.5 family, such as Qwen2.5-Coder and Qwen2.5-Math, achieve high performance on coding and mathematical problem sets in their respective benchmark evaluations.

Applications and Use Cases

As a base model, Qwen2.5-7B is designed predominantly for further post-training and development, rather than direct conversational deployment. Typical application areas include:

  • Natural language understanding and long-form text generation, supported by extended context handling and contextual understanding.
  • Multilingual applications such as translation and instruction following in over 29 languages.
  • Structured data processing, including the extraction, completion, and generation of tabular or JSON-formatted outputs.
  • Advanced tool integration and function calling for building AI agents, as supported by frameworks and toolkits.
  • Specialized reasoning tasks, including coding (through Qwen2.5-Coder expert models) and mathematical problem solving (using Qwen2.5-Math).

Its architecture and design support performance in agent-based environments, role-play scenarios, and settings that involve prompt engineering and persona management.

Model Family and Development Timeline

The Qwen2.5 series comprises a range of models—both base and instruction-tuned—from 0.5B to 72B parameters, with expert offshoots like Qwen2.5-Coder and Qwen2.5-Math optimizing for code and mathematical domains, respectively. Key milestones include the Qwen2.5 release in September 2024, Qwen2 in June 2024, and the subsequent introduction of Qwen3 in April 2025, which includes additional capabilities in reasoning and agent integration as detailed in the Qwen3 technical report.

Limitations and Licensing

Qwen2.5-7B is distributed as a base model and is specifically not recommended for direct end-user conversation tasks without further post-training or fine-tuning. Within its parameter class, there are inherent limitations compared to larger LLMs, particularly in handling complex tasks or nuanced dialog. Licensing for Qwen2.5-7B follows the Apache 2.0 standard, providing permissive access for research and development.

Helpful Resources

About Qwen 2: Qwen 2 (and 2.5) is a family of advanced AI models developed by Alibaba, designed to excel in various tasks including general language understanding, coding, and mathematics.

More in the Qwen 2 Family

Alibaba Cloud /

Qwen2.5 VL 3B

A 3B-parameter multimodal language model that processes images, videos, and text with capabilities for visual reasoning, document analysis, and computer interface interaction.
Alibaba Cloud /

Qwen2.5 VL 7B

A 7-billion parameter multimodal model capable of processing images, documents, and videos with native dynamic resolution support and structured output capabilities.
Alibaba Cloud /

Qwen2.5 VL 72B

A 72-billion parameter multimodal model that processes images, videos, and text for document analysis, object detection, and visual agent tasks.
Alibaba Cloud /

QwQ 32B Preview

Experimental research model focused on advancing AI reasoning capabilities by thinking through problems step by step, continuously questioning assumptions, and exploring different paths of thought. Impressive benchmark scores in math and coding tasks.
Alibaba Cloud /

QwQ 32B

A 32.5-billion parameter transformer model optimized for mathematical reasoning, coding, and complex problem-solving through reinforcement learning training.
Alibaba Cloud /

Qwen 2.5 Math 1.5B

A 1.5 billion parameter specialized model focused on mathematical reasoning and problem-solving capabilities in English and Chinese.
Deepseek AI /

DeepSeek R1 Distill Qwen 1.5B

A 1.5 billion parameter language model created through distillation techniques, focusing on mathematical reasoning and chain-of-thought problem-solving capabilities.
Agentica /

DeepCoder 1.5B Preview

A 1.5B parameter code generation model fine-tuned with reinforcement learning and iterative context lengthening for long-context reasoning tasks.
Alibaba Cloud /

Qwen 2.5 Math 7B

A 7.62-billion parameter mathematical reasoning model supporting English and Chinese with chain-of-thought and tool-integrated reasoning capabilities.
Alibaba Cloud /

Qwen 2.5 Math PRM 7B

A 7B parameter process reward model that evaluates mathematical reasoning steps individually rather than just final answers, achieving 67.6% accuracy on mathematical benchmarks.
Deepseek AI /

DeepSeek R1 Distill Qwen 7B

A 7.62B parameter distilled language model based on Qwen2.5-Math-7B, trained via knowledge distillation for mathematical and logical reasoning tasks.
Alibaba Cloud /

Qwen 2.5 Math 72B

A 72.7 billion parameter language model specialized for solving mathematical problems in English and Chinese using chain-of-thought and tool-integrated reasoning.
Alibaba Cloud /

Qwen 2.5 Math PRM 72B

A 72.8 billion parameter process reward model that evaluates intermediate mathematical reasoning steps using consensus filtering and achieves 78.3% F1 on ProcessBench.
Alibaba Cloud /

Qwen 2.5 Coder 7B

A 7.61 billion parameter transformer model designed for code generation and reasoning across 92 programming languages with a 128,000-token context window.
Alibaba Cloud /

Qwen 2.5 Coder 32B

Qwen2.5-Coder builds on Qwen2.5 by training on 5.5 trillion additional tokens of code data including source code, text-code grounding data, and synthetic data. This leads to significant improvements in code-related tasks.
Alibaba Cloud /

Qwen2.5 7B 1M

A 7-billion parameter transformer model with 1-million token context capacity utilizing Dual Chunk Attention for long-range text processing.
Alibaba Cloud /

Qwen 2.5 14B

A 14.7 billion parameter transformer model supporting 29 languages with 128K context window, designed as foundation for fine-tuning applications.
Alibaba Cloud /

Qwen2.5 14B 1M

A 14.7B parameter transformer model with 1-million token context capacity, utilizing dual chunk attention and progressive training for extended sequence processing.
Deepseek AI /

DeepSeek R1 Distill Qwen 14B

A 14B dense language model distilled from a mixture-of-experts architecture, optimized for mathematical reasoning and code generation tasks.
Agentica /

DeepCoder 14B Preview

A 14-billion parameter code reasoning model fine-tuned using distributed reinforcement learning with long-context capabilities up to 64,000 tokens.
Deep Cogito /

Cogito V1 Preview 14B

A 14.8 billion parameter instruction-tuned language model trained using Iterated Distillation and Amplification with hybrid reasoning capabilities across 30+ languages.
Alibaba Cloud /

Qwen 2.5 32B

A 32.5 billion parameter multilingual transformer model trained on 18 trillion tokens with 128K context length supporting text generation, coding, and mathematical reasoning.
Deepseek AI /

DeepSeek R1 Distill Qwen 32B

A 32B parameter language model created through knowledge distillation, optimized for mathematical reasoning, code generation, and complex problem-solving tasks.
Deep Cogito /

Cogito V1 Preview 32B

A 32-billion parameter instruction-tuned model based on Qwen2.5 architecture featuring dual operational modes and iterated distillation alignment methodology.
Alibaba Cloud /

Qwen 2.5 72B

This update to Qwen 2, pretrained on 18 trillion tokens, shows impressive performance in benchmarks (MMLU 85+, HumanEval 85+, MATH 80+), outperforming models of similar size. Qwen 2.5 has multilingual support for over 29 languages and can understand and generate structured data.
Alibaba Cloud /

Qwen 2 7B

A 7.6 billion parameter multilingual decoder-only Transformer model designed for further post-training with 32,000-token context and strong coding capabilities.
Alibaba Cloud /

Qwen 2 72B

A 72-billion parameter transformer model supporting extended context lengths up to 128K tokens and demonstrating strong multilingual capabilities across diverse benchmarks.

More from Alibaba Cloud

Alibaba Cloud /

Qwen3 0.6B

A 0.6B parameter language model featuring dual thinking modes, multilingual capabilities, and 32K context length through strong-to-weak distillation training.
Alibaba Cloud /

Qwen3 1.7B

A 1.7 billion parameter multilingual transformer supporting dual-mode reasoning with step-by-step "thinking" and rapid "non-thinking" response capabilities.
Alibaba Cloud /

Qwen3 4B

A 4-billion parameter transformer model featuring dual reasoning modes, extensive multilingual training, and competitive performance across mathematical, coding, and logical reasoning benchmarks.
Alibaba Cloud /

Qwen3 8B

Dense 8.2 billion parameter transformer model featuring hybrid thinking capabilities with 32K token context and multilingual support across 119 languages.
Alibaba Cloud /

Qwen3 14B

A 14.8 billion parameter transformer model featuring hybrid thinking/non-thinking reasoning modes, 32K context length, and multilingual capabilities across 119 languages.
Alibaba Cloud /

Qwen3 32B

A 32.8 billion parameter language model featuring hybrid thinking modes for both rapid responses and step-by-step reasoning across multilingual tasks.
Alibaba Cloud /

Qwen3 30B A3B

Mixture-of-experts model with 30.5B total parameters, 3.3B activated per token, featuring hybrid reasoning modes and multilingual support across 119 languages.
Alibaba Cloud /

Qwen3 235B A22B

Large language model with Mixture-of-Experts architecture featuring dual operational modes for both rapid inference and complex reasoning tasks.
Alibaba Cloud /

Qwen 1.5 32B

Foundation 32B Qwen 1.5 model from Alibaba Cloud.
Alibaba Cloud /

Qwen 1.5 72B

Foundation 72B Qwen 1.5 model from Alibaba Cloud.