Skip to main content
Browse Models

Alibaba Cloud

Qwen 2.5 14B

Released

2024-09-19

Family

Qwen 2

Type

Foundation Model

Downloads

External Download

You are about to open a link to an external source. Verify the URL before continuing.

Download

Instruct model, 4-bit GGUF (Q4_K_M)

GGUF · Qwen2.5-14B-Instruct-Q4_K_M.gguf

Instruct model, 5-bit GGUF (Q5_K_M)

GGUF · Qwen2.5-14B-Instruct-Q5_K_M.gguf

Instruct model, 6-bit GGUF (Q6_K)

GGUF · Qwen2.5-14B-Instruct-Q6_K.gguf

Instruct model, 8-bit GGUF (Q8_0)

GGUF · Qwen2.5-14B-Instruct-Q8_0.gguf

Instruct model, 16-bit GGUF (F16)

GGUF · Qwen2.5-14B-Instruct-f16.gguf

Ollama Model (q4_K_M)

Ollama

Model Report

Overview

Qwen 2.5-14B is a large language model developed by the Qwen Team at Alibaba Group. Released in September 2024, it is part of the Qwen 2.5 model family, which builds upon previous Qwen iterations and introduces several specialized variants. Qwen 2.5 models are available as open-source, dense, decoder-only transformer models and span a range of parameter sizes and capabilities. The 14B version targets developers and researchers seeking a scalable base for further fine-tuning and domain-specific adaptation.

Qwen2.5 specification table

Figure 1. A tabular summary of the Qwen2.5 model family, highlighting key architectural and licensing details for each model variant.

An introduction to the Qwen2.5 project, detailing its motivations, architecture, and capabilities. · Source

Model Architecture and Features

Qwen2.5-14B is a causal language model built on the transformer architecture. It utilizes Rotary Position Embeddings (RoPE), the SwiGLU activation function, RMSNorm for layer normalization, and introduces Attention QKV biasing to enhance expressivity. The model contains approximately 14.7 billion parameters in total, with 13.1 billion non-embedding parameters. Architecturally, it is organized into 48 layers and implements Grouped Query Attention (GQA), allocating 40 heads for query vectors and 8 heads for key/value vectors.

Qwen2.5-14B supports an extended context window of up to 128,000 tokens, facilitating applications that demand significant memory and document processing capabilities. For generation tasks, it can produce output sequences of up to 8,000 tokens. The design of Qwen 2.5 series emphasizes adaptability to post-training techniques such as supervised fine-tuning and reinforcement learning from human feedback, and the base models are not recommended for direct conversational use prior to such refinement. Technical specifications and architectural choices are further detailed in the Qwen2.5 announcement and technical documentation.

Pretraining Data and Multilingual Support

Qwen2.5 models, including the 14B variant, are pretrained on a large-scale corpus comprising up to 18 trillion tokens. The pretraining data covers diverse domains and is curated for high data quality, aiming to foster robust language understanding and generation capabilities. The Qwen2.5 family also features expert models, such as Qwen2.5-Coder-7B and Qwen2.5-Math-7B, which are trained on specialist datasets, including 5.5 trillion code-related tokens for the Coder variant and synthetic math-focused data for the Math variant.

A hallmark of Qwen2.5 is its strong multilingual support. The model is capable of working with over 29 languages, including Chinese, English, French, Spanish, Portuguese, German, Italian, Russian, Japanese, Korean, Vietnamese, Thai, and Arabic. This enables enhanced instruction following and translation across a broad linguistic landscape, as documented in the Qwen technical resources and Hugging Face model repository.

Performance and Benchmarking

The Qwen2.5-14B model demonstrates competitive performance across a range of language understanding, reasoning, coding, and mathematical benchmarks. Notably, Qwen2.5-14B often matches or outperforms other models in its parameter class and, in some instances, rivals larger models on select tasks. Evaluation results highlight strong knowledge representation on the MMLU benchmark (with scores above 85), as well as high proficiency in coding (HumanEval 85+) and mathematics (MATH 80+).

Qwen2.5-14B and 32B benchmark results

Figure 2. Benchmark results comparing Qwen2.5-14B and Qwen2.5-32B to peer models on tasks including knowledge, reasoning, coding, and alignment.

In comprehensive comparisons, the Qwen2.5-14B model's instruction-following, long text generation, structured data understanding, and JSON output capabilities are evident. The model demonstrates robustness to various prompt structures, which improves both role-play and condition-setting for downstream chatbot applications. Detailed evaluation metrics and benchmarks are available in the Qwen2.5 performance summary.

Applications and Usage

While the base Qwen2.5-14B model is not intended for direct deployment in conversational settings without further fine-tuning, it serves as a foundation for a variety of advanced applications after supervised training. Use cases include complex logical reasoning, advanced mathematics, and code synthesis. The model is suitable for tasks such as multi-turn dialogues, instruction following, creative writing, and role-playing, particularly when preference alignment or agent-based functionality is required.

Specialized variants, such as Qwen2.5-Coder-7B and Qwen2.5-Math-7B, are tailored for intensive coding assistance—including debugging and code suggestion—and mathematical problem-solving using advanced reasoning strategies. After domain-specific or instruction tuning, Qwen2.5-14B may also serve as the core model in multi-agent systems and tool-augmented pipelines, enabled by its ability to handle diverse prompt styles and system settings. Further details on deployment and optimal settings can be found in the Qwen2.5 GitHub repository.

Model Family, Licensing, and Limitations

The Qwen2.5 series is positioned between Qwen2 and Qwen3 in the Qwen model family. Later generations, such as Qwen3, introduce additional architectural refinements, extended language coverage, and features like "thinking mode" and streamlined instruction tuning. Qwen2.5 models are distributed under the Apache 2.0 license, with the exception of select 3B and 72B variants.

Qwen2.5-14B, as a base model, is not recommended for out-of-the-box conversational use and should be fine-tuned for most real-world applications. In some specialized benchmarks, certain competing models may demonstrate higher scores, and context management settings may require attention when deployed in different environments. The model’s licensing, source code, and weights are openly available in their respective Hugging Face and GitHub repositories.

Helpful Links

About Qwen 2: Qwen 2 (and 2.5) is a family of advanced AI models developed by Alibaba, designed to excel in various tasks including general language understanding, coding, and mathematics.

More in the Qwen 2 Family

Alibaba Cloud /

Qwen2.5 VL 3B

A 3B-parameter multimodal language model that processes images, videos, and text with capabilities for visual reasoning, document analysis, and computer interface interaction.
Alibaba Cloud /

Qwen2.5 VL 7B

A 7-billion parameter multimodal model capable of processing images, documents, and videos with native dynamic resolution support and structured output capabilities.
Alibaba Cloud /

Qwen2.5 VL 72B

A 72-billion parameter multimodal model that processes images, videos, and text for document analysis, object detection, and visual agent tasks.
Alibaba Cloud /

QwQ 32B Preview

Experimental research model focused on advancing AI reasoning capabilities by thinking through problems step by step, continuously questioning assumptions, and exploring different paths of thought. Impressive benchmark scores in math and coding tasks.
Alibaba Cloud /

QwQ 32B

A 32.5-billion parameter transformer model optimized for mathematical reasoning, coding, and complex problem-solving through reinforcement learning training.
Alibaba Cloud /

Qwen 2.5 Math 1.5B

A 1.5 billion parameter specialized model focused on mathematical reasoning and problem-solving capabilities in English and Chinese.
Deepseek AI /

DeepSeek R1 Distill Qwen 1.5B

A 1.5 billion parameter language model created through distillation techniques, focusing on mathematical reasoning and chain-of-thought problem-solving capabilities.
Agentica /

DeepCoder 1.5B Preview

A 1.5B parameter code generation model fine-tuned with reinforcement learning and iterative context lengthening for long-context reasoning tasks.
Alibaba Cloud /

Qwen 2.5 Math 7B

A 7.62-billion parameter mathematical reasoning model supporting English and Chinese with chain-of-thought and tool-integrated reasoning capabilities.
Alibaba Cloud /

Qwen 2.5 Math PRM 7B

A 7B parameter process reward model that evaluates mathematical reasoning steps individually rather than just final answers, achieving 67.6% accuracy on mathematical benchmarks.
Deepseek AI /

DeepSeek R1 Distill Qwen 7B

A 7.62B parameter distilled language model based on Qwen2.5-Math-7B, trained via knowledge distillation for mathematical and logical reasoning tasks.
Alibaba Cloud /

Qwen 2.5 Math 72B

A 72.7 billion parameter language model specialized for solving mathematical problems in English and Chinese using chain-of-thought and tool-integrated reasoning.
Alibaba Cloud /

Qwen 2.5 Math PRM 72B

A 72.8 billion parameter process reward model that evaluates intermediate mathematical reasoning steps using consensus filtering and achieves 78.3% F1 on ProcessBench.
Alibaba Cloud /

Qwen 2.5 Coder 7B

A 7.61 billion parameter transformer model designed for code generation and reasoning across 92 programming languages with a 128,000-token context window.
Alibaba Cloud /

Qwen 2.5 Coder 32B

Qwen2.5-Coder builds on Qwen2.5 by training on 5.5 trillion additional tokens of code data including source code, text-code grounding data, and synthetic data. This leads to significant improvements in code-related tasks.
Alibaba Cloud /

Qwen 2.5 7B

A 7.61-billion parameter transformer-based language model with 128K token context length, trained on 18 trillion tokens supporting 29+ languages.
Alibaba Cloud /

Qwen2.5 7B 1M

A 7-billion parameter transformer model with 1-million token context capacity utilizing Dual Chunk Attention for long-range text processing.
Alibaba Cloud /

Qwen2.5 14B 1M

A 14.7B parameter transformer model with 1-million token context capacity, utilizing dual chunk attention and progressive training for extended sequence processing.
Deepseek AI /

DeepSeek R1 Distill Qwen 14B

A 14B dense language model distilled from a mixture-of-experts architecture, optimized for mathematical reasoning and code generation tasks.
Agentica /

DeepCoder 14B Preview

A 14-billion parameter code reasoning model fine-tuned using distributed reinforcement learning with long-context capabilities up to 64,000 tokens.
Deep Cogito /

Cogito V1 Preview 14B

A 14.8 billion parameter instruction-tuned language model trained using Iterated Distillation and Amplification with hybrid reasoning capabilities across 30+ languages.
Alibaba Cloud /

Qwen 2.5 32B

A 32.5 billion parameter multilingual transformer model trained on 18 trillion tokens with 128K context length supporting text generation, coding, and mathematical reasoning.
Deepseek AI /

DeepSeek R1 Distill Qwen 32B

A 32B parameter language model created through knowledge distillation, optimized for mathematical reasoning, code generation, and complex problem-solving tasks.
Deep Cogito /

Cogito V1 Preview 32B

A 32-billion parameter instruction-tuned model based on Qwen2.5 architecture featuring dual operational modes and iterated distillation alignment methodology.
Alibaba Cloud /

Qwen 2.5 72B

This update to Qwen 2, pretrained on 18 trillion tokens, shows impressive performance in benchmarks (MMLU 85+, HumanEval 85+, MATH 80+), outperforming models of similar size. Qwen 2.5 has multilingual support for over 29 languages and can understand and generate structured data.
Alibaba Cloud /

Qwen 2 7B

A 7.6 billion parameter multilingual decoder-only Transformer model designed for further post-training with 32,000-token context and strong coding capabilities.
Alibaba Cloud /

Qwen 2 72B

A 72-billion parameter transformer model supporting extended context lengths up to 128K tokens and demonstrating strong multilingual capabilities across diverse benchmarks.

More from Alibaba Cloud

Alibaba Cloud /

Qwen3 0.6B

A 0.6B parameter language model featuring dual thinking modes, multilingual capabilities, and 32K context length through strong-to-weak distillation training.
Alibaba Cloud /

Qwen3 1.7B

A 1.7 billion parameter multilingual transformer supporting dual-mode reasoning with step-by-step "thinking" and rapid "non-thinking" response capabilities.
Alibaba Cloud /

Qwen3 4B

A 4-billion parameter transformer model featuring dual reasoning modes, extensive multilingual training, and competitive performance across mathematical, coding, and logical reasoning benchmarks.
Alibaba Cloud /

Qwen3 8B

Dense 8.2 billion parameter transformer model featuring hybrid thinking capabilities with 32K token context and multilingual support across 119 languages.
Alibaba Cloud /

Qwen3 14B

A 14.8 billion parameter transformer model featuring hybrid thinking/non-thinking reasoning modes, 32K context length, and multilingual capabilities across 119 languages.
Alibaba Cloud /

Qwen3 32B

A 32.8 billion parameter language model featuring hybrid thinking modes for both rapid responses and step-by-step reasoning across multilingual tasks.
Alibaba Cloud /

Qwen3 30B A3B

Mixture-of-experts model with 30.5B total parameters, 3.3B activated per token, featuring hybrid reasoning modes and multilingual support across 119 languages.
Alibaba Cloud /

Qwen3 235B A22B

Large language model with Mixture-of-Experts architecture featuring dual operational modes for both rapid inference and complex reasoning tasks.
Alibaba Cloud /

Qwen 1.5 32B

Foundation 32B Qwen 1.5 model from Alibaba Cloud.
Alibaba Cloud /

Qwen 1.5 72B

Foundation 72B Qwen 1.5 model from Alibaba Cloud.