Skip to main content
Browse Models

Alibaba Cloud

Qwen2.5 VL 3B

Released

2025-01-26

Family

Qwen 2

Type

Foundation Model

Downloads

External Download

You are about to open a link to an external source. Verify the URL before continuing.

Download

4-bit GGUF (Q4_K_M)

GGUF · Qwen2.5-VL-3B-Instruct-Q4_K_M.gguf

6-bit GGUF (Q6_K)

GGUF · Qwen2.5-VL-3B-Instruct-Q6_K.gguf

8-bit GGUF (Q8_0)

GGUF · Qwen2.5-VL-3B-Instruct-Q8_0.gguf

Model Report

Overview

Qwen2.5-VL-3B-Instruct is a multimodal, instruction-tuned large language model developed by the Qwen team at Alibaba Cloud. As part of the Qwen2.5-VL series, which includes larger 7B and 72B parameter variants, the 3B model is designed for efficient, on-device deployment while maintaining capabilities in image, video, and multimodal understanding. Released in January 2025, Qwen2.5-VL-3B-Instruct includes architectural and functional refinements over its predecessor, Qwen2-VL, particularly in visual comprehension, agentic behavior, and temporal processing for long videos, with further developments described in the Qwen2.5-VL technical report.

A woman and a dog on the beach

Figure 1. Qwen2.5-VL-3B processes and describes scenes involving people, animals, and diverse objects, supporting detailed multimodal understanding.

Model Architecture and Innovations

Qwen2.5-VL-3B-Instruct employs a multi-stage training strategy and incorporates several architectural elements. The core of the system integrates a Vision Transformer (ViT) encoder with a Qwen2.5-series large language model decoder, connected via cross-modal layers. The ViT architecture incorporates window attention for efficiency, with four of its layers utilizing full attention and the remainder operating in windowed mode for computational scalability. The model natively supports dynamic resolution inputs for both image and video data, enabling it to process media in their original aspect ratios for spatial sensitivity.

A critical innovation is the extension of dynamic resolution into the temporal dimension for video understanding. Qwen2.5-VL-3B samples video frames at dynamic rates and applies Multimodal Rotary Position Embedding (mRoPE) with absolute time alignment, allowing it to discern and reason over long video sequences and accurately localize events.

Qwen2.5-VL model architecture diagram

Figure 2. Technical diagram of the Qwen2.5-VL model, showing the interaction between the vision encoder, dynamic resolution handling, and temporal processing for images and videos.

Additionally, the model is configured for native box and point representation: it can directly output bounding box coordinates and keypoints in the context of the original image frame, bypassing traditional normalization techniques and contributing to fidelity in object localization.

Technical Capabilities

The 3B-Instruct variant provides multimodal reasoning across images, documents, and video streams. It demonstrates performance in object recognition, chart analysis, layout understanding, and text extraction, including optical character recognition (OCR) for multilingual and multi-orientational scenarios.

Bird detection with overlaid dots and labels

Figure 3. Bird counting demonstration: Qwen2.5-VL-3B detects and counts all birds in the image, including partially visible ones.

In video-related tasks, the model performs event detection, segment localization, and structured captioning, supporting variable resolution and video length.

Demonstrating Qwen2.5-VL-3B's video reasoning capabilities by providing detailed object analysis and information extraction from video frames. · Source

The model’s agentic features allow it to interact with computer and mobile interfaces by interpreting user interfaces and executing actions in applications—including document editing, image manipulation, and task automation.

Computer agent demonstration: Qwen2.5-VL-3B performs photo editing tasks in a desktop application based on user instructions. · Source

Performance and Evaluation

Qwen2.5-VL-3B-Instruct has been evaluated against established vision-language benchmarks. Results from evaluations against established vision-language benchmarks indicate its performance in comparison to larger models in its class and similarly efficient open models. On tasks such as multi-modal reasoning (MMMU), chart and diagram interpretation (DocVQA), visual question answering (AI2D, InfoVQA, TextVQA), mathematics (MathVista, MathVision), and video reasoning (VideoMME, MVBench), the 3B model's reported scores are comparable to those of some higher-parameter open models.

Qwen2.5-VL model illustration

Figure 4. Model banner for the Qwen2.5-VL series, highlighting its multimodal focus.

The model's structured output capabilities enable extraction of information from invoices, forms, and receipts, and its QwenVL HTML format provides detailed document layout extraction, supporting downstream applications in finance, logistics, and commercial documentation.

Training Data and Methodology

Training of Qwen2.5-VL-3B-Instruct follows a three-stage process. Initially, the vision encoder is trained independently on a corpus of image-text pairs to establish foundational multimodal representations. In the subsequent stage, all model parameters are unfrozen and are further pre-trained on a broader dataset incorporating images, OCR, document formats, and visual question answering (VQA) data. The final stage involves locking the visual encoder weights and fine-tuning the language model on curated instruction datasets formatted in ChatML, encompassing both text and multimodal conversational data.

Pretraining leverages approximately 1.4 trillion tokens, including image, video, and text modalities. The data is composed of cleaned web data, curated open-source datasets, and synthetic sources, with a knowledge cutoff in June 2023. During fine-tuning, datasets span standard dialog, multi-image comparison, document parsing, video comprehension, and agent interaction.

Typical Applications

Qwen2.5-VL-3B-Instruct is suited for a range of applications requiring fine-grained visual analysis, document and chart parsing, information extraction from receipts or invoices, object counting, keypoint detection, and temporal localization within video. It supports accessibility solutions, process automation in mobile and desktop environments, multimedia content moderation, and educational applications that rely on multimodal understanding.

Object localization and helmet detection

Figure 5. Demonstrating precise object grounding: the model localizes multiple motorcyclists, indicating helmet usage with bounding boxes.

Mobile agent example: Qwen2.5-VL-3B assists in booking a ticket in a mobile application by interpreting UI elements and automating input. · Source
Structured video captioning: the model identifies and describes segmented activity events with precise timestamps. · Source

Limitations and Licensing

While Qwen2.5-VL-3B-Instruct processes context up to 32,768 tokens by default, certain extensions to context length (such as YaRN) may negatively impact the spatial or temporal precision required for some tasks. Careful parameter configuration is necessary in these cases, particularly when performing OCR on small images, where excessive upscaling may degrade performance due to shifts in data distribution.

Qwen2.5-VL-3B-Instruct is released under the Apache-2.0 license, which permits broad research and development use.

External Resources

About Qwen 2: Qwen 2 (and 2.5) is a family of advanced AI models developed by Alibaba, designed to excel in various tasks including general language understanding, coding, and mathematics.

More in the Qwen 2 Family

Alibaba Cloud /

Qwen2.5 VL 7B

A 7-billion parameter multimodal model capable of processing images, documents, and videos with native dynamic resolution support and structured output capabilities.
Alibaba Cloud /

Qwen2.5 VL 72B

A 72-billion parameter multimodal model that processes images, videos, and text for document analysis, object detection, and visual agent tasks.
Alibaba Cloud /

QwQ 32B Preview

Experimental research model focused on advancing AI reasoning capabilities by thinking through problems step by step, continuously questioning assumptions, and exploring different paths of thought. Impressive benchmark scores in math and coding tasks.
Alibaba Cloud /

QwQ 32B

A 32.5-billion parameter transformer model optimized for mathematical reasoning, coding, and complex problem-solving through reinforcement learning training.
Alibaba Cloud /

Qwen 2.5 Math 1.5B

A 1.5 billion parameter specialized model focused on mathematical reasoning and problem-solving capabilities in English and Chinese.
Deepseek AI /

DeepSeek R1 Distill Qwen 1.5B

A 1.5 billion parameter language model created through distillation techniques, focusing on mathematical reasoning and chain-of-thought problem-solving capabilities.
Agentica /

DeepCoder 1.5B Preview

A 1.5B parameter code generation model fine-tuned with reinforcement learning and iterative context lengthening for long-context reasoning tasks.
Alibaba Cloud /

Qwen 2.5 Math 7B

A 7.62-billion parameter mathematical reasoning model supporting English and Chinese with chain-of-thought and tool-integrated reasoning capabilities.
Alibaba Cloud /

Qwen 2.5 Math PRM 7B

A 7B parameter process reward model that evaluates mathematical reasoning steps individually rather than just final answers, achieving 67.6% accuracy on mathematical benchmarks.
Deepseek AI /

DeepSeek R1 Distill Qwen 7B

A 7.62B parameter distilled language model based on Qwen2.5-Math-7B, trained via knowledge distillation for mathematical and logical reasoning tasks.
Alibaba Cloud /

Qwen 2.5 Math 72B

A 72.7 billion parameter language model specialized for solving mathematical problems in English and Chinese using chain-of-thought and tool-integrated reasoning.
Alibaba Cloud /

Qwen 2.5 Math PRM 72B

A 72.8 billion parameter process reward model that evaluates intermediate mathematical reasoning steps using consensus filtering and achieves 78.3% F1 on ProcessBench.
Alibaba Cloud /

Qwen 2.5 Coder 7B

A 7.61 billion parameter transformer model designed for code generation and reasoning across 92 programming languages with a 128,000-token context window.
Alibaba Cloud /

Qwen 2.5 Coder 32B

Qwen2.5-Coder builds on Qwen2.5 by training on 5.5 trillion additional tokens of code data including source code, text-code grounding data, and synthetic data. This leads to significant improvements in code-related tasks.
Alibaba Cloud /

Qwen 2.5 7B

A 7.61-billion parameter transformer-based language model with 128K token context length, trained on 18 trillion tokens supporting 29+ languages.
Alibaba Cloud /

Qwen2.5 7B 1M

A 7-billion parameter transformer model with 1-million token context capacity utilizing Dual Chunk Attention for long-range text processing.
Alibaba Cloud /

Qwen 2.5 14B

A 14.7 billion parameter transformer model supporting 29 languages with 128K context window, designed as foundation for fine-tuning applications.
Alibaba Cloud /

Qwen2.5 14B 1M

A 14.7B parameter transformer model with 1-million token context capacity, utilizing dual chunk attention and progressive training for extended sequence processing.
Deepseek AI /

DeepSeek R1 Distill Qwen 14B

A 14B dense language model distilled from a mixture-of-experts architecture, optimized for mathematical reasoning and code generation tasks.
Agentica /

DeepCoder 14B Preview

A 14-billion parameter code reasoning model fine-tuned using distributed reinforcement learning with long-context capabilities up to 64,000 tokens.
Deep Cogito /

Cogito V1 Preview 14B

A 14.8 billion parameter instruction-tuned language model trained using Iterated Distillation and Amplification with hybrid reasoning capabilities across 30+ languages.
Alibaba Cloud /

Qwen 2.5 32B

A 32.5 billion parameter multilingual transformer model trained on 18 trillion tokens with 128K context length supporting text generation, coding, and mathematical reasoning.
Deepseek AI /

DeepSeek R1 Distill Qwen 32B

A 32B parameter language model created through knowledge distillation, optimized for mathematical reasoning, code generation, and complex problem-solving tasks.
Deep Cogito /

Cogito V1 Preview 32B

A 32-billion parameter instruction-tuned model based on Qwen2.5 architecture featuring dual operational modes and iterated distillation alignment methodology.
Alibaba Cloud /

Qwen 2.5 72B

This update to Qwen 2, pretrained on 18 trillion tokens, shows impressive performance in benchmarks (MMLU 85+, HumanEval 85+, MATH 80+), outperforming models of similar size. Qwen 2.5 has multilingual support for over 29 languages and can understand and generate structured data.
Alibaba Cloud /

Qwen 2 7B

A 7.6 billion parameter multilingual decoder-only Transformer model designed for further post-training with 32,000-token context and strong coding capabilities.
Alibaba Cloud /

Qwen 2 72B

A 72-billion parameter transformer model supporting extended context lengths up to 128K tokens and demonstrating strong multilingual capabilities across diverse benchmarks.

More from Alibaba Cloud

Alibaba Cloud /

Qwen3 0.6B

A 0.6B parameter language model featuring dual thinking modes, multilingual capabilities, and 32K context length through strong-to-weak distillation training.
Alibaba Cloud /

Qwen3 1.7B

A 1.7 billion parameter multilingual transformer supporting dual-mode reasoning with step-by-step "thinking" and rapid "non-thinking" response capabilities.
Alibaba Cloud /

Qwen3 4B

A 4-billion parameter transformer model featuring dual reasoning modes, extensive multilingual training, and competitive performance across mathematical, coding, and logical reasoning benchmarks.
Alibaba Cloud /

Qwen3 8B

Dense 8.2 billion parameter transformer model featuring hybrid thinking capabilities with 32K token context and multilingual support across 119 languages.
Alibaba Cloud /

Qwen3 14B

A 14.8 billion parameter transformer model featuring hybrid thinking/non-thinking reasoning modes, 32K context length, and multilingual capabilities across 119 languages.
Alibaba Cloud /

Qwen3 32B

A 32.8 billion parameter language model featuring hybrid thinking modes for both rapid responses and step-by-step reasoning across multilingual tasks.
Alibaba Cloud /

Qwen3 30B A3B

Mixture-of-experts model with 30.5B total parameters, 3.3B activated per token, featuring hybrid reasoning modes and multilingual support across 119 languages.
Alibaba Cloud /

Qwen3 235B A22B

Large language model with Mixture-of-Experts architecture featuring dual operational modes for both rapid inference and complex reasoning tasks.
Alibaba Cloud /

Qwen 1.5 32B

Foundation 32B Qwen 1.5 model from Alibaba Cloud.
Alibaba Cloud /

Qwen 1.5 72B

Foundation 72B Qwen 1.5 model from Alibaba Cloud.