Alibaba Cloud
Qwen2.5 VL 72B
Downloads
Model Report
Overview
Qwen2.5-VL 72B is a configuration within the Qwen2.5-VL family, a suite of large multimodal generative AI models developed by the Qwen team at Alibaba Cloud. Released in early 2025, Qwen2.5-VL unifies vision and language understanding. Its capabilities include image and video comprehension, document parsing, object grounding, structured data extraction, and visual-agent interactions, as detailed in the Qwen2.5-VL Technical Report and on the Qwen2.5 VL Blog. Containing 72 billion parameters, the Qwen2.5-VL-72B-Instruct functions as a configuration for multimodal tasks, succeeding the earlier Qwen2-VL models and introducing architectural, training, and functional enhancements.

Figure 1. Banner for Qwen2.5-VL illustrating the model’s launch.
Model Architecture
Qwen2.5-VL 72B is built upon a unified architecture that integrates a large language model from the Qwen2 series with a vision encoder, enabling comprehensive visual-language reasoning. Key architectural features include:
- Dynamic Resolution Processing: The vision encoder adopts a native vision transformer (ViT) optimized for dynamic resolutions. Images and videos of differing spatial and temporal dimensions are tokenized according to actual scale, accommodating variable input sizes without conventional normalization, as described in the Qwen2.5-VL Technical Report.
- Temporal and Spatial Alignment: Videos are sampled at dynamically adjustable frame rates; absolute time encoding aligns multimodal rotary position embeddings (mRoPE) with real video durations, improving event localization and summarization in long-form video analysis, as presented in the Qwen2.5-VL Technical Report.
- Efficient Attention Mechanisms: The vision encoder structure strategically combines full attention layers with windowed attention for efficient computation, further stabilized by RMSNorm and SwiGLU techniques to harmonize architectural consistency with the Qwen2 LLM, according to the Qwen2.5 VL Blog.

Figure 2. Schematic of Qwen2.5-VL’s video understanding pipeline and modular network design.
- Structured Output and Grounding: The model supports outputting bounding boxes, points, and structured JSON data for detected objects, which is for tasks such as precise object localization and form data extraction.
Training Methodology and Datasets
The training of Qwen2.5-VL employs a three-phase approach, optimizing each component for multimodal understanding, as detailed in the Qwen2.5-VL Technical Report:
- Stage 1: The ViT encoder is pre-trained with image-text pairs, focusing on image classification, OCR, and semantic alignment. Initial weights are adapted from large vision models but use rotary 2D positional embeddings.
- Stage 2: Multimodal joint training unfreezes all network parameters, introducing diverse data including visual question answering, multitask datasets, and continued text-only learning for language robustness.
- Stage 3: Parameters for the vision encoder are frozen, while the LLM undergoes instruction fine-tuning using conversational, document, video, and agent-based datasets in the ChatML format.
The curriculum exposes Qwen2.5-VL to over 1.4 trillion tokens, blending textual and visual data. Specialized agent datasets enable the model to reason through UI operations and decision-making tasks, while OCR and document parsing data ensure reliable recognition under varied orientations and languages.
Technical Capabilities
Qwen2.5-VL 72B’s capabilities extend across multiple modalities and tasks, including:
- Object Detection and Grounding: The model provides object localization with bounding boxes and labels, supporting hierarchical grounding for complex scenes.

Figure 3. Output showing detection and helmet classification among motorcyclists.
- Fine-Grained Keypoint Detection: Qwen2.5-VL can identify and label keypoints, supporting annotation of entities like body parts in sports imagery.

Figure 4. Model output demonstrating head and hand keypoint detection for basketball players prompted for body part localization.
- Robust Object Counting and Classification: The model accurately counts and identifies multiple instances, including partially occluded objects, as demonstrated in benchmarks and practical outputs.

Figure 5. Benchmark comparison between compact Qwen2.5-VL models and peers.

Figure 6. 16 birds automatically detected and labeled in a natural scene, prompted for total bird count.

Figure 7. Detections and object count output for a set of summer-themed items. Prompt: count and label all objects.
- Visual Content Structuring: The model delivers structured outputs for complex scenes, including documents, receipts, and engineering tables.

Figure 8. Performance comparison for structured recognition in documents and tables.

Figure 9. Multiple cupcakes grounded and described with bounding boxes, demonstrating attribute extraction. Prompt: enumerate coordinates and features of all cupcakes.

Figure 10. Shipping label and building address matched for key information extraction and verification tasks.
- Document Understanding and OCR: Upgraded OCR enables precise, multilingual, and multi-orientation recognition in complex documents, such as scanned receipts and structured ledgers.

Figure 11. Parsing a technical report and outputting HTML-like layout, demonstrating structured textual and graphic extraction.
- Video Comprehension: Qwen2.5-VL can process and summarize long-form videos, identify events at second-level granularity, and output structured descriptions for video segments.
- Agentic and Interactive Abilities: The model can operate as a visual agent for dynamic reasoning and tool manipulation in computer and mobile environments.
Benchmark Evaluation
Qwen2.5-VL-72B-Instruct delivers strong performance on established benchmarks spanning math, science, document understanding, object recognition, and video analysis, as reported in the official technical report and model documentation.

Figure 12. Comparative benchmark scores of Qwen2.5-VL-72B versus other vision-language models on diverse tasks.
The model demonstrates noteworthy results for document parsing, general visual question answering, multilingual OCR, and agent benchmarks. Evaluations indicate robust video reasoning and generalization to multiple languages and domains. Performance on challenging complex-problem sets, such as MMMU, remains a focus for future improvement.
Applications and Use Cases
Qwen2.5-VL 72B addresses a set of practical demands:
- Document Analysis: Extracts, parses, and structures data from invoices, receipts, forms, and technical diagrams, supporting business and financial workflows.
- Visual Question Answering: Answers queries about images and videos, recognizes and localizes objects, and responds to prompts integrating both textual and visual clues.
- Multilingual OCR: Processes texts embedded in images across major Asian and European languages under various orientations.
- Video Summarization: Understands lengthy video footage, locating and describing key events at a fine temporal resolution.
- Visual Agents: Functions as a digital agent for computer or phone operations, automating UI tasks, application management, and tool use in interactive settings.
Limitations and Model Family
While Qwen2.5-VL-72B offers multimodal capabilities, there are documented limitations. The model’s performance can be affected by out-of-distribution small images in OCR tasks, and its handling of extended text inputs beyond default context limits may reduce temporal or spatial localization fidelity, as noted in the Qwen2.5-VL Technical Report. Video URL compatibility is also subject to backend library constraints.
Within the Qwen2.5-VL family, smaller variants such as Qwen2.5-VL 3B and Qwen2.5-VL 7B provide resource-efficient options. Comparisons to the precursor Qwen2-VL series and experimental models such as QvQ-72B-Preview show ongoing development in visual reasoning and fine-grained multimodal alignment, as reported on the Qwen2.5-VL GitHub.
Release, Licensing, and Resources
Qwen2.5-VL was announced on January 26, 2025, with technical reports, quantized models, and open-source materials following in subsequent months, according to the Qwen2.5 VL blog. The series is released under the Apache-2.0 license, suitable for research, development, and further innovation in vision-language AI.
Helpful Links
More in the Qwen 2 Family
Qwen2.5 VL 3B
Qwen2.5 VL 7B
QwQ 32B Preview
QwQ 32B
Qwen 2.5 Math 1.5B
DeepSeek R1 Distill Qwen 1.5B
DeepCoder 1.5B Preview
Qwen 2.5 Math 7B
Qwen 2.5 Math PRM 7B
DeepSeek R1 Distill Qwen 7B
Qwen 2.5 Math 72B
Qwen 2.5 Math PRM 72B
Qwen 2.5 Coder 7B
Qwen 2.5 Coder 32B
Qwen 2.5 7B
Qwen2.5 7B 1M
Qwen 2.5 14B
Qwen2.5 14B 1M
DeepSeek R1 Distill Qwen 14B
DeepCoder 14B Preview
Cogito V1 Preview 14B
Qwen 2.5 32B
DeepSeek R1 Distill Qwen 32B
Cogito V1 Preview 32B
Qwen 2.5 72B
Qwen 2 7B
Qwen 2 72B
More from Alibaba Cloud
Qwen3 0.6B
Qwen3 1.7B
Qwen3 4B
Qwen3 8B
Qwen3 14B
Qwen3 32B
Qwen3 30B A3B
Qwen3 235B A22B
Qwen 1.5 32B
Qwen 1.5 72B
Compatible Apps

Open WebUI
A polished, self-hosted chat interface for LLMs with Ollama integration, multimodal prompts, and extensive workspace customization.
Web UI
Chat UIs · Beginner Friendly
llama.cpp
GGUF model inference with a polished web UI and OpenAI-format API. CPU-only build — high hardware compatibility, works on any machine without a GPU.
Web UI · API · CLI
LLM Inference · Chat UIs
llama.cpp (CUDA)
GGUF model inference with a polished web UI and OpenAI-format API. CUDA build — GPU-accelerated for NVIDIA GPUs.
Web UI · API · CLI
LLM Inference · Chat UIs

Text Generation Web UI
A feature-rich interface for running and experimenting with open-weight LLMs, including multiple inference backends, plugins, and tuning controls.
Web UI · API
Chat UIs · LLM Inference