Skip to main content
Browse Models

Deepseek AI

DeepSeek VL2 Small

Released

2024-12-13

Family

DeepSeek VL2

Type

Foundation Model

Model Report

Overview

DeepSeek-VL2 model performance and scaling chart

Figure 1. Performance chart positioning DeepSeek-VL2 model variants (including Small) alongside alternative multimodal models as a function of activated parameters. DeepSeek-VL2 models are highlighted, illustrating relative performance scaling across model sizes.

DeepSeek-VL2-Small is part of the DeepSeek-VL2 family of large Mixture-of-Experts (MoE) vision-language models developed by DeepSeek-AI. Designed for multimodal understanding, the VL2 series builds upon earlier iterations to provide performance in tasks such as visual question answering, optical character recognition (OCR), document and chart analysis, and visual grounding, with a focus on computational efficiency and scalability. The DeepSeek-VL2-Small model comprises 2.8 billion activated parameters and leverages the DeepSeekMoE-16B architecture, contributing to its performance in both language and vision domains according to the official technical report.

Architecture and Innovations

DeepSeek-VL2 models, including the Small variant, follow a LLaVA-style modular architecture comprising a vision encoder, a vision-language adaptor, and a Mixture-of-Experts language model. The vision encoder is based on SigLIP-SO400M-384, responsible for extracting image features. These features pass through a two-layer multilayer perceptron (MLP) that projects the visual data into the activation space used by the language model. The core language model, DeepSeekMoE, employs a sparse mixture-of-experts framework optimized with Multi-head Latent Attention (MLA), which compresses key-value caches into compact latent representations for more efficient inference.

A key innovation is the dynamic tiling vision encoding strategy, which segments large and high-resolution images into local tiles (384x384 pixels) and a global thumbnail. This enables processing of varying aspect ratios and ultra-high-resolution content. Each tile undergoes separate encoding, followed by a 2x2 pixel shuffle operation to reduce the token length, which contributes to maintaining a manageable computational footprint. When more than two images are input, the dynamic tiling is disabled to prioritize throughput and context length.

The MoE framework within DeepSeek-VL2-Small uses 64 routed experts and 2 shared experts, employing a softmax-based gating mechanism with Top-K routing. This allows the model to dynamically select a subset of experts per input, facilitating both scalability and specialization among expert modules.

Training Data and Methodology

Training of DeepSeek-VL2-Small involves a three-stage pipeline. The initial phase centers on vision-language alignment, using datasets such as ShareGPT4V to optimize the vision-language connector while keeping the language model parameters frozen. This stage focuses on building cross-modal bridges.

The second pretraining phase expands to joint optimization of all model components, employing an interleaved mixture of vision-language and text-only samples. The vision-language data includes sources like WIT, WikiHow, and portions of OBELICS for broad coverage, as well as curated Chinese content for multilingual performance. Specialized datasets target tasks such as OCR (e.g., LaTeX OCR, RenderedText), table and chart understanding (e.g., PubTabNet, FinTabNet), and visual grounding.

Supervised fine-tuning constitutes the final stage, where instruction-following and conversational abilities are refined using a comprehensive in-house dataset encompassing question-answer pairs, document understanding dialogues, table-based questions, reasoning, and visual grounding. This stage supports the model's responses in terms of contextual relevance and accuracy, particularly in multilingual scenarios.

Performance and Evaluation

DeepSeek-VL2-Small demonstrates performance on a range of established multimodal benchmarks, particularly in OCR and document understanding, general visual question-answering, and visual grounding tasks. On the DocVQA benchmark for document comprehension, DeepSeek-VL2-Small achieves a score of 92.3, outperforming comparably sized alternatives such as InternVL2-2B and Qwen2-VL-2B, as reported in the official benchmark tables. In visual reasoning and grounding, results on datasets like RefCOCO and RefCOCOg also indicate competitive accuracy, with scores exceeding 90 on key test sets.

The model supports sequence lengths up to 4096 tokens, permitting the integration of conversations and multiple image contexts. Visual grounding capabilities are specifically enhanced, enabling both object localization from textual prompts and in-context grounding that requires cross-referencing between images and regions of interest.

Applications and Limitations

Typical application domains for DeepSeek-VL2-Small include visual question answering, dense captioning, chart and table analysis, and GUI perception, as detailed in the DeepSeek-VL2 documentation. It is also suited for tasks like visual reasoning, meme understanding, and multi-image conversational settings, where grounded responses referencing visual content are essential.

Certain limitations remain. The current version limits context to a small number of images per conversation, which restricts complex multi-image tasks. Some challenges persist in handling low-quality or unseen content and in advanced reasoning or creative storytelling. The basic demonstration interface is not yet fully optimized for high-throughput deployment, suggesting considerations for production use.

Model Availability, Licensing, and Resources

DeepSeek-VL2-Small is released under the MIT License, with use governed by the DeepSeek Model License. Commercial utilization is supported. Model checkpoints, pre-trained weights, code, and detailed documentation are available on the official DeepSeek-AI GitHub repository and the Hugging Face model hub.

Helpful Links

About DeepSeek VL2: DeepSeek-VL2 is an advanced series of Mixture-of-Experts (MoE) Vision-Language Models that enhance multimodal understanding by efficiently activating between 1.0B to 4.5B parameters per task, achieving state-of-the-art performance across various vision-language tasks.

More from Deepseek AI

Deepseek AI /

DeepSeek R1 Distill Qwen 1.5B

A 1.5 billion parameter language model created through distillation techniques, focusing on mathematical reasoning and chain-of-thought problem-solving capabilities.
Deepseek AI /

DeepSeek R1 Distill Qwen 7B

A 7.62B parameter distilled language model based on Qwen2.5-Math-7B, trained via knowledge distillation for mathematical and logical reasoning tasks.
Deepseek AI /

DeepSeek R1 Distill Qwen 14B

A 14B dense language model distilled from a mixture-of-experts architecture, optimized for mathematical reasoning and code generation tasks.
Deepseek AI /

DeepSeek R1 Distill Qwen 32B

A 32B parameter language model created through knowledge distillation, optimized for mathematical reasoning, code generation, and complex problem-solving tasks.
Deepseek AI /

DeepSeek R1 Distill Llama 8B

Distilled 8B-parameter model optimized for mathematical reasoning and code generation through knowledge transfer from larger reinforcement learning-trained teacher models.
Deepseek AI /

DeepSeek R1 Distill Llama 70B

A 70B parameter dense language model distilled from DeepSeek-R1 using Llama 3.3 architecture, optimized for mathematical and coding reasoning tasks.
Deepseek AI /

DeepSeek R1 (0528)

A 671B-parameter MoE model with 37B active parameters featuring enhanced reasoning capabilities through reinforcement learning and chain-of-thought training methodologies.
Deepseek AI /

DeepSeek R1

A 671B parameter Mixture-of-Experts model trained with reinforcement learning to enhance reasoning capabilities in mathematics, coding, and logical tasks.
Deepseek AI /

DeepSeek V3 (0324)

Large-scale MoE language model utilizing 671B parameters with 37B activated per token, featuring enhanced reasoning and multilingual capabilities.
Deepseek AI /

DeepSeek V3

A 671-billion parameter Mixture-of-Experts language model with 37 billion active parameters per token, featuring auxiliary-loss-free load balancing and FP8 mixed-precision training.
Deepseek AI /

DeepSeek V2.5

A 236-billion parameter mixture-of-experts language model with multi-head latent attention, activating 21 billion parameters per token for bilingual text generation.
Deepseek AI /

DeepSeek V2

A 236-billion parameter Mixture-of-Experts language model that activates only 21 billion parameters per token for efficient multilingual text generation.
Deepseek AI /

DeepSeek Coder V2

Open-source Mixture-of-Experts model with 236B total parameters specialized for code generation, mathematical reasoning, and programming across 338 languages.
Deepseek AI /

DeepSeek Coder V2 Lite

A 16B parameter Mixture-of-Experts model designed for code generation, completion, and reasoning across 338 programming languages with 128K token context length.