Alibaba Cloud
Qwen 2.5 Math 1.5B
Downloads
Model Report
Overview
Qwen2.5-Math-1.5B is a specialized large language model designed for mathematical reasoning and problem-solving in both English and Chinese. Developed by the Qwen Team, Qwen2.5-Math-1.5B is part of the Qwen2.5-Math series, which was released in September 2024 as an enhancement to the earlier Qwen2-Math models. The series also includes larger models such as Qwen-2_5-math-7b and Qwen-2_5-math-72b, as well as instruction-tuned variants and a reward model for reinforcement learning. Qwen2.5-Math-1.5B is engineered with a focus on logic, symbolic manipulation, and bilingual mathematical processing, leveraging both chain-of-thought and tool-integrated reasoning paradigms.

Figure 1. Qwen2.5-Math-72B-Instruct performance on Zero-shot@1 accuracy on mathematical benchmarks in September 2024. The chart shows progression among open-weight and closed-source mathematical models over the course of 2024.
Model Architecture and Technical Innovations
Qwen2.5-Math-1.5B is built on the Qwen2.5 base architecture, benefiting from enhancements in language understanding, symbolic reasoning, and code generation over previous iterations. The parameter initialization for this model comes directly from the Qwen2.5 series, ensuring a robust foundation for mathematical tasks.
The development pipeline for Qwen2.5-Math incorporates several stages, including synthetic data generation using larger instruction-tuned models, aggregation of high-quality mathematical data (with a focus on expanding Chinese language resources), and specialized parameter tuning for mathematical domains. Model training maintains a context length of 4,096 tokens, offering sufficient capacity for complex, multi-step mathematical reasoning.

Figure 2. Flowchart depicting the specialization pipeline from Qwen2-Math to Qwen2.5-Math, illustrating stages such as data synthesis, supervised fine-tuning, reward model training, and policy optimization.
Training Corpus and Methodology
The pre-training of Qwen2.5-Math-1.5B utilizes the Qwen Math Corpus v2, comprising over one trillion tokens, which is a notable increase from the previous version's 700 billion. This corpus includes extensive mathematical content in both English and Chinese, collected from web data, educational resources, and curated code repositories. The training process further incorporates synthetic data generated by high-capacity Qwen2-Math models to enrich coverage and complexity.
A math-specific reward model, Qwen2.5-Math-RM-72B, is employed for constructing supervised fine-tuning data through rejection sampling and drives reinforcement learning post-SFT via Group Relative Policy Optimization. Task-specific instruction tuning introduces both chain-of-thought and tool-integrated reasoning data in English and Chinese.
Rigorous data decontamination protocols are enforced to prevent overlap between training and benchmark datasets. These include 13-gram text matching, normalization to remove extraneous symbols, and subsequence ratio thresholds. This ensures unbiased model evaluation, particularly on standardized mathematical datasets such as GSM8K, MATH, and various Chinese academic examinations.

Figure 3. Evaluation metrics for Qwen2.5-Math and prior models, showing performance across both English and Chinese mathematical benchmarks at equivalent model scales.
Reasoning Modes and Capabilities
Qwen2.5-Math-1.5B supports two principal reasoning modalities: Chain-of-Thought (CoT) and Tool-Integrated Reasoning (TIR). CoT enables stepwise natural language explanations, enhancing the interpretability and logical flow of mathematical solutions. Tool-Integrated Reasoning connects model outputs to a local or embedded Python interpreter, facilitating precise numerical computations and symbolic algebraic manipulation.
The TIR approach allows the model to solve complex mathematical questions by generating executable code. This yields accuracy on benchmark datasets and enables solutions to problems that are difficult to resolve solely through text-based reasoning.

Figure 4. Chatbot demonstration of Qwen2.5-Math's Tool-Integrated Reasoning: the model generates a solution plan, outputs Python code, and executes it for the answer.
A demo implementation within Qwen-Agent showcases this process, accepting code-execution requests from users and enabling local solution verification. Additionally, multimodal demos built on Qwen2-VL expand capabilities to process images or handwritten sketches of problems.
Performance and Evaluation
Qwen2.5-Math-1.5B demonstrates results across standard mathematical evaluation benchmarks in both English and Chinese environments. On the MATH benchmark, it achieves approximately 80% accuracy when utilizing the Python interpreter with TIR, and shows consistent gains over the prior Qwen2-Math series.

Figure 5. Performance of Qwen2.5-Math-Instruct and other models on English mathematical benchmarks under chain-of-thought and tool-integrated reasoning modalities.

Figure 6. Qwen2.5-Math-Instruct's performance compared against other models on Chinese mathematical benchmarks using both reasoning approaches.
In rigorous competition settings such as AIME 2024 and AMC 2023, the model solves a substantial portion of challenging problems. For instance, in CoT mode with reward aggregation, Qwen2.5-Math-1.5B-Instruct answers 29 out of 40 AMC 2023 questions. Larger variants within the Qwen2.5-Math series further improve these metrics, with performance comparable to several other models on both English and Chinese tasks.

Figure 7. Performance of Qwen2.5-Math-Instruct and other models on AIME 2024 and AMC 2023 benchmarks; results are grouped by reasoning strategy.

Figure 8. Scatter plot illustrating Qwen2.5-Math series models and peers, displaying MATH benchmark accuracy as a function of parameter scale.
Intended Usage and Limitations
Qwen2.5-Math-1.5B is specifically intended for mathematical reasoning tasks in English and Chinese. Its design is tailored for academic benchmarks, math education, and problem-solving involving symbolic computation and logical deduction. The model is not recommended for general-purpose language tasks outside mathematics, as highlighted in project documentation.
Both the base and instruction-tuned versions are provided: the base model is optimized for few-shot inference and fine-tuning, while the instruction-tuned model is suitable for interactive chat-like use. The recommended inference settings utilize the latest versions of popular transformer libraries for compatibility.
References and Further Reading
More in the Qwen 2 Family
Qwen2.5 VL 3B
Qwen2.5 VL 7B
Qwen2.5 VL 72B
QwQ 32B Preview
QwQ 32B
DeepSeek R1 Distill Qwen 1.5B
DeepCoder 1.5B Preview
Qwen 2.5 Math 7B
Qwen 2.5 Math PRM 7B
DeepSeek R1 Distill Qwen 7B
Qwen 2.5 Math 72B
Qwen 2.5 Math PRM 72B
Qwen 2.5 Coder 7B
Qwen 2.5 Coder 32B
Qwen 2.5 7B
Qwen2.5 7B 1M
Qwen 2.5 14B
Qwen2.5 14B 1M
DeepSeek R1 Distill Qwen 14B
DeepCoder 14B Preview
Cogito V1 Preview 14B
Qwen 2.5 32B
DeepSeek R1 Distill Qwen 32B
Cogito V1 Preview 32B
Qwen 2.5 72B
Qwen 2 7B
Qwen 2 72B
More from Alibaba Cloud
Qwen3 0.6B
Qwen3 1.7B
Qwen3 4B
Qwen3 8B
Qwen3 14B
Qwen3 32B
Qwen3 30B A3B
Qwen3 235B A22B
Qwen 1.5 32B
Qwen 1.5 72B
Compatible Apps

Open WebUI
A polished, self-hosted chat interface for LLMs with Ollama integration, multimodal prompts, and extensive workspace customization.
Web UI
Chat UIs · Beginner Friendly
llama.cpp
GGUF model inference with a polished web UI and OpenAI-format API. CPU-only build — high hardware compatibility, works on any machine without a GPU.
Web UI · API · CLI
LLM Inference · Chat UIs
llama.cpp (CUDA)
GGUF model inference with a polished web UI and OpenAI-format API. CUDA build — GPU-accelerated for NVIDIA GPUs.
Web UI · API · CLI
LLM Inference · Chat UIs

Text Generation Web UI
A feature-rich interface for running and experimenting with open-weight LLMs, including multiple inference backends, plugins, and tuning controls.
Web UI · API
Chat UIs · LLM Inference