Mistral AI
Mistral Small 3.1 (2503)
Downloads
Model Report
Overview
Mistral Small 3.1 (2503) is a large-scale, open-source generative AI model developed by Mistral AI and released under the Apache 2.0 license. As a multimodal and multilingual model, Mistral Small 3.1 is designed to deliver advanced performance across both text and visual inputs. Building upon its predecessor, Mistral Small 3, this iteration introduces improvements in text generation accuracy, visual understanding, and contextual reasoning. The model is available in both base and instruction-tuned versions, supporting a broad spectrum of applications that require high-quality language understanding and image analysis.

Figure 1. Performance scatter plot comparing Mistral Small 3.1 with Gemma 3-it, GPT-4o Mini, and Claude-3.5 Haiku. The plot highlights Mistral Small 3.1's superior GPQA-Diamond scores combined with lower latency.
Model Architecture and Capabilities
Mistral Small 3.1 is a transformer-based model comprising 24 billion parameters, offered in both a pretrained base variant and an instruction-finetuned version. The model utilizes the Tekken tokenizer, featuring a vocabulary size of 131,000 tokens, and supports input sequences up to 128,000 tokens, enabling comprehension of long documents and complex conversational contexts.
As a multimodal model, Mistral Small 3.1 is proficient in processing both textual and visual data. Its vision system is capable of detailed analysis, including image-based document classification, content extraction, and scene description. The model's multilingual proficiency spans dozens of languages, including major European, Asian, and Middle Eastern languages, making it suitable for global-scale deployments. Further, it features advanced function-calling and agent-centric capabilities, facilitating structured outputs such as JSON for downstream automation and workflow integration.

Figure 2. Output from Mistral Small 3.1 analyzing a political map of Europe, demonstrating country identification, color parsing, and city recognition from visual data.
Performance and Benchmarking
Mistral Small 3.1 exhibits competitive performance on a broad array of benchmarks when compared with both open and proprietary models in a similar parameter range. On academic evaluation suites, such as MMLU (Massive Multitask Language Understanding), GPQA (Graduate Level Question Answering), and multilingual tests, it performs at or above the level of leading models like Gemma 3-it (27B), GPT-4o Mini, and Claude 3.5 Haiku.
The model demonstrates strong results on both general and specialized tasks. For instance, its instruction-tuned version achieves 80.6% on standard MMLU, 44.4% on GPQA Main (5-shot CoT), and 64.0% on MMMU for multimodal instruction. In multilingual settings, it averages 71.2% accuracy across diverse language groupings. Its long-context reasoning capabilities are reflected in high scores on the LongBench v2 and RULER benchmarks, where it outperforms comparable models in maintaining accuracy over extended sequences. Additionally, the model delivers low inference latency, supporting high-throughput and responsive applications even in resource-constrained environments, as illustrated by public benchmark analyses.
Training Data and Methodology
While specific details regarding the training data composition remain undisclosed, Mistral Small 3.1 is reported to have built upon the methodologies established in Mistral Small 3, employing large-scale web, scientific, and technical datasets to support its broad reasoning and multilingual abilities. The instruction-tuned variant is further refined to follow complex system prompts and user instructions accurately, using high-quality supervised and reinforcement learning-based strategies to enhance alignment and safety. The vision capabilities are achieved through integration of image-text pretraining, enabling robust interpretation of a wide range of document and natural images.
Applications and Use Cases
The versatility of Mistral Small 3.1 allows deployment across a spectrum of practical applications. Its enhanced instruction-following skills make it suitable as a conversational assistant, supporting dialogue in multiple languages with context persistence over long exchanges. The model's image understanding enables automated document verification, technical diagnostics, quality inspection, and visual customer service scenarios. It is well-suited for agentic deployments requiring on-the-fly decision-making, such as executing structured function calls or integrating with data platforms via JSON outputs.
Furthermore, the model supports domain-specific fine-tuning, enabling its adaptation for specialized subject matter expertise—including legal and medical advisory services and technical troubleshooting—while still maintaining strong foundational reasoning skills. Its efficiency and lightweight design both facilitate deployment in local and edge environments, which is advantageous for privacy-sensitive or latency-critical use cases.
Limitations and Licensing
Despite broad capabilities, Mistral Small 3.1 has certain constraints. The model cannot generate images, access the internet, or transcribe audio and video inputs. Additionally, while it is available in a Transformers-compatible format, optimal operation is recommended using the original weight format, as complete behavioral parity with Transformers-based implementations has not yet been guaranteed.
Mistral Small 3.1 is made available under the Apache 2.0 license, permitting commercial and non-commercial use, modification, and redistribution. This supports a wide range of research and enterprise use cases, promoting transparency and flexibility for adopters.
Related Models in the Mistral Family
Mistral Small 3.1 extends the Mistral family of models, following the release of Mistral Small 3. The Mistral suite serves as the foundation for several derivative models developed by the community, such as DeepHermes 24B by Nous Research, which targets advanced reasoning and instruction following. The continued development of the Mistral series reflects a commitment to accessible, high-performance open-source AI for research and production environments.
External Resources
- Mistral Small 3.1 Base on Hugging Face
- Mistral Small 3.1 Instruct on Hugging Face
- Official Mistral Small 3.1 announcement
- Prior model: Mistral Small 3
- Mistral Common GitHub repository
- vLLM inference library documentation
- DeepHermes 24B by Nous Research
- Hugging Face task documentation: Image-text-to-text
- Apache 2.0 license terms
More in the Mistral Family
Mistral Large 2
Behemoth 123B v1.2
Mistral Small (2409)
Mistral Small 3.2 (2506)
Harbinger 24B
Devstral Small 1.0
Mistral Small 3 (2501)
Cydonia 24B v2
Dolphin 3.0 Mistral 24B
Mistral NeMo 12B
Rocinante 12B v1.1
More from Mistral AI
Mistral 7B
Codestral 22B v0.1
Mixtral 8x7B
Mixtral 8x22B
Compatible Apps

Open WebUI
A polished, self-hosted chat interface for LLMs with Ollama integration, multimodal prompts, and extensive workspace customization.
Web UI
Chat UIs · Beginner Friendly
llama.cpp
GGUF model inference with a polished web UI and OpenAI-format API. CPU-only build — high hardware compatibility, works on any machine without a GPU.
Web UI · API · CLI
LLM Inference · Chat UIs
llama.cpp (CUDA)
GGUF model inference with a polished web UI and OpenAI-format API. CUDA build — GPU-accelerated for NVIDIA GPUs.
Web UI · API · CLI
LLM Inference · Chat UIs

Text Generation Web UI
A feature-rich interface for running and experimenting with open-weight LLMs, including multiple inference backends, plugins, and tuning controls.
Web UI · API
Chat UIs · LLM Inference