Skip to main content
Browse Models

Mistral AI

Mistral Large 2

Released

2024-07-24

Family

Mistral

Type

Foundation Model

Downloads

External Download

You are about to open a link to an external source. Verify the URL before continuing.

Download

Ollama Model (q4_K_M)

Ollama

Instruct Model - 4 BPW EXL2

EXL2

Model Report

Overview

Mistral Large 2 is a dense generative large language model (LLM) developed by Mistral AI and released in July 2024. Designed for general-purpose language modeling, Mistral Large 2 is characterized by an extensive parameter count, a substantial context window, comprehensive multilingual abilities, and high performance in tasks involving code generation and mathematical reasoning. This model is a subsequent offering from Mistral AI, with observed improvements in accuracy, efficiency, and agentic capabilities compared to its predecessors, while remaining publicly available under a dedicated research license as outlined by Mistral AI.

Scatter plots benchmarking code and math performance

Figure 1. Comparison of code and math performance for Mistral Large 2 and peer models, showing performance relative to parameter count.

Model Architecture and Training

Mistral Large 2 is constructed as a dense transformer-based language model, comprising 123 billion parameters. It features a context window of 128,000 tokens, enabling it to handle extended documents and complex dialogue scenarios. The model was trained on a large proportion of multilingual texts and a substantial corpus of source code, leveraging advancements in both data curation and architecture engineering, as described in the official release blog.

Mistral Large 2’s architecture is optimized for single-node inference, supporting large-throughput operations. Special attention was given to mitigating hallucination, enhancing mathematical reasoning, and instruct-following performance through targeted fine-tuning regimens. Training data emphasized diversity in language and code domain coverage, supporting both general and specialized tasks.

Multilingual and Coding Capabilities

One of the defining characteristics of Mistral Large 2 is its robust multilingual proficiency. The model supports dozens of languages, including but not limited to English, French, German, Spanish, Italian, Dutch, Russian, Chinese, Japanese, Korean, Arabic, and Hindi. Performance on multilingual evaluation benchmarks such as MMLU demonstrates strong results across a wide range of languages, positioning the model as a versatile tool for global applications. This multilingual support is highlighted in direct comparisons with other large open models, as shown in benchmarking data from the Mistral Large 2 release.

Multilingual MMLU benchmarking plot

Figure 2. Performance of Mistral Large 2 on multilingual MMLU benchmarks, illustrating efficiency relative to model size.

In addition, Mistral Large 2 exhibits high proficiency in code-related tasks. Trained on at least 80 distinct programming languages—including Python, Java, C++, JavaScript, Bash, Swift, and Fortran—it demonstrates high capabilities in both code generation and completion, consistently outperforming prior Mistral models and performing competitively with contemporary models in code-focused benchmarks.

Bar chart of code generation benchmark scores

Figure 3. Code generation accuracy across different models, with Mistral Large 2 performing competitively on Human Eval and MBPP benchmarks.

Code accuracy benchmark table for multiple programming languages

Figure 4. Table comparing Mistral Large 2’s MultiPL-E benchmark scores across major programming languages, highlighting broad proficiency.

Performance and Benchmarks

Mistral Large 2 demonstrates performance across a variety of standardized language modeling tasks, as evidenced by publicly reported evaluations in code generation, mathematical reasoning, general instruction following, and multilingual benchmarks.

For general knowledge and logic, the model attains an MMLU accuracy of 84%. It positions itself relative to competing open models. On code-generation tasks, it reports 92% on the HumanEval benchmark and high results on HumanEval Plus and MBPP datasets. In mathematical reasoning, Mistral Large 2’s GSM8K score is 93%, and it achieves high marks on zero- and few-shot problem-solving tasks.

Math benchmarks bar chart

Figure 5. Comparison of Mistral Large 2 and other models on GSM8K and Math Instruct benchmarks for mathematical reasoning.

Instruction tuning has resulted in high alignment and conversational capacities. The model’s WildBench and Arena Hard results reflect its capacity for robust handling of open-ended dialogue and challenging prompts.

Instruction following benchmark chart

Figure 6. Performance on alignment and instruction benchmarks Wild Bench and Arena Hard, demonstrating Mistral Large 2’s conversational capabilities.

Mistral Large 2 also introduces refinements in agentic and function-calling abilities. The model is equipped for both parallel and sequential native function execution, and is capable of outputting structured responses in JSON format. Its accuracy on function-calling tasks is on par with other models, facilitating use in complex retrieval and integration scenarios.

Function-calling benchmark bar chart

Figure 7. Function calling benchmark accuracy for Mistral Large 2 and peer models, indicating strong agentic capabilities.

The design also emphasizes concise outputs. Benchmark results from MT Bench reveal that Mistral Large 2 produces succinct and relevant responses, an attribute useful for enterprise and high-throughput applications.

MT Bench and output length bar charts

Figure 8. Left: MT Bench evaluation by GPT-4o, showing high score for Mistral Large 2. Right: average generation length chart, illustrating the model’s tendency for concise outputs.

Applications and Use Cases

The advanced feature set of Mistral Large 2 enables a wide range of applications. Its strong coding skills are leveraged in software engineering and data analysis workflows, while mathematical reasoning supports domains such as scientific research and education. General-purpose language capabilities facilitate summarization, translation, knowledge retrieval, content generation, and instructional dialogue across multiple languages.

The model’s support for agentic workflows—native function calling and complex reasoning—enables integration into business process automation, virtual assistants, and research platforms. Function calling and structured JSON output are particularly oriented towards use in enterprise data pipelines and interactive applications.

Limitations and Licensing

Mistral Large 2, as released, does not include integrated moderation mechanisms. The developers encourage community engagement to establish appropriate guardrails for use in environments requiring output moderation, as stated in the official release materials.

The model is distributed under the Mistral AI Research License, which authorizes non-commercial research, personal, or academic use. Commercial usage, including commercial deployment or business-related development, requires a dedicated commercial license from Mistral AI. The license clarifies attribution, derivative works, distribution requirements, and provides no warranty. Users are solely responsible for content generated using the model, and must not imply endorsement by Mistral AI.

Model Availability and Ecosystem

Mistral Large 2 is available for research purposes with downloadable weights and integration support across frameworks such as transformers and mistral_inference.

Within the broader Mistral model family, Mistral Large 2 builds on techniques pioneered in earlier models, such as Codestral and Mistral Nemo, and consolidates previous innovations into a single advanced architecture. Existing specialist models, general-purpose models, and fine-tuning workflows remain accessible for non-commercial research and are progressively updated in line with new releases.

External Resources

For additional details and the most current resources, refer to the following:

About Mistral: The Mistral family of AI models, developed by Paris-based Mistral AI, includes the original 2023 Mistral 7B release, as well as the more recent Mistral Small, Nemo, and Large weights.

More in the Mistral Family

TheDrummer /

Behemoth 123B v1.2

A 123-billion parameter language model optimized for conversational AI, creative prose generation, and role-playing applications with enhanced narrative consistency.
Mistral AI /

Mistral Small (2409)

A 22B parameter enterprise-grade small model, a convenient mid-point between Mistral NeMo 12B and Mistral Large 2. This version delivers significant improvements in human alignment, reasoning capabilities, and code over the previous version.
Mistral AI /

Mistral Small 3.2 (2506)

A 24-billion parameter multimodal model featuring improved instruction following, function calling, and reduced repetition over its predecessor.
Mistral AI /

Mistral Small 3.1 (2503)

A 24-billion parameter multimodal transformer supporting text and vision tasks with 128K token context length under Apache 2.0 license.
LatitudeGames /

Harbinger 24B

A 24-billion parameter language model fine-tuned on Mistral Small 3.1 Instruct, specialized for interactive storytelling and text-based adventures.
Mistral AI /

Devstral Small 1.0

A 23.6B parameter coding assistant finetuned for agentic software engineering tasks with 128K context window and 46.8% SWE-Bench performance.
Mistral AI /

Mistral Small 3 (2501)

A 24-billion parameter instruction-tuned language model with multilingual capabilities, 32K context window, and optimized low-latency inference performance.
TheDrummer /

Cydonia 24B v2

A fine-tuned 23.6 billion parameter Mistral-based model designed for long-context conversations and maintaining narrative coherence across extended dialogues.
Cognitive Computations /

Dolphin 3.0 Mistral 24B

A 24-billion parameter instruction-tuned model built on Mistral architecture with deliberately removed content filters to maximize user control over outputs.
Mistral AI /

Mistral NeMo 12B

A 12B parameter multi-lingual model that supports function calling built in collaboration with NVIDIA and trained using the new Tekken tokenizer. By some metrics, it is state-of-the-art in its size category. NeMo was trained with quantisation awareness, enabling FP8 inference without any performance loss.
TheDrummer /

Rocinante 12B v1.1

A 12.2 billion parameter text generation model optimized for creative storytelling, role-playing scenarios, and adventure-based interactive fiction applications.