Skip to main content
Browse Models

Deep Cogito

Cogito V1 Preview 32B

Released

2025-03-31

Family

Qwen 2

Type

Fine-Tuned Model

Downloads

External Download

You are about to open a link to an external source. Verify the URL before continuing.

Download

4-bit GGUF (Q4_K_M)

GGUF · deepcogito_cogito-v1-preview-qwen-32B-Q4_K_M.gguf

5-bit GGUF (Q5_K_M)

GGUF · deepcogito_cogito-v1-preview-qwen-32B-Q5_K_M.gguf

6-bit GGUF (Q6_K)

GGUF · deepcogito_cogito-v1-preview-qwen-32B-Q6_K.gguf

8-bit GGUF (Q8_0)

GGUF · deepcogito_cogito-v1-preview-qwen-32B-Q8_0.gguf

Ollama Model (q4_K_M)

Ollama

Model Report

Overview

Cogito V1 Preview 32B is a large language model (LLM) developed by Deep Cogito, representing the 32-billion-parameter variant in a series of generative models designed for advanced language understanding and reasoning. The model was released on April 8, 2025, as part of the broader Cogito family, which encompasses models ranging from 3B up to 70B parameters. Cogito V1 Preview 32B is positioned as a research-driven system that integrates new approaches to alignment and self-improvement within contemporary transformer-based neural architectures.

Benchmarking chart for Cogito 32B and comparison models

Figure 1. Benchmark comparison of Cogito 32B (in standard and thinking modes) against Qwen2.5 32B and Qwen QwQ 32B across general, math, and multilingual tasks.

Model Architecture and Alignment Strategy

Cogito V1 Preview 32B is built upon the Qwen2.5 32B architecture, employing a transformer-based neural network with 32.8 billion parameters. The model extends the capabilities of its Qwen lineage by introducing a unique alignment strategy known as Iterated Distillation and Amplification (IDA). IDA aims to systematically improve the reasoning capacities of the model through alternating cycles of "amplification"—where more complex subroutines are executed to achieve higher-level reasoning—and "distillation," which transfers these improved reasoning strategies into the model's own weights.

This approach contrasts with more conventional alignment techniques such as reinforcement learning from human feedback (RLHF) and model distillation from larger systems. IDA emphasizes scalable self-improvement while reducing dependency on expensive human or computational oversight, supporting efficient refinement of the model's internal thought processes.

Technical Features and Modes of Operation

Cogito V1 Preview 32B is an instruction-tuned, text-in/text-out model optimized for a variety of advanced applications, including code generation, function calling, and autonomous agentic behaviors. Notably, the model offers two distinct operational modes: the standard mode, which provides direct responses, and a specialized "reasoning mode" (or "thinking mode") that engages in explicit self-reflection before generating output. Reasoning mode can be invoked by appending a directive such as "Enable deep thinking subroutine." to the system prompt or by enabling a flag in the tokenizer's chat template method.

Multilingual proficiency is integral to the design, as the model is trained on data across more than 30 languages. Additionally, Cogito 32B supports a long context window of up to 128,000 tokens, enhancing its ability to manage lengthy documents and complex conversational states.

A distinguishing feature is the model's native support for tool calling across both its standard and extended reasoning modes. This allows for seamless integration of external tools via single, parallel, or multiple calls, and enables flexible composition of outputs with tool results—beyond what is typically present in other models of comparable size.

Training Procedures and Data Management

Training for Cogito V1 Preview 32B commences with existing base checkpoints from Qwen2.5 32B and, in some cases, Llama models, upon which the IDA methodology is applied. Throughout this iterative process, expensive cognitive subroutines are learned and consolidated, improving each successive checkpoint's ability to handle complex reasoning. Evaluation benchmarks used in development are rigorously excluded from training material by using string-matching methods to prevent data contamination and maintain the validity of performance assessments.

The distillation phase internalizes external reasoning subroutines, leading to a feedback loop of iterative improvement. This contrasts with more static pretraining and fine-tuning paradigms and is designed to make the model's "thinking process" more efficient and scalable over time.

Table comparing benchmark results for Cogito 32B and peers

Figure 2. Performance data comparing Cogito 32B (standard and thinking) to Qwen2.5 32B and Qwen QwQ 32B across non-reasoning and reasoning benchmarks (including MMLU, Math, and multilingual tasks).

Performance Benchmarks and Model Evaluation

Cogito V1 Preview 32B demonstrates competitive results on standard language model benchmarks, including MMLU, MMLU-Pro, GSM8K, MATH, MMMLU, and MGSM. In published comparisons against other major open-source models—such as Qwen2.5 32B, Qwen QwQ 32B, as well as Llama3 and DeepSeek models—the Cogito 32B model typically matches or outperforms alternatives in tasks spanning general language understanding, mathematics, and multilingual reasoning.

The model's reasoning mode ("Cogito 32B Thinking") further enhances its ability to solve tasks that require more complex sequential deduction, a feature illustrated in side-by-side benchmark tables. For example, relative improvements over Qwen2.5 32B are observed across several key evaluation metrics, including global averages computed from diverse test sets. Native tool calling capabilities are cited as a factor contributing to gains over models like Llama, where such features are less comprehensively integrated.

It is noted by Deep Cogito that while benchmark metrics provide a useful reference, their correspondence with end-user experience is inherently limited. Real-world utility, especially in dynamic or multi-turn dialogues, may diverge from results indicated by static benchmarks.

Use Cases and Model Applications

Cogito V1 Preview 32B is engineered for a broad spectrum of practical applications. Coding and code comprehension are primary targets, supported by robust capabilities in natural language instruction following, general agentic behavior, and complex function or tool invocation. The model finds utility in STEM disciplines as well, benefiting from its extensive multilingual training and its long-context capacity.

Typical interaction scenarios include multi-step reasoning tasks, data analysis pipelines incorporating tool calls, and high-fidelity language generation in multilingual environments. The reasoning mode is especially suitable for tasks where explicit self-reflection enhances output quality or traceability.

Limitations and Model License

Despite its advancements, Cogito V1 Preview 32B is not tailored for extremely long or deeply recursive reasoning chains, a tradeoff intended to balance model efficiency with practical wait times and to streamline the distillation process. Benchmark accuracy is acknowledged as an imperfect proxy for real-world performance, and results may vary based on deployment context.

The model is distributed under the Apache 2.0 License, permitting wide commercial and research use.

Related Models and Future Directions

The Cogito V1 Preview 32B is part of a scalable family with available checkpoints at 3B, 8B, 14B, 32B, and 70B parameters. According to development roadmaps, larger model variants—such as those in the 109B to over 600B parameter range—are under active development. Each release leverages the IDA framework to incrementally improve intelligence and alignment.

Helpful Links

About Qwen 2: Qwen 2 (and 2.5) is a family of advanced AI models developed by Alibaba, designed to excel in various tasks including general language understanding, coding, and mathematics.

More in the Qwen 2 Family

Alibaba Cloud /

Qwen2.5 VL 3B

A 3B-parameter multimodal language model that processes images, videos, and text with capabilities for visual reasoning, document analysis, and computer interface interaction.
Alibaba Cloud /

Qwen2.5 VL 7B

A 7-billion parameter multimodal model capable of processing images, documents, and videos with native dynamic resolution support and structured output capabilities.
Alibaba Cloud /

Qwen2.5 VL 72B

A 72-billion parameter multimodal model that processes images, videos, and text for document analysis, object detection, and visual agent tasks.
Alibaba Cloud /

QwQ 32B Preview

Experimental research model focused on advancing AI reasoning capabilities by thinking through problems step by step, continuously questioning assumptions, and exploring different paths of thought. Impressive benchmark scores in math and coding tasks.
Alibaba Cloud /

QwQ 32B

A 32.5-billion parameter transformer model optimized for mathematical reasoning, coding, and complex problem-solving through reinforcement learning training.
Alibaba Cloud /

Qwen 2.5 Math 1.5B

A 1.5 billion parameter specialized model focused on mathematical reasoning and problem-solving capabilities in English and Chinese.
Deepseek AI /

DeepSeek R1 Distill Qwen 1.5B

A 1.5 billion parameter language model created through distillation techniques, focusing on mathematical reasoning and chain-of-thought problem-solving capabilities.
Agentica /

DeepCoder 1.5B Preview

A 1.5B parameter code generation model fine-tuned with reinforcement learning and iterative context lengthening for long-context reasoning tasks.
Alibaba Cloud /

Qwen 2.5 Math 7B

A 7.62-billion parameter mathematical reasoning model supporting English and Chinese with chain-of-thought and tool-integrated reasoning capabilities.
Alibaba Cloud /

Qwen 2.5 Math PRM 7B

A 7B parameter process reward model that evaluates mathematical reasoning steps individually rather than just final answers, achieving 67.6% accuracy on mathematical benchmarks.
Deepseek AI /

DeepSeek R1 Distill Qwen 7B

A 7.62B parameter distilled language model based on Qwen2.5-Math-7B, trained via knowledge distillation for mathematical and logical reasoning tasks.
Alibaba Cloud /

Qwen 2.5 Math 72B

A 72.7 billion parameter language model specialized for solving mathematical problems in English and Chinese using chain-of-thought and tool-integrated reasoning.
Alibaba Cloud /

Qwen 2.5 Math PRM 72B

A 72.8 billion parameter process reward model that evaluates intermediate mathematical reasoning steps using consensus filtering and achieves 78.3% F1 on ProcessBench.
Alibaba Cloud /

Qwen 2.5 Coder 7B

A 7.61 billion parameter transformer model designed for code generation and reasoning across 92 programming languages with a 128,000-token context window.
Alibaba Cloud /

Qwen 2.5 Coder 32B

Qwen2.5-Coder builds on Qwen2.5 by training on 5.5 trillion additional tokens of code data including source code, text-code grounding data, and synthetic data. This leads to significant improvements in code-related tasks.
Alibaba Cloud /

Qwen 2.5 7B

A 7.61-billion parameter transformer-based language model with 128K token context length, trained on 18 trillion tokens supporting 29+ languages.
Alibaba Cloud /

Qwen2.5 7B 1M

A 7-billion parameter transformer model with 1-million token context capacity utilizing Dual Chunk Attention for long-range text processing.
Alibaba Cloud /

Qwen 2.5 14B

A 14.7 billion parameter transformer model supporting 29 languages with 128K context window, designed as foundation for fine-tuning applications.
Alibaba Cloud /

Qwen2.5 14B 1M

A 14.7B parameter transformer model with 1-million token context capacity, utilizing dual chunk attention and progressive training for extended sequence processing.
Deepseek AI /

DeepSeek R1 Distill Qwen 14B

A 14B dense language model distilled from a mixture-of-experts architecture, optimized for mathematical reasoning and code generation tasks.
Agentica /

DeepCoder 14B Preview

A 14-billion parameter code reasoning model fine-tuned using distributed reinforcement learning with long-context capabilities up to 64,000 tokens.
Deep Cogito /

Cogito V1 Preview 14B

A 14.8 billion parameter instruction-tuned language model trained using Iterated Distillation and Amplification with hybrid reasoning capabilities across 30+ languages.
Alibaba Cloud /

Qwen 2.5 32B

A 32.5 billion parameter multilingual transformer model trained on 18 trillion tokens with 128K context length supporting text generation, coding, and mathematical reasoning.
Deepseek AI /

DeepSeek R1 Distill Qwen 32B

A 32B parameter language model created through knowledge distillation, optimized for mathematical reasoning, code generation, and complex problem-solving tasks.
Alibaba Cloud /

Qwen 2.5 72B

This update to Qwen 2, pretrained on 18 trillion tokens, shows impressive performance in benchmarks (MMLU 85+, HumanEval 85+, MATH 80+), outperforming models of similar size. Qwen 2.5 has multilingual support for over 29 languages and can understand and generate structured data.
Alibaba Cloud /

Qwen 2 7B

A 7.6 billion parameter multilingual decoder-only Transformer model designed for further post-training with 32,000-token context and strong coding capabilities.
Alibaba Cloud /

Qwen 2 72B

A 72-billion parameter transformer model supporting extended context lengths up to 128K tokens and demonstrating strong multilingual capabilities across diverse benchmarks.