Skip to main content
Browse Models

Deep Cogito

Cogito V1 Preview 14B

Released

2025-03-31

Family

Qwen 2

Type

Fine-Tuned Model

Downloads

External Download

You are about to open a link to an external source. Verify the URL before continuing.

Download

4-bit GGUF (Q4_K_M)

GGUF · deepcogito_cogito-v1-preview-qwen-14B-Q4_K_M.gguf

5-bit GGUF (Q5_K_M)

GGUF · deepcogito_cogito-v1-preview-qwen-14B-Q5_K_M.gguf

6-bit GGUF (Q6_K)

GGUF · deepcogito_cogito-v1-preview-qwen-14B-Q6_K.gguf

8-bit GGUF (Q8_0)

GGUF · deepcogito_cogito-v1-preview-qwen-14B-Q8_0.gguf

Ollama Model (q4_K_M)

Ollama

Model Report

Overview

Cogito V1 Preview 14B is an instruction-tuned, generative large language model (LLM) developed by Deep Cogito. As part of the Cogito family, this model is designed for advanced text generation tasks in a wide variety of languages and domains. The Cogito series encompasses several model sizes, including 3B, 8B, 14B, 32B, and 70B, with plans for future expansion to larger parameter counts and updated model checkpoints. These models are engineered to deliver robust performance in text-centric applications while incorporating innovative architectural and training strategies.

Benchmark comparison of Cogito 14B with other 14B models across a range of tasks

Figure 1. Benchmark table comparing Cogito 14B (Standard and Thinking modes) with Qwen2.5 14B and Deepseek R1 Distill 14B across general, math, and multi-lingual domains. Data includes both non-reasoning and reasoning results.

Model Architecture and Technical Details

Cogito V1 Preview 14B is fundamentally based on the Qwen 2.5 14B architecture, utilizing a parameter count of approximately 14.8 billion and employing BF16 tensor type for efficient computation. The model supports a substantial context window of 128,000 tokens, enabling the handling of lengthy and complex textual inputs. Its multilingual training corpus spans over 30 languages, reflecting a focus on interoperability and global application.

One distinguishing aspect of the Cogito models is their "hybrid reasoning" capability. Unlike standard LLMs, Cogito can directly answer user prompts or, alternatively, engage in a reflective reasoning process before responding. This specialized mode—referred to as the "deep thinking subroutine"—can be activated via a specific system prompt or through configuration parameters in the inference pipeline. Such design choices facilitate the model's flexibility in accommodating both rapid-response and in-depth reasoning tasks.

Training Methodology: Iterated Distillation and Amplification

The most notable innovation in the Cogito V1 Preview 14B training process is its adoption of Iterated Distillation and Amplification (IDA), a strategy developed to achieve scalable and efficient model alignment. IDA training consists of two recurring stages: "amplification," where the model generates more sophisticated solutions by invoking additional computation and auxiliary reasoning subroutines, and "distillation," in which these enhanced behaviors are internalized into the model's core parameters. This paradigm promotes continuous self-improvement and incremental capability gains over successive training cycles.

IDA distinguishes itself from traditional approaches like Reinforcement Learning from Human Feedback (RLHF) by minimizing reliance on direct human oversight and enabling self-supervised advancement. The Cogito models begin with pretrained Llama or Qwen checkpoints and progressively absorb higher-level cognitive patterns through iterative amplification/distillation processes. According to published overviews and research on IDA, this method aims to surpass the limitations inherent to human oversight in training highly capable language models.

Performance Benchmarks and Evaluation

Empirical evaluations indicate that Cogito V1 Preview 14B offers competitive performance relative to other 14B-parameter language models. Benchmarking spans domains including general knowledge, mathematics, and multilingual capability, and encompasses both standard ("non-reasoning") and advanced ("reasoning") modes. For example, when assessed on tasks such as MMLU, GSM8K, and MMMLU, Cogito 14B demonstrates notable improvements in key categories compared to base Qwen 2.5 14B and Deepseek R1 Qwen 14B models.

The provided evaluation framework ensures impartiality by excluding benchmark test sets from the model's training data, using string-matching removal strategies. Detailed performance reports and live benchmark results are disseminated in official technical documentation, revealing that related larger models (such as Cogito 70B) also show strong results in cross-model comparisons, including those involving models like Llama 4 109B MoE and Llama 3.3 70B.

Functional Capabilities and Applications

Cogito V1 Preview 14B is optimized for a range of practical applications, with particular emphasis on coding, STEM-related tasks, complex instruction following, and agentic use cases. The model supports a variety of tool-calling functionalities, allowing for direct, parallel, and multi-tool invocation within both standard and extended "thinking" modes. This flexibility is particularly relevant for tasks requiring programmatic interaction, automated reasoning, and workflow orchestration.

The model's hybrid reasoning and multilingual training allow it to provide helpful responses across a wide array of user prompts, from day-to-day query answering to more sophisticated reasoning or computational tasks. While Cogito 14B supports advanced reasoning subroutines, the design intentionally avoids optimization for very long reasoning chains, reflecting an emphasis on practical performance and minimal latency in real-world deployment contexts.

Limitations and License

Despite robust capabilities, Cogito V1 Preview 14B does present certain limitations. The model is not specifically optimized for extremely deep or extended reasoning sequences, favoring instead scenarios that require concise and efficient responses. Additionally, while benchmark scores are indicative of relative technical merit, real-world performance can diverge from standardized evaluation metrics due to the nuance of user queries and application-specific requirements.

The model and its associated weights are distributed under the Apache 2.0 License, granting broad rights for commercial use and encouraging further research and development within the community. All released models in the Cogito V1 Preview family are available under this permissive open license.

Further Reading and Resources

About Qwen 2: Qwen 2 (and 2.5) is a family of advanced AI models developed by Alibaba, designed to excel in various tasks including general language understanding, coding, and mathematics.

More in the Qwen 2 Family

Alibaba Cloud /

Qwen2.5 VL 3B

A 3B-parameter multimodal language model that processes images, videos, and text with capabilities for visual reasoning, document analysis, and computer interface interaction.
Alibaba Cloud /

Qwen2.5 VL 7B

A 7-billion parameter multimodal model capable of processing images, documents, and videos with native dynamic resolution support and structured output capabilities.
Alibaba Cloud /

Qwen2.5 VL 72B

A 72-billion parameter multimodal model that processes images, videos, and text for document analysis, object detection, and visual agent tasks.
Alibaba Cloud /

QwQ 32B Preview

Experimental research model focused on advancing AI reasoning capabilities by thinking through problems step by step, continuously questioning assumptions, and exploring different paths of thought. Impressive benchmark scores in math and coding tasks.
Alibaba Cloud /

QwQ 32B

A 32.5-billion parameter transformer model optimized for mathematical reasoning, coding, and complex problem-solving through reinforcement learning training.
Alibaba Cloud /

Qwen 2.5 Math 1.5B

A 1.5 billion parameter specialized model focused on mathematical reasoning and problem-solving capabilities in English and Chinese.
Deepseek AI /

DeepSeek R1 Distill Qwen 1.5B

A 1.5 billion parameter language model created through distillation techniques, focusing on mathematical reasoning and chain-of-thought problem-solving capabilities.
Agentica /

DeepCoder 1.5B Preview

A 1.5B parameter code generation model fine-tuned with reinforcement learning and iterative context lengthening for long-context reasoning tasks.
Alibaba Cloud /

Qwen 2.5 Math 7B

A 7.62-billion parameter mathematical reasoning model supporting English and Chinese with chain-of-thought and tool-integrated reasoning capabilities.
Alibaba Cloud /

Qwen 2.5 Math PRM 7B

A 7B parameter process reward model that evaluates mathematical reasoning steps individually rather than just final answers, achieving 67.6% accuracy on mathematical benchmarks.
Deepseek AI /

DeepSeek R1 Distill Qwen 7B

A 7.62B parameter distilled language model based on Qwen2.5-Math-7B, trained via knowledge distillation for mathematical and logical reasoning tasks.
Alibaba Cloud /

Qwen 2.5 Math 72B

A 72.7 billion parameter language model specialized for solving mathematical problems in English and Chinese using chain-of-thought and tool-integrated reasoning.
Alibaba Cloud /

Qwen 2.5 Math PRM 72B

A 72.8 billion parameter process reward model that evaluates intermediate mathematical reasoning steps using consensus filtering and achieves 78.3% F1 on ProcessBench.
Alibaba Cloud /

Qwen 2.5 Coder 7B

A 7.61 billion parameter transformer model designed for code generation and reasoning across 92 programming languages with a 128,000-token context window.
Alibaba Cloud /

Qwen 2.5 Coder 32B

Qwen2.5-Coder builds on Qwen2.5 by training on 5.5 trillion additional tokens of code data including source code, text-code grounding data, and synthetic data. This leads to significant improvements in code-related tasks.
Alibaba Cloud /

Qwen 2.5 7B

A 7.61-billion parameter transformer-based language model with 128K token context length, trained on 18 trillion tokens supporting 29+ languages.
Alibaba Cloud /

Qwen2.5 7B 1M

A 7-billion parameter transformer model with 1-million token context capacity utilizing Dual Chunk Attention for long-range text processing.
Alibaba Cloud /

Qwen 2.5 14B

A 14.7 billion parameter transformer model supporting 29 languages with 128K context window, designed as foundation for fine-tuning applications.
Alibaba Cloud /

Qwen2.5 14B 1M

A 14.7B parameter transformer model with 1-million token context capacity, utilizing dual chunk attention and progressive training for extended sequence processing.
Deepseek AI /

DeepSeek R1 Distill Qwen 14B

A 14B dense language model distilled from a mixture-of-experts architecture, optimized for mathematical reasoning and code generation tasks.
Agentica /

DeepCoder 14B Preview

A 14-billion parameter code reasoning model fine-tuned using distributed reinforcement learning with long-context capabilities up to 64,000 tokens.
Alibaba Cloud /

Qwen 2.5 32B

A 32.5 billion parameter multilingual transformer model trained on 18 trillion tokens with 128K context length supporting text generation, coding, and mathematical reasoning.
Deepseek AI /

DeepSeek R1 Distill Qwen 32B

A 32B parameter language model created through knowledge distillation, optimized for mathematical reasoning, code generation, and complex problem-solving tasks.
Deep Cogito /

Cogito V1 Preview 32B

A 32-billion parameter instruction-tuned model based on Qwen2.5 architecture featuring dual operational modes and iterated distillation alignment methodology.
Alibaba Cloud /

Qwen 2.5 72B

This update to Qwen 2, pretrained on 18 trillion tokens, shows impressive performance in benchmarks (MMLU 85+, HumanEval 85+, MATH 80+), outperforming models of similar size. Qwen 2.5 has multilingual support for over 29 languages and can understand and generate structured data.
Alibaba Cloud /

Qwen 2 7B

A 7.6 billion parameter multilingual decoder-only Transformer model designed for further post-training with 32,000-token context and strong coding capabilities.
Alibaba Cloud /

Qwen 2 72B

A 72-billion parameter transformer model supporting extended context lengths up to 128K tokens and demonstrating strong multilingual capabilities across diverse benchmarks.