Skip to main content
Browse Models

TheDrummer

Rocinante 12B v1.1

Released

2024-08-15

Family

Mistral

Type

Fine-Tuned Model

Downloads

External Download

You are about to open a link to an external source. Verify the URL before continuing.

Download

4-bit GGUF (Q4_K_M)

GGUF · Rocinante-12B-v1.1-Q4_K_M.gguf

5-bit GGUF (Q5_K_M)

GGUF · Rocinante-12B-v1.1-Q5_K_M.gguf

6-bit GGUF (Q6_K)

GGUF · Rocinante-12B-v1.1-Q6_K.gguf

8-bit GGUF (Q8_0)

GGUF · Rocinante-12B-v1.1-Q8_0.gguf

16-bit GGUF (F16)

GGUF · Rocinante-12B-v1.1-f16.gguf

Model Report

Overview

Rocinante 12B v1.1 is a large-scale generative language model developed by BeaverAI, designed with an emphasis on creative storytelling and producing adventure-based interactive fiction. Building upon its predecessor, Rocinante 12B v1, this iteration refines its text generation capabilities, aiming to deliver more vivid narratives and nuanced prose. Through its architecture and parameter count, the model is positioned for diverse creative tasks, particularly in the realms of role-play (RP), story generation, and adventure-oriented instruction scenarios, as documented on its Hugging Face model page.

Stylized, luminous horse rearing up

Figure 1. The Rocinante model's emblematic stylized horse, reflecting the model’s adventurous spirit.

Technical Capabilities

Rocinante 12B v1.1 features a parameter count of approximately 12.2 billion, which supports advanced and adaptable language modeling, as detailed in its Hugging Face Model Card. The model employs the BF16 tensor format, leveraging efficient computation and supporting high-performance inference. Available in multiple quantized variants—including GGUF and EXL2 formats at 4, 6, and 8 bits per weight—it facilitates practical deployment across a range of environments and for diverse use cases.

Testers have observed that Rocinante 12B v1.1 generates text described as expressive and diverse, according to Hugging Face user feedback. Compared to its earlier version, v1.1 focuses on generating text with narrative depth and flow.

Model Architecture and Training

The architecture of Rocinante 12B v1.1 follows the contemporary trends of large language models, optimized for generative performance across creative and interactive applications. With 12.2 billion parameters, the model utilizes transformer-based neural network layers, which are fundamental in enabling context-aware and consistent long-form text generation, as described in the technical documentation. The inclusion of multiple quantized weights ensures adaptability for hardware with varying computational capabilities.

While specific details about the pre-training corpus and fine-tuning procedures have not been disclosed, the public documentation highlights the model’s alignment towards conversational and narrative-dense exchanges, particularly in role-play and adventure-centric applications.

Model performance leaderboard

Figure 2. Leaderboard chart positioning Rocinante 12B v1 among peer models on a suite of natural language generation benchmarks.

Applications and Use Cases

Rocinante 12B v1.1 is primarily oriented towards applications in creative writing, narrative simulation, and adventure-driven instruction. Its generative outputs are tailored to support immersive role-playing (RP) experiences and story creation, making it suitable for interactive fiction platforms as well as custom chat-based environments, as demonstrated in application scenarios. The model accommodates a range of prompting styles and chat templates, including ChatML for RP, Alpaca for instructive adventures, and Mistral for NeMo-driven dialogue. The adaptability in template compatibility encourages experimentation and customization to suit varied creative narratives and user preferences.

Recommended inference settings—such as the DRY sampler with temperature parameters typically ranging between 0.7 and 1.2—allow fine-tuning of output randomness, enabling either steady or highly creative text generation. Such settings help users to balance controlled storytelling with spontaneous narrative expansion, as further outlined in Hugging Face usage recommendations.

Model Family and Comparative Context

Rocinante 12B v1.1 is a direct successor of Rocinante 12B v1. The original model was characterized by an exploratory text generation style, providing what the documentation describes as a “pure off-the-rails experience.” In contrast, v1.1 incorporates modifications intended to enhance coherence, distinctiveness, and the breadth of vocabulary while maintaining characteristics consistent with the model's theme.

The iterative development reflects a growing demand for models that can balance creative liberty with narrative reliability, offering both improvisational storytelling and structured adventure dialogue. Benchmarking data and leaderboard comparisons are available on its Hugging Face model page.

Limitations and Deployment

At the time of reporting, Rocinante 12B v1.1 is not listed as being deployed on public inference endpoints, as indicated in its Hugging Face documentation. Detailed licensing information has not been specified in publicly available sources. While support for various quantization schemes facilitates broader compatibility and experimentation, deployment remains at the discretion of users and researchers integrating the model into custom solutions.

Chalk-style 'REMEMBER THE CANT' graphic

Figure 3. Acknowledgement graphic from the model's credits section, representing contributions to the Rocinante project.

The development of Rocinante 12B v1.1 involved multiple contributors, with acknowledgments given to “Garg” for compute resources, “MarinaraSpaghetti” for model merging, and “Statuo” for quantization formats, as documented in the model credits.

External Resources

About Mistral: The Mistral family of AI models, developed by Paris-based Mistral AI, includes the original 2023 Mistral 7B release, as well as the more recent Mistral Small, Nemo, and Large weights.

More in the Mistral Family

Mistral AI /

Mistral Large 2

123 billion parameter model from Paris-based Mistral AI, significantly more capable than its predecessor in code generation, mathematics, reasoning, multilingual support, and function calling.
TheDrummer /

Behemoth 123B v1.2

A 123-billion parameter language model optimized for conversational AI, creative prose generation, and role-playing applications with enhanced narrative consistency.
Mistral AI /

Mistral Small (2409)

A 22B parameter enterprise-grade small model, a convenient mid-point between Mistral NeMo 12B and Mistral Large 2. This version delivers significant improvements in human alignment, reasoning capabilities, and code over the previous version.
Mistral AI /

Mistral Small 3.2 (2506)

A 24-billion parameter multimodal model featuring improved instruction following, function calling, and reduced repetition over its predecessor.
Mistral AI /

Mistral Small 3.1 (2503)

A 24-billion parameter multimodal transformer supporting text and vision tasks with 128K token context length under Apache 2.0 license.
LatitudeGames /

Harbinger 24B

A 24-billion parameter language model fine-tuned on Mistral Small 3.1 Instruct, specialized for interactive storytelling and text-based adventures.
Mistral AI /

Devstral Small 1.0

A 23.6B parameter coding assistant finetuned for agentic software engineering tasks with 128K context window and 46.8% SWE-Bench performance.
Mistral AI /

Mistral Small 3 (2501)

A 24-billion parameter instruction-tuned language model with multilingual capabilities, 32K context window, and optimized low-latency inference performance.
TheDrummer /

Cydonia 24B v2

A fine-tuned 23.6 billion parameter Mistral-based model designed for long-context conversations and maintaining narrative coherence across extended dialogues.
Cognitive Computations /

Dolphin 3.0 Mistral 24B

A 24-billion parameter instruction-tuned model built on Mistral architecture with deliberately removed content filters to maximize user control over outputs.
Mistral AI /

Mistral NeMo 12B

A 12B parameter multi-lingual model that supports function calling built in collaboration with NVIDIA and trained using the new Tekken tokenizer. By some metrics, it is state-of-the-art in its size category. NeMo was trained with quantisation awareness, enabling FP8 inference without any performance loss.