TheDrummer
Rocinante 12B v1.1
Downloads
Model Report
Overview
Rocinante 12B v1.1 is a large-scale generative language model developed by BeaverAI, designed with an emphasis on creative storytelling and producing adventure-based interactive fiction. Building upon its predecessor, Rocinante 12B v1, this iteration refines its text generation capabilities, aiming to deliver more vivid narratives and nuanced prose. Through its architecture and parameter count, the model is positioned for diverse creative tasks, particularly in the realms of role-play (RP), story generation, and adventure-oriented instruction scenarios, as documented on its Hugging Face model page.

Figure 1. The Rocinante model's emblematic stylized horse, reflecting the model’s adventurous spirit.
Technical Capabilities
Rocinante 12B v1.1 features a parameter count of approximately 12.2 billion, which supports advanced and adaptable language modeling, as detailed in its Hugging Face Model Card. The model employs the BF16 tensor format, leveraging efficient computation and supporting high-performance inference. Available in multiple quantized variants—including GGUF and EXL2 formats at 4, 6, and 8 bits per weight—it facilitates practical deployment across a range of environments and for diverse use cases.
Testers have observed that Rocinante 12B v1.1 generates text described as expressive and diverse, according to Hugging Face user feedback. Compared to its earlier version, v1.1 focuses on generating text with narrative depth and flow.
Model Architecture and Training
The architecture of Rocinante 12B v1.1 follows the contemporary trends of large language models, optimized for generative performance across creative and interactive applications. With 12.2 billion parameters, the model utilizes transformer-based neural network layers, which are fundamental in enabling context-aware and consistent long-form text generation, as described in the technical documentation. The inclusion of multiple quantized weights ensures adaptability for hardware with varying computational capabilities.
While specific details about the pre-training corpus and fine-tuning procedures have not been disclosed, the public documentation highlights the model’s alignment towards conversational and narrative-dense exchanges, particularly in role-play and adventure-centric applications.

Figure 2. Leaderboard chart positioning Rocinante 12B v1 among peer models on a suite of natural language generation benchmarks.
Applications and Use Cases
Rocinante 12B v1.1 is primarily oriented towards applications in creative writing, narrative simulation, and adventure-driven instruction. Its generative outputs are tailored to support immersive role-playing (RP) experiences and story creation, making it suitable for interactive fiction platforms as well as custom chat-based environments, as demonstrated in application scenarios. The model accommodates a range of prompting styles and chat templates, including ChatML for RP, Alpaca for instructive adventures, and Mistral for NeMo-driven dialogue. The adaptability in template compatibility encourages experimentation and customization to suit varied creative narratives and user preferences.
Recommended inference settings—such as the DRY sampler with temperature parameters typically ranging between 0.7 and 1.2—allow fine-tuning of output randomness, enabling either steady or highly creative text generation. Such settings help users to balance controlled storytelling with spontaneous narrative expansion, as further outlined in Hugging Face usage recommendations.
Model Family and Comparative Context
Rocinante 12B v1.1 is a direct successor of Rocinante 12B v1. The original model was characterized by an exploratory text generation style, providing what the documentation describes as a “pure off-the-rails experience.” In contrast, v1.1 incorporates modifications intended to enhance coherence, distinctiveness, and the breadth of vocabulary while maintaining characteristics consistent with the model's theme.
The iterative development reflects a growing demand for models that can balance creative liberty with narrative reliability, offering both improvisational storytelling and structured adventure dialogue. Benchmarking data and leaderboard comparisons are available on its Hugging Face model page.
Limitations and Deployment
At the time of reporting, Rocinante 12B v1.1 is not listed as being deployed on public inference endpoints, as indicated in its Hugging Face documentation. Detailed licensing information has not been specified in publicly available sources. While support for various quantization schemes facilitates broader compatibility and experimentation, deployment remains at the discretion of users and researchers integrating the model into custom solutions.

Figure 3. Acknowledgement graphic from the model's credits section, representing contributions to the Rocinante project.
The development of Rocinante 12B v1.1 involved multiple contributors, with acknowledgments given to “Garg” for compute resources, “MarinaraSpaghetti” for model merging, and “Statuo” for quantization formats, as documented in the model credits.
External Resources
More in the Mistral Family
Mistral Large 2
Behemoth 123B v1.2
Mistral Small (2409)
Mistral Small 3.2 (2506)
Mistral Small 3.1 (2503)
Harbinger 24B
Devstral Small 1.0
Mistral Small 3 (2501)
Cydonia 24B v2
Dolphin 3.0 Mistral 24B
Mistral NeMo 12B
More from TheDrummer
Anubis 70B v1
Anubis 70B v1.1
Compatible Apps

Open WebUI
A polished, self-hosted chat interface for LLMs with Ollama integration, multimodal prompts, and extensive workspace customization.
Web UI
Chat UIs · Beginner Friendly
llama.cpp
GGUF model inference with a polished web UI and OpenAI-format API. CPU-only build — high hardware compatibility, works on any machine without a GPU.
Web UI · API · CLI
LLM Inference · Chat UIs
llama.cpp (CUDA)
GGUF model inference with a polished web UI and OpenAI-format API. CUDA build — GPU-accelerated for NVIDIA GPUs.
Web UI · API · CLI
LLM Inference · Chat UIs

Text Generation Web UI
A feature-rich interface for running and experimenting with open-weight LLMs, including multiple inference backends, plugins, and tuning controls.
Web UI · API
Chat UIs · LLM Inference