Skip to main content
Browse Models

Sao10K

70B L3.3 Cirrus x1

Released

2025-01-06

Family

Llama 3

Type

Fine-Tuned Model

Downloads

External Download

You are about to open a link to an external source. Verify the URL before continuing.

Download

4-bit GGUF (Q4_K_M)

GGUF · 70B-L3.3-Cirrus-x1-Q4_K_M.gguf

Model Report

Overview

The 70B L3.3 Cirrus x1 is a large-scale generative language model developed by Sao10K, notable for its distinctive data composition and extended training methodology. Building on the Llama model family, Cirrus x1 integrates advanced training strategies aimed at improving output stability and performance in text generation tasks. Its development reflects ongoing experimentation within the open-source language modeling community and emphasizes both technical rigor and iterative model refinement.

Introductory illustration associated with the 70B L3.3 Cirrus x1 model.

Figure 1. Introductory illustration associated with the public release of 70B L3.3 Cirrus x1.

Model Architecture and Lineage

At its core, 70B L3.3 Cirrus x1 is a finetuned variant within the Llama model ecosystem, specifically derived from Llama 3 70B and further refined from Llama 3.3 70B. This developmental lineage ensures compatibility with the Llama-3-Instruct prompt formatting and leverages advancements in underlying transformer architectures to manage large parameter counts efficiently. The model weights are stored using the BF16 tensor format, which supports efficient computation while maintaining numerical stability during inference.

With a parameter count of 70.6 billion, Cirrus x1 belongs to the class of high-capacity transformer-based language models designed for nuanced text understanding and generation. The model's development included merging with checkpoints using tailored techniques to enhance output consistency and minimize performance regressions observed in earlier iterations.

Training Methodology and Data Composition

A distinguishing characteristic of 70B L3.3 Cirrus x1 is its training regimen, which closely echoes the approach used for the "Freya" model while introducing unique adjustments in dataset curation and extended training duration. The implementation involved iterative checkpoint merging across multiple epochs—utilizing experimental tools such as dare_ties—to reinforce model stability and counteract overfitting tendencies observed in long-duration training scenarios.

The primary training phase was conducted over approximately 22 hours utilizing an 8xH100 node configuration, followed by an additional 3 hours focused on cross-epoch checkpoint merging. This latter phase employed a 2xH200 node setup, which allowed for experimental refinement of the trained network and integration of multiple learned representations to further bolster output diversity and reliability.

Evaluation and Performance

Formal benchmark scores and comprehensive leaderboard evaluations for 70B L3.3 Cirrus x1 have not been explicitly published. However, the model's stability and stylistic characteristics have been documented and are noted as key improvements compared to previous versions. Contributions to various leaderboard frameworks are possible if a suitable dataset is specified for evaluation within model card metadata.

Empirical observations by the developer suggest that the model produces outputs with a "nice style," addressing known generation issues present in earlier releases. Occasional artifacts or generation anomalies may still occur, but these are reported as being easily correctable through minor fine-tuning or prompt engineering adjustments.

Typical Use Cases and Configuration

70B L3.3 Cirrus x1 is optimized for text generation scenarios that demand high coherence and context retention over extended passages. It employs the Llama-3-Instruct prompt format, making it suitable for both general instruction-following tasks and creative composition. The recommended inference configuration includes a temperature setting of 1.1 to enable creative outputs and a minimum probability (min_p) value of 0.05 to ensure response diversity. While the model is compatible with standard sampling algorithms, alternative samplers such as DRY or XTC may be used at the operator's discretion, though detailed support for these is not provided by the developer.

Screenshot: conversation output from 70B L3.3 Cirrus x1 model.

Figure 2. Example of model-generated dialogue discussing various cloud types. The conversation demonstrates contextual understanding and creative naming.

Limitations and Model Family

Despite its robust training and checkpoint merging strategy, 70B L3.3 Cirrus x1 may occasionally exhibit output inconsistencies that require post-processing or prompt refinement. Formal licensing details are not specified in the publicly available technical documentation. The model serves as both a foundation for derivative models—having been used in further finetunes, merges, and quantizations—and an experimental platform for continued research into large language model synthesis and stability.

For researchers and practitioners seeking comparative context, the model directly draws upon the data and approach utilized for the Freya model and represents a downstream extension of major releases within the Llama series. Metrics for derivative models and their relationship to Cirrus x1 can be explored through associated model card documentation.

External Resources

For further technical details and ongoing developments, the following resources may be helpful:

About Llama 3: The Llama 3 family of AI models, developed by Meta, represents a significant advancement in open-source large language models, offering parameter sizes up to 405 billion and supporting context windows of up to 128k tokens. Llama 3.1, 3.2, and 3.3 optimize this performance through distillation learning and improved multimodal capabilities.

More in the Llama 3 Family

Meta /

Llama 3 8B

Large language model with 8 billion parameters featuring transformer architecture, trained on 15 trillion tokens for text generation and coding tasks.
Meta /

Llama 3 70B

State-of-the-art 70B foundation model from Meta, trained on over 15 trillion tokens.
Meta /

Llama 3.1 8B

Llama 3.1 is a new state-of-the-art large language model from Meta.
Sao10K /

Llama 3.1 8B Stheno v3.4

An 8-billion parameter language model fine-tuned for multi-turn dialogue, creative writing, and roleplaying using curated conversational datasets and synthetic data.
Deepseek AI /

DeepSeek R1 Distill Llama 8B

Distilled 8B-parameter model optimized for mathematical reasoning and code generation through knowledge transfer from larger reinforcement learning-trained teacher models.
Deep Cogito /

Cogito V1 Preview 8B

A Llama 3.1-based model trained with Iterated Distillation and Amplification, featuring dual reasoning modes and tool calling capabilities.
Meta /

Llama 3.1 70B

The Llama 3.1 series of open models rivals top closed models in performance. It was trained on over 15 trillion tokens using over 16K H100 GPUs. These models display state-of-the-art capabilities in general knowledge, steerability, math, tool use, and translation.
Deep Cogito /

Cogito V1 Preview 70B

A 70B parameter instruction-tuned model based on Llama 3.1 architecture featuring dual reasoning modes and multilingual tool-calling capabilities.
Meta /

Llama 3.2 3B

The next iteration in the Llama series of open models. This lightweight model was designed to run on edge devices, even mobile.
Cognitive Computations /

Dolphin 3.0 Llama3.2 3B

An uncensored instruct-tuned 3.2B parameter language model that grants users full control over system prompts and behavioral alignment.
Deep Cogito /

Cogito V1 Preview 3B

A 3B-parameter multilingual instruction-tuned model based on Llama 3.2 that supports tool-calling and features dual operational modes for standard and extended reasoning.
Meta /

Llama 3.3 70B

Llama 3.3 is a text-only 70B instruction-tuned model that provides enhanced performance relative to Llama 3.1 70B and to Llama 3.2 90B when used for text-only applications. For some applications, Llama 3.3 70B approaches the performance of Llama 3.1 405B.
Sao10K /

L3.3 70B Euryale v2.3

A 70-billion parameter language model fine-tuned from Llama 3.3 for creative writing and role-playing applications using custom datasets.
TheDrummer /

Anubis 70B v1

A 70.6-billion parameter text generation model fine-tuned from Llama 3.3, designed for creative writing and role-playing applications.
TheDrummer /

Anubis 70B v1.1

A 70.6 billion parameter Llama 3.3-based model fine-tuned for character consistency and dynamic dialogue in creative text generation applications.
LatitudeGames /

Wayfarer Large 70B Llama 3.3

A 70.6-billion parameter language model fine-tuned for adventure role-play scenarios, emphasizing conflict, tension, and narrative stakes in second-person storytelling.
Deepseek AI /

DeepSeek R1 Distill Llama 70B

A 70B parameter dense language model distilled from DeepSeek-R1 using Llama 3.3 architecture, optimized for mathematical and coding reasoning tasks.