CoolFace
Modelpublic

myxy/recursive-compressor-v1.2-2.6b-c4c1-instruct

sourceHugging Facellama2updated 3mo agoView on Hugging Face
0likes11downloads
Model Card

RecursiveCompressorLM (Instruct)

English | 日本語

日本語

モデル概要

RecursiveCompressor アーキテクチャによる日英バイリンガル対話モデルです。 事前学習版の重みから対話データセットでファインチューニングしました。

  • —アーキテクチャ: RecursiveCompressorLM (PreTrainedModel)
  • —パラメータ数: 2,601,891,840
  • —ベースモデル: myxy/recursive-compressor-2-7b
  • —コンテキスト長: 2048(学習時)
  • —トークナイザ: elyza/ELYZA-japanese-Llama-2-7b-fast
  • —dtype: bfloat16

対話フォーマット

本モデルは以下のテンプレートで学習されています:

<s>[INST]質問1[/INST]回答1</s><s>[INST]質問2[/INST]回答2</s>...
推論例
python
prompt = "[INST]日本の首都はどこですか?[/INST]"
# → モデルが回答を生成

訓練データ

事前学習に加え、以下の対話データセットで追加学習しました:

データセット言語ライセンス・由来
shi3z/ja_conv_wikipedia_llama2pro8b_30k日本語Llama2-Pro 8B 生成(Llama 2 ToS継承)
shi3z/ja_conv_wikipedia_orion14B_100K日本語Orion-14B 生成
HuggingFaceH4/ultrachat_200k英語MIT(GPT系で生成された合成対話)

訓練設定

  • —ベースモデルからの warm-start
  • —オプティマイザ: Muon (隠れ層2D重み) + AdamW (embedding/head/bias/norm)
  • —学習率: 5e-5
  • —並列方式: パイプライン並列 (PyTorch PipelineStage + Schedule1F1B, 6 GPUs)
  • —バッチ: マイクロバッチ6 × バッチサイズ6
  • —混合精度: fp32マスター重み + bfloat16 autocast

用途

  • —想定用途: 日本語/英語の質問応答、対話、文章生成
  • —非想定用途: 医療・法律・金融などの高リスク判断、安全性が重要な用途、事実確認を要する用途

使い方

本モデルは独自アーキテクチャ RecursiveCompressorLM を使用するため、 以下のリポジトリをクローンしてその中のクラス定義を読み込む必要があります:

リポジトリ: https://github.com/myxyy/RecursiveCompressorHF

bash
git clone https://github.com/myxyy/RecursiveCompressorHF.git -b v1.2
cd RecursiveCompressorHF
uv sync

HuggingFaceの generate() メソッドに対応しています:

python
import torch
from transformers import AutoTokenizer, TextStreamer
from recursive_compressor_lm import RecursiveCompressorLM

model = RecursiveCompressorLM.from_pretrained(
    "myxy/recursive-compressor-v1.2-2.6b-c4c1-instruct",
    torch_dtype=torch.bfloat16,
).to("cuda").eval()
tokenizer = AutoTokenizer.from_pretrained("elyza/ELYZA-japanese-Llama-2-7b-fast")

prompt = "[INST]猫の足は何本ですか?[/INST]"
input_ids = tokenizer.encode(prompt, return_tensors="pt").to("cuda")

output_ids = model.generate(
    input_ids,
    max_new_tokens=256,
    do_sample=True,
    temperature=0.8,
    top_p=0.9,
    streamer=TextStreamer(tokenizer, skip_prompt=True, skip_special_tokens=True),
)
#print(tokenizer.decode(output_ids[0], skip_special_tokens=True))

リポジトリ内の predict.py / predict_stream.py も同等の機能を提供します(インタラクティブREPL等)。

注: beam searchは未対応です。do_sample=Falseで貪欲生成、do_sample=Trueで確率サンプリングが使えます。

制限・バイアス

  • —ベースモデルのバイアスを継承
  • —対話データの大半が他のLLMによる合成データのため、それらモデルのスタイル・誤りを反映しうる
  • —事実関係の誤り(hallucination)が頻発する可能性
  • —安全性フィルタは未実装。有害な出力を生成する可能性がある
  • —マルチターン対話で長い文脈を保持する能力は限定的

ライセンス

Llama 2 Community License に従います。

  • —ライセンス全文: https://llama.meta.com/llama2/license/
  • —「Built with Meta Llama 2」と明示する必要があります
  • —MAU 7億超の事業者は別途ライセンス取得が必要

トークナイザがLlama 2派生であること、および対話学習データの一部 (shi3z) がLlama2-Proで生成されていることから、本モデルはこのライセンスに従います。

引用

TODO


English

Model Description

A bilingual (Japanese / English) conversational model built on the RecursiveCompressor architecture. Fine-tuned from the pretrain variant on dialogue datasets.

  • —Architecture: RecursiveCompressorLM (extends PreTrainedModel)
  • —Parameters: 2,601,891,840
  • —Base model: MODEL_REPO_ID/pretrain
  • —Training context length: 2048
  • —Tokenizer: elyza/ELYZA-japanese-Llama-2-7b-fast
  • —dtype: bfloat16

Conversation Format

The model is trained with the following template:

<s>[INST]question1[/INST]answer1</s><s>[INST]question2[/INST]answer2</s>...
Inference Example
python
prompt = "[INST]What is the capital of Japan?[/INST]"
# → model generates the answer

Training Data

In addition to the pretrain corpus, fine-tuned on these dialogue datasets:

DatasetLanguageLicense / Origin
shi3z/ja_conv_wikipedia_llama2pro8b_30kJapaneseGenerated by Llama2-Pro 8B (inherits Llama 2 ToS)
shi3z/ja_conv_wikipedia_orion14B_100KJapaneseGenerated by Orion-14B
HuggingFaceH4/ultrachat_200kEnglishMIT (synthetic dialogues generated by GPT models)

Training Setup

  • —Warm-started from the pretrain model
  • —Optimizers: Muon (2D hidden weights) + AdamW (embedding/head/bias/norm)
  • —Learning rate: 5e-5
  • —Parallelism: pipeline parallel (PyTorch PipelineStage + Schedule1F1B, 6 GPUs)
  • —Batch: 6 microbatches × batch size 6
  • —Mixed precision: fp32 master weights + bfloat16 autocast

Usage

This model uses the custom RecursiveCompressorLM architecture, so you need to clone the repository to import the class definitions:

Repository: https://github.com/myxyy/RecursiveCompressorHF

bash
git clone https://github.com/myxyy/RecursiveCompressorHF.git -b v1.2
cd RecursiveCompressorHF
uv sync

The model supports HuggingFace's generate() method:

python
import torch
from transformers import AutoTokenizer, TextStreamer
from recursive_compressor_lm import RecursiveCompressorLM

model = RecursiveCompressorLM.from_pretrained(
    "myxy/recursive-compressor-v1.2-2.6b-c4c1-instruct",
    torch_dtype=torch.bfloat16,
).to("cuda").eval()
tokenizer = AutoTokenizer.from_pretrained("elyza/ELYZA-japanese-Llama-2-7b-fast")

prompt = "[INST]How many legs does a cat have?[/INST]"
input_ids = tokenizer.encode(prompt, return_tensors="pt").to("cuda")

output_ids = model.generate(
    input_ids,
    max_new_tokens=256,
    do_sample=True,
    temperature=0.8,
    top_p=0.9,
    streamer=TextStreamer(tokenizer, skip_prompt=True, skip_special_tokens=True),
)
#print(tokenizer.decode(output_ids[0], skip_special_tokens=True))

See predict.py / predict_stream.py in the repository for an interactive REPL.

Note: beam search is not supported (use do_sample=False for greedy or do_sample=True for sampling).

Intended Use

  • —Intended: Japanese / English question answering, dialogue, text generation
  • —Not intended: High-stakes decisions in medical / legal / financial domains; safety-critical use; fact verification

Limitations & Bias

  • —Inherits biases from the base model
  • —Most of the dialogue training data is synthetic data generated by other LLMs, so this model may reflect the style and errors of those models
  • —May frequently produce factually incorrect statements (hallucination)
  • —No safety filtering; may produce harmful outputs
  • —Limited ability to maintain long conversational context across many turns

License

Llama 2 Community License.

  • —Full text: https://llama.meta.com/llama2/license/
  • —"Built with Meta Llama 2" attribution required
  • —Organizations with > 700M MAU must seek a separate license

The tokenizer is derived from Llama 2, and part of the instruction tuning data (shi3z) was generated by Llama2-Pro, so this model inherits the license.

Citation

TODO

Acknowledgments

  • —Meta AI for Llama 2
  • —ELYZA for the Japanese-extended tokenizer
  • —shi3z, HuggingFaceH4 for dialogue datasets