CoolFace
Modelpublic

hatsukaze86/llm-jp-3-13b-it

sourceHugging Faceapache-2.0updated 2y agoView on Hugging Face
0likes
Model Card

Uploaded model

  • —Developed by: hatsukaze86
  • —License: apache-2.0
  • —Finetuned from model : llm-jp/llm-jp-3-13b

This llama model was trained 2x faster with Unsloth and Huggingface's TRL library.

<img src="https://raw.githubusercontent.com/unslothai/unsloth/main/images/unsloth%20made%20with%20love.png" width="200"/>

最終課題コンペ用 Fine-tuning テンプレート(unsloth)

最終課題コンペにて Fine-tuning を行ないたい方に向けの Fine-tuning コードです。 Unsloth を使うことで Google Colab の無料で利用可能な T4 でも動作可能になっています。 環境設定の難易度が高いので、慎重に取り組んでいただければと思います。

terminalでのconda環境構築

事前にterminalで環境構築の必要があります。Google Colabでは不要です。

# conda環境の構築
wget "https://github.com/conda-forge/miniforge/releases/latest/download/Miniforge3-$(uname)-$(uname -m).sh"  

# このコマンドではいくつか質問があるので答えて下さい。おそらくインストール先のデフォルトは/root/miniforge3かと思います  
bash Miniforge3-$(uname)-$(uname -m).sh  

# 以下、インストール先が/root/miniforge3であることを前提とします  
export PATH=/root/miniforge3/bin:$PATH  
conda init  

# ここで一度、terminalを立ち上げ直す必要があります。  
# 以下のリンク先に従い環境を作ります。  
# https://docs.unsloth.ai/get-started/installation/conda-install  
conda create --name unsloth_env python=3.10 pytorch-cuda=12.1 pytorch cudatoolkit xformers -c pytorch -c nvidia -c xformers -y  
conda activate unsloth_env  
pip install "unsloth[colab-new] @ git+https://github.com/unslothai/unsloth.git"  
pip install --no-deps "trl<0.9.0" peft accelerate bitsandbytes  

# jupyter notebook用のセットアップ。  
conda install -c conda-forge ipykernel  
python -m ipykernel install --user --name=unsloth_env --display-name "Python (unsloth_env)"  

Google Colab の場合は上記の環境構築手順を行なわず、単にこのセルから実行していってください。

!pip uninstall unsloth -y !pip install --upgrade --no-cache-dir "unsloth[colab-new] @ git+https://github.com/unslothai/unsloth.git"

Google Colab のデフォルトで入っているパッケージをアップグレード(Moriyasu さんありがとうございます)

!pip install --upgrade torch !pip install --upgrade xformers

notebookでインタラクティブな表示を可能とする(ただし、うまく動かない場合あり)

!pip install ipywidgets --upgrade

Install Flash Attention 2 for softcapping support

import torch if torch.cuda.getdevicecapability()[0] >= 8: !pip install --no-deps packaging ninja einops "flash-attn>=2.6.3"

llm-jp/llm-jp-3-13bを4bit量子化のqLoRA設定でロード。

from transformers import AutoModelForCausalLM, AutoTokenizer, BitsAndBytesConfig from unsloth import FastLanguageModel import torch maxseqlength = 512 # unslothではRoPEをサポートしているのでコンテキスト長は自由に設定可能 dtype = None # Noneにしておけば自動で設定 loadin4bit = True # 今回は8Bクラスのモデルを扱うためTrue

modelid = "llm-jp/llm-jp-3-13b" newmodel_id = "llm-jp-3-13b-it" #Fine-Tuningしたモデルにつけたい名前、it: Instruction Tuning

FastLanguageModel インスタンスを作成

model, tokenizer = FastLanguageModel.frompretrained( modelname=modelid, dtype=dtype, loadin4bit=loadin4bit, trustremote_code=True, )

SFT用のモデルを用意

model = FastLanguageModel.getpeftmodel( model, r = 32, targetmodules = ["qproj", "kproj", "vproj", "oproj", "gateproj", "upproj", "downproj",], loraalpha = 32, loradropout = 0.05, bias = "none", usegradientcheckpointing = "unsloth", randomstate = 3407, userslora = False, loftqconfig = None, maxseqlength = maxseq_length, )

Hugging Face Token を指定

下記の URL から Hugging Face Token を取得できますので下記の HF_TOKEN に入れてください。

https://huggingface.co/settings/tokens

HF_TOKEN = "your-token" #@param {type:"string"}

あるいは Google Colab シークレットを使う場合、左のサイドバーより🔑マークをクリック

HF_TOKEN という名前で Value に Hugging Face Token を入れてください。

ノートブックからのアクセスのトグルをオンにし、下記の二行のコードのコメントアウトを外してください。

from google.colab import userdata

HFTOKEN=userdata.get('HFTOKEN')

学習に用いるデータセットの指定

今回はLLM-jp の公開している Ichikara Instruction を使います。データにアクセスするためには申請が必要ですので、使いたい方のみ申請をしてください。

Ichikara Instruciton を Hugging Face Hub にて公開することはお控えください。

下記のリンクから申請を終えた先に Google Drive があり、Distribution20241221_all というフォルダごとダウンロードしてください。

今回は「ichikara-instruction-003-001-1.json」を使います。必要であれば展開(!unzip など)し、データセットのパスを適切に指定してください。

omnicampusの開発環境では取得したデータを左側にドラッグアンドドロップしてお使いください。

Google Colab の場合も左のサイドバーよりドラッグ&ドロップでアップデートしてください。

https://liat-aip.sakura.ne.jp/wp/llmのための日本語インストラクションデータ作成/llmのための日本語インストラクションデータ-公開/

関根聡, 安藤まや, 後藤美知子, 鈴木久美, 河原大輔, 井之上直也, 乾健太郎. ichikara-instruction: LLMのための日本語インストラクションデータの構築. 言語処理学会第30回年次大会(2024)

from datasets import load_dataset

dataset = loaddataset("json", datafiles="/content/ichikara-instruction-003-001-1.json")

パスの指定にご注意ください。アップロードしたファイルを右クリックし、「パスをコピー」をクリック、上記の data_files と合致していることをご確認ください。Omnicampus のディレクトリ構造とは異なるかもしれません。

学習時のプロンプトフォーマットの定義

prompt = """### 指示 {}

回答

{}"""

""" formattingpromptsfunc: 各データをプロンプトに合わせた形式に合わせる """ EOSTOKEN = tokenizer.eostoken # トークナイザーのEOSトークン(文末トークン) def formattingpromptsfunc(examples): input = examples["text"] # 入力データ output = examples["output"] # 出力データ text = prompt.format(input, output) + EOSTOKEN # プロンプトの作成 return { "formattedtext" : text, } # 新しいフィールド "formatted_text" を返す pass

# 各データにフォーマットを適用

dataset = dataset.map( formattingpromptsfunc, num_proc= 4, # 並列処理数を指定 )

dataset

データを確認

print(dataset["train"]["formattedtext"][3]) """ trainingarguments: 学習の設定

  • —output_dir: -トレーニング後のモデルを保存するディレクトリ
  • —perdevicetrainbatchsize:
  • —デバイスごとのトレーニングバッチサイズ
  • —perdeviceevalbatchsize:
  • —デバイスごとの評価バッチサイズ
  • —gradientaccumulationsteps:
  • —勾配を更新する前にステップを積み重ねる回数
  • —optim:
  • —オプティマイザの設定
  • —numtrainepochs:
  • —エポック数
  • —eval_strategy:
  • —評価の戦略 ("no"/"steps"/"epoch")
  • —eval_steps:
  • —eval_strategyが"steps"のとき、評価を行うstep間隔
  • —logging_strategy:
  • —ログ記録の戦略
  • —logging_steps:
  • —ログを出力するステップ間隔
  • —warmup_steps:
  • —学習率のウォームアップステップ数
  • —save_steps:
  • —モデルを保存するステップ間隔
  • —savetotallimit:
  • —保存しておくcheckpointの数
  • —max_steps:
  • —トレーニングの最大ステップ数
  • —learning_rate:
  • —学習率
  • —fp16:
  • —16bit浮動小数点の使用設定(第8回演習を参考にすると良いです)
  • —bf16:
  • —BFloat16の使用設定
  • —groupbylength:
  • —入力シーケンスの長さによりバッチをグループ化 (トレーニングの効率化)
  • —report_to:
  • —ログの送信先 ("wandb"/"tensorboard"など) """ from trl import SFTTrainer from transformers import TrainingArguments from unsloth import isbfloat16supported

trainer = SFTTrainer( model = model, tokenizer = tokenizer, traindataset=dataset["train"], maxseqlength = maxseqlength, datasettextfield="formattedtext", packing = False, args = TrainingArguments( perdevicetrainbatchsize = 2, gradientaccumulationsteps = 4, numtrainepochs = 1, loggingsteps = 10, warmupsteps = 10, savesteps=100, savetotallimit=2, maxsteps=-1, learningrate = 2e-4, fp16 = not isbfloat16supported(), bf16 = isbfloat16supported(), groupbylength=True, seed = 3407, outputdir = "outputs", reportto = "none", ), ) #@title 現在のメモリ使用量を表示 gpustats = torch.cuda.getdeviceproperties(0) startgpumemory = round(torch.cuda.maxmemoryreserved() / 1024 / 1024 / 1024, 3) maxmemory = round(gpustats.totalmemory / 1024 / 1024 / 1024, 3) print(f"GPU = {gpustats.name}. Max memory = {maxmemory} GB.") print(f"{startgpumemory} GB of memory reserved.") #@title 学習実行 trainerstats = trainer.train()

ELYZA-tasks-100-TVの読み込み。事前にファイルをアップロードしてください

データセットの読み込み。

omnicampusの開発環境では、左にタスクのjsonlをドラッグアンドドロップしてから実行。

import json datasets = [] with open("./elyza-tasks-100-TV_0.jsonl", "r") as f: item = "" for line in f: line = line.strip() item += line if item.endswith("}"): datasets.append(json.loads(item)) item = ""

学習したモデルを用いてタスクを実行

from tqdm import tqdm

推論するためにモデルのモードを変更

FastLanguageModel.for_inference(model)

results = [] for dt in tqdm(datasets): input = dt["input"]

prompt = f"""### 指示\n{input}\n### 回答\n"""

inputs = tokenizer([prompt], return_tensors = "pt").to(model.device)

outputs = model.generate(**inputs, maxnewtokens = 512, usecache = True, dosample=False, repetitionpenalty=1.2) prediction = tokenizer.decode(outputs[0], skipspecial_tokens=True).split('\n### 回答')[-1]

results.append({"taskid": dt["taskid"], "input": input, "output": prediction})

jsonlで保存

with open(f"{newmodelid}output.jsonl", 'w', encoding='utf-8') as f: for result in results: json.dump(result, f, ensureascii=False) f.write('\n')

モデルとトークナイザーをHugging Faceにアップロード。

一旦privateでアップロードしてください。

最終成果物が決まったらpublicにするようお願いします。

現在公開しているModelInferenceTemplate.ipynbはunslothを想定していないためそのままでは動かない可能性があります。

model.pushtohubmerged( newmodelid, tokenizer=tokenizer, savemethod="lora", token=HF_TOKEN, private=True )

model.pushtohub(newmodelid, token=HF_TOKEN, private=True) # Online saving

tokenizer.pushtohub(newmodelid, token=HF_TOKEN) # Online saving