datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
oasst2-33k-ja
oasst2-33k-ja
This repository provides an instruction tuning dataset developed by LLM-jp, a collaborative project launched in Japan.
The dataset comprises a Japanese translation of an English subset from oasst2, translated using DeepL.
The English subset can be found here.
For the creation of this dataset, we processed data from kunishou/oasst2-135k-ja.
Send Questions to
llm-jp(at)nii.ac.jp
Model Card Authors
The names are listed in alphabetical order.… See the full description on the dataset page: https://huggingface.co/datasets/llm-jp/oasst2-33k-ja.oasst2-135k-jaUpdate:
2023/12/25oasst2-135k-jaをチャット形式に変換したoasst2-chat-68k-jaを公開しました。
This dataset was created by automatically translating "OpenAssistant/oasst2" into Japanese by DeepL.
"OpenAssistant/oasst2" を DeepL翻訳を用いて日本語に自動翻訳したデータセットになります。
以下のコードを用いることで、 Instruction と Output (prompterの命令とassistantの回答)の形式に変換することができます。
ファインチューニングで使用する場合はこちらのコードで変換して下さい(変換には5分程度かかります)。
変換コード参考https://github.com/h2oai/h2o-llmstudio/blob/5ebfd3879e226b4e1afd0a0b45eb632e60412129/app_utils/utils.py#L1888
pip install… See the full description on the dataset page: https://huggingface.co/datasets/kunishou/oasst2-135k-ja.oasst2-chat-68k-jaoasst2-135k-jaをチャット形式に変換したデータセットになります。マルチターン会話でのファインチューニングをする際にご活用下さい(1レコードのトークン長が大きいのでそれなりの計算リソースが必要になります)。フォーマットは ShareGPT 形式になっています。ファインチューニングをする際はこちらの記事を参考にして下さい。
OpenAssistant/oasst2https://huggingface.co/datasets/OpenAssistant/oasst2
oasst2_french_dpo_pairs
Dataset Card for oasst2_french_dpo_pairs
This dataset was created from OpenAssistant/oasst2 by keeping only the french data and producing dpo pairs with their rank.
Dataset Card Contact
ntnq
oasst2_gl
OASST2 Galician Subset
Dataset description
This dataset is a Galician translation/adaptation of a subset of the OASST2 conversational dataset. It is intended for instruction tuning, dialogue modeling, and related experiments in Galician.
This release contains 1,786 instances in JSONL format. It does not include the full original OASST2 dataset. The data preserves the original conversation-oriented structure, where messages are linked through tree and parent identifiers.… See the full description on the dataset page: https://huggingface.co/datasets/proxectonos/oasst2_gl.oasst2-chat-5k-ja
説明
これはoasst2-chat-68k-jaの前半約41000件をSwallow-MX-8x7b-NVE-instruct-v2で評価した中で最高評価を獲得した物のみを抽出したデータセットです。
マルチターン会話でのファインチューニングをする際にご活用下さい。
Pret-a-porter
データセット
Variant
Link
instruction-v0.1
Kendamarron/pret-a-porter-instruction-v0.1
math-problem-v0.1
Kendamarron/pret-a-porter-math-problem-v0.1
jimba-instruction-simplify-200
Kendamarron/jimba-instruction-simplify-200
chat-with-cosmopedia
aixsatoshi/Chat-with-cosmopedia
longcontext-aozora-summary… See the full description on the dataset page: https://huggingface.co/datasets/sudy-super/oasst2-chat-5k-ja.oasst2-33k-en
oasst2-33k-en
This repository provides an instruction tuning dataset developed by LLM-jp, a collaborative project launched in Japan.
The dataset comprises an English subset from oasst2.
Send Questions to
llm-jp(at)nii.ac.jp
Model Card Authors
The names are listed in alphabetical order.
Hirokazu Kiyomaru, Takashi Kodama.
tasksource_oasst2_pairwise_rlhf_reward-PreferenceShareGPToasst2_thoasst2_ruВсе русские части из датасета oasst2
oasst2_chat_llama3.2g-ronimo_oasst2_top1_en-ShareGPT
