datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
oasst1-89k-ja
This dataset was created by automatically translating "OpenAssistant/oasst1" into Japanese.
The "ng_translation" flag indicates that the translation was not successful, and "1" means that the translation failed.Therefore, for data with "1", "text" and "text_en" contain the same text.
Update:
2023/11/12oasst1-89k-jaをチャット形式に変換したoasst1-chat-44k-jaを公開しました。
2023/10/21自動翻訳によるコード関連データの翻訳誤り2000箇所程度を手動で修正しました。
修正イメージを表示
修正前
もちろん!これは、Flask… See the full description on the dataset page: https://huggingface.co/datasets/kunishou/oasst1-89k-ja.oasst1-21k-ja
oasst1-21k-ja
This repository provides an instruction tuning dataset developed by LLM-jp, a collaborative project launched in Japan.
This dataset is a Japanese translation of an English subset of oasst1 using DeepL.
English subset is here.
Send Questions to
llm-jp(at)nii.ac.jp
Model Card Authors
The names are listed in alphabetical order.
Hirokazu Kiyomaru, Hiroshi Matsuda, Jun Suzuki, Namgi Han, Saku Sugawara, Shota Sasaki, Shuhei Kurita, Taishi… See the full description on the dataset page: https://huggingface.co/datasets/llm-jp/oasst1-21k-ja.oasst1-chat-44k-jaoasst1-89k-jaをチャット形式に変換したデータセットになります。マルチターン会話でのファインチューニングをする際にご活用下さい(1レコードのトークン長が大きいのでそれなりの計算リソースが必要になります)。フォーマットは ShareGPT 形式になっています。ファインチューニングをする際はこちらの記事を参考にして下さい。
oasst1-ja-89k Repositoryhttps://github.com/kunishou/oasst1-89k-ja
OpenAssistant/oasst1https://huggingface.co/datasets/OpenAssistant/oasst1
openassistant_oasst1_h2ogpt
h2oGPT Data Card
Summary
H2O.ai's openassistant_oasst1_h2ogpt is an open-source instruct-type dataset for fine-tuning of large language models, licensed for commercial use.
Number of rows: 48307
Number of columns: 3
Column names: ['input', 'prompt_type', 'source']
Source
Original Open Assistant data in tree structure
This flattened dataset created by script in h2oGPT repository
openassistant_oasst1_h2ogpt_llama2_chat
h2oGPT Data Card
Summary
H2O.ai's openassistant_oasst1_h2ogpt_llama2_chat is an open-source instruct-type dataset for fine-tuning of large language models, licensed for commercial use.
Number of rows: 44219
Number of columns: 5
Column names: ['id', 'prompt_type', 'input', 'output', 'source']
Source
Original Open Assistant data in tree structure
This flattened dataset created by script in h2oGPT repository
openassistant_oasst1_h2ogpt_graded
h2oGPT Data Card
Summary
H2O.ai's openassistant_oasst1_h2ogpt_graded is an open-source instruct-type dataset for fine-tuning of large language models, licensed for commercial use.
Number of rows: 30368
Number of columns: 5
Column names: ['input', 'source', 'prompt_type', 'grade_deberta', 'id']
Source
Original Open Assistant data in tree structure
This flattened dataset created by script in h2oGPT repository
recall-rewrite-oasst1
Recall Rewrite OASST1: knowledge-aligned SFT data
Data release for the paper "Stick to What You Know: A Study of Knowledge-Aligned Supervised Fine-Tuning"
(Becker, Kemmler, Thulke, Schäfer, Dugast, Ney; accepted at EMNLP 2026, Main Conference).
Knowledge-aligned SFT constrains supervised fine-tuning targets to what the base model already knows.
Recall Rewrite implements this without external evidence: every gold response of the SFT set is
decomposed into atomic claims, each… See the full description on the dataset page: https://huggingface.co/datasets/apptek-com/recall-rewrite-oasst1.h2ogpt-oig-oasst1-instruct-cleaned-v1
h2oGPT Data Card
Summary
H2O.ai's h2ogpt-oig-oasst1-instruct-cleaned-v1 is an open-source instruct-type dataset for fine-tuning of large language models, licensed for commercial use.
Number of rows: 349837
Number of columns: 3
Column names: ['input', 'source', 'prompt_type']
Source
Original LAION OIG Dataset
LAION OIG data detoxed and filtered down by scripts in h2oGPT repository
Original Open Assistant data in tree structure
This flattened dataset… See the full description on the dataset page: https://huggingface.co/datasets/h2oai/h2ogpt-oig-oasst1-instruct-cleaned-v1.oasst1-21k-en
oasst1-21k-en
This repository provides an instruction tuning dataset developed by LLM-jp, a collaborative project launched in Japan.
This dataset is an English subset of oasst1.
Send Questions to
llm-jp(at)nii.ac.jp
Model Card Authors
The names are listed in alphabetical order.
Hirokazu Kiyomaru, Hiroshi Matsuda, Jun Suzuki, Namgi Han, Saku Sugawara, Shota Sasaki, Shuhei Kurita, Taishi Nakamura, Takashi Kodama, Takumi Okamoto.
h2ogpt-oig-oasst1-instruct-cleaned-v2
h2oGPT Data Card
Summary
H2O.ai's h2ogpt-oig-oasst1-instruct-cleaned-v2 is an open-source instruct-type dataset for fine-tuning of large language models, licensed for commercial use.
Number of rows: 350581
Number of columns: 3
Column names: ['input', 'source', 'prompt_type']
Source
Original LAION OIG Dataset
LAION OIG data detoxed and filtered down by scripts in h2oGPT repository
Original Open Assistant data in tree structure
This flattened dataset… See the full description on the dataset page: https://huggingface.co/datasets/h2oai/h2ogpt-oig-oasst1-instruct-cleaned-v2.openassistant_oasst1
h2oGPT Data Card
Summary
H2O.ai's openassistant_oasst1 is an open-source instruct-type dataset for fine-tuning of large language models, licensed for commercial use.
Number of rows: 46283
Number of columns: 3
Column names: ['input', 'prompt_type', 'source']
Source
Original Open Assistant data in tree structure
This flattened dataset created by script in h2oGPT repository
oasst1-guanaco-extended-sharegptoasst1-guanaco-extendedThis is the Guanaco Extended dataset derived from OpenAssistant/oasst1.
Guanaco only uses the first (highest rank; rank 0) response from the assistant at each reply level as their dataset.
h2ogpt-oig-oasst1-instruct-cleaned-v3
h2oGPT Data Card
Summary
H2O.ai's h2ogpt-oig-oasst1-instruct-cleaned-v3 is an open-source instruct-type dataset for fine-tuning of large language models, licensed for commercial use.
Number of rows: 269406
Number of columns: 4
Column names: ['input', 'source', 'prompt_type', 'id']
Source
Original LAION OIG Dataset
LAION OIG data detoxed and filtered down by scripts in h2oGPT repository
Original Open Assistant data in tree structure
This flattened dataset… See the full description on the dataset page: https://huggingface.co/datasets/h2oai/h2ogpt-oig-oasst1-instruct-cleaned-v3.oasst1-89k-ja
学習データセット
OpenAssistant/oasst1を日本語化したデータセットであるkunishou/oasst1-89k-jaをすべて利用した。
このデータセットは、LINE社のInstruction Tuningに利用されている。
oasst1_ruВсе русские части из датасета oasst1
anthropic_hh_oasst1_splitoasst1-zh-pilot
oasst1 Chinese Translation Pilot (10 samples)
This is a pilot release of 10 parallel English→Chinese samples translated from
OpenAssistant/oasst1. It is
intended as a methodology demonstration and quality evaluation artifact, not as
a training-ready dataset.
Why this exists
We are evaluating whether LLM-assisted translation of open instruction-tuning datasets
into low-resource languages can be done at a quality bar that the ML community will
accept. Chinese is our first… See the full description on the dataset page: https://huggingface.co/datasets/AgenticCommons/oasst1-zh-pilot.oasst1_enA copy of the HuggingFaceH4/oasst1_en dataset, an English version of oasst1 filtered and processed by HuggingFace H4.
If you used this dataset, please cite
@misc{köpf2023openassistantconversationsdemocratizing,
title={OpenAssistant Conversations -- Democratizing Large Language Model Alignment},
author={Andreas Köpf and Yannic Kilcher and Dimitri von Rütte and Sotiris Anagnostidis and Zhi-Rui Tam and Keith Stevens and Abdullah Barhoum and Nguyen Minh Duc and Oliver Stanley and… See the full description on the dataset page: https://huggingface.co/datasets/XSpace2000/oasst1_en.oasst1-alpaca-jsontulu_oasst1
