datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
jimba-instuction-1k-betacyberagent/calm2-7b-chatの出力を人手でチェック・修正することで作成した日本語Instructionデータセットです。
詳しくはこちらの記事を御覧ください。
https://zenn.dev/kendama/articles/dc727218a2eae6
Nous-Instuct-PT
Nous-Instuct-PT
Synthetic instruction and pre-training style dataset prepared for Hugging Face Hub. The repository contains three train configs with intentionally different supervision styles and non-repeating prompt text across shards.
This dataset has three configs:
Config
Focus
Samples
train-001
Instruction-following and task completion
960
train-002
Transformation, labeling, repair, and ranking
576
nous
Larger mixed supervision corpus
4,200
Loading… See the full description on the dataset page: https://huggingface.co/datasets/Surpem/Nous-Instuct-PT.
