datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
gaokao-sft-chinese-strict-abcd-v3
Gaokao SFT Chinese Strict ABCD V3
This dataset is the cleaned Chinese SFT release that keeps only single-choice samples where A, B, C, and D all have explicit option-level analysis.
Composition
Total samples: 88466
Train samples: 86670
Validation samples: 1796
Subject Counts
{
"biology": 33104,
"chemistry": 35796,
"english": 174,
"general_exam": 7982,
"geography": 888,
"history": 229,
"physics": 9986,
"politics": 307
}
Fields
id… See the full description on the dataset page: https://huggingface.co/datasets/callofthenight1/gaokao-sft-chinese-strict-abcd-v3.atcoder_abc_contests
Notification
Atcoder is selling this data now. If you are interested in accessing it please contact them.
Dataset Summary
This dataset aims to facilitate the creation of sophisticated, multi-turn dialogue datasets focused on coding
that could be used for training reasoning Large Language Models (LLMs), particularly for Supervised Fine-Tuning (SFT) and Knowledge Distillation techniques.
It also serves as a robust foundation for problem-solving in Large Language… See the full description on the dataset page: https://huggingface.co/datasets/Nan-Do/atcoder_abc_contests.atcoder_abc_contests_small
Dataset Summary
This dataset aims to facilitate the creation of sophisticated, multi-turn dialogue datasets focused on coding
that could be used for training reasoning Large Language Models (LLMs), particularly for Supervised Fine-Tuning (SFT) and Knowledge Distillation techniques.
It also serves as a robust foundation for problem-solving in Large Language Models (LLMs).
The dataset includes both accepted and failed solutions from Atcoders's (ABC) contests.
In total, it features… See the full description on the dataset page: https://huggingface.co/datasets/Nan-Do/atcoder_abc_contests_small.gaokao-sft-chinese-strict-abcd
Gaokao SFT Chinese Balanced
This is the strict balanced Chinese SFT dataset version.
Only multiple-choice samples with explicit A/B/C/D option-level explanations are kept in this balanced release.
Composition
Total samples: 645
Train samples: 632
Validation samples: 13
Subject Counts
{
"biology": 199,
"chemistry": 170,
"english": 13,
"geography": 34,
"history": 118,
"physics": 77,
"politics": 34
}
Fields
id
lang
subject
source… See the full description on the dataset page: https://huggingface.co/datasets/callofthenight1/gaokao-sft-chinese-strict-abcd.ABC-Lakh-MIDI-Dataset
🎵 Mader ABC V2 - Music Dataset
This dataset contains music pieces derived from the Lakh MIDI Dataset and stylistically aligned with the Million Song Dataset (for musical genres), converted to ABC notation and tokenized for NLP-style applications on music (classification, generation, clustering, ...).
It provides two configurations, each with instrument and genre metadata:
abc_texts – text in ABC format
abc_tokens – token sequence
Each configuration contains one entry per… See the full description on the dataset page: https://huggingface.co/datasets/Gapagapi1/ABC-Lakh-MIDI-Dataset.
