datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
muri-it-language-split
MURI-IT: Multilingual Instruction Tuning Dataset for 200 Languages via Multilingual Reverse Instructions
MURI-IT is a large-scale multilingual instruction tuning dataset containing 2.2 million instruction-output pairs across 200 languages. It is designed to address the challenges of instruction tuning in low-resource languages with Multilingual Reverse Instructions (MURI), which ensures that the output is human-written, high-quality, and authentic to the cultural and linguistic… See the full description on the dataset page: https://huggingface.co/datasets/akoksal/muri-it-language-split.mala-monolingual-split
MaLA Corpus: Massive Language Adaptation Corpus
This version contains train and validation splits.
Dataset Summary
The MaLA Corpus (Massive Language Adaptation) is a comprehensive, multilingual dataset designed to support the continual pre-training of large language models. It covers 939 languages and consists of over 74 billion tokens, making it one of the largest datasets of its kind. With a focus on improving the representation of low-resource languages, the… See the full description on the dataset page: https://huggingface.co/datasets/MaLA-LM/mala-monolingual-split.toricgt-curated-splits
ToricGT Curated Graph Reasoning Splits
Curated working dataset repository for ToricGT.
The upload contains only curated split Parquet files and metadata generated locally.
Raw upstream downloads are not uploaded. Each row preserves source dataset, license, split, hashes, and graph JSON fields for audit.
Hebrew/Jewish-text records are sourced from Sefaria and UniMorph Hebrew sources.
Files
train.parquet
validation.parquet
test.parquet
all.parquet if… See the full description on the dataset page: https://huggingface.co/datasets/AmelieSchreiber/toricgt-curated-splits.sft_processed_large_split
sft_processed_large — profile-disjoint split
This is the train / val / test split of Xuhui/sft_processed_large, the
OdysSim midtraining corpus (21.4M interactions across 63 datasets).
Split structure
split
rows
how it's built
train
21.20M
what's left after val + test are carved out
val
28K
per-dataset random sample, in-distribution; for checkpoint selection
test
128K
profile-disjoint where the dataset's profile space supports it; for OOD generalization… See the full description on the dataset page: https://huggingface.co/datasets/Xuhui/sft_processed_large_split.UltraData-Math-L3-Textbook-Exercise-Synthetic-split
UltraData-Math L3 Textbook Exercise Synthetic Split
Source dataset: openbmb/UltraData-Math
Source config: UltraData-Math-L3-Textbook-Exercise-Synthetic
Each row contains:
uid
question
answer
The original content field was split using the literal markers
The exercise: and The solution:.
fictionalqa_training_splits
Training splits view of the FictionalQA dataset
The FictionalQA dataset
Repository: https://github.com/jwkirchenbauer/fictionalqa
Paper: https://arxiv.org/abs/2506.05639
Dataset Description
This dataset is a derivative of the main dataset hf.co/datasets/jwkirchenbauer/fictionalqa. Please see that dataset's README for a detailed description of the assets.
The dataset splits (configs) provided here are the exact ones materialized and used in the experiments for… See the full description on the dataset page: https://huggingface.co/datasets/jwkirchenbauer/fictionalqa_training_splits.sommelier-xlam-single-call-splits
sommelier xlam single-call splits
Deterministic, deduplicated, single-tool-call train/validation/test
splits derived from
Salesforce/xlam-function-calling-60k
(APIGen, CC-BY-4.0), produced by the
sommelier pipeline for
reproducible tool-calling fine-tuning. These are the exact splits used to
train and evaluate
abdelstark/llama-3.1-nemotron-nano-8b-xlam-tool-calling-lora.
Why single-call
The upstream dataset mixes single-call and multi-call examples (~52.6%… See the full description on the dataset page: https://huggingface.co/datasets/abdelstark/sommelier-xlam-single-call-splits.pallas_splitted_18cgutenberg_clean_en_splits
Dataset Card for Project Gutenberg Cleaned with splits (English Only) Dataset
This dataset is a cleaned English-language subset of the Project Gutenberg Dataset manu/project_gutenberg, originally containing ~70,000 digitized books.
The original dataset includes multiple languages, duplicate entries, and boilerplate content, all of which were removed for practicality and cleaner downstream use.
This dataset containg 38.026 books.
Dataset Splits
The dataset is divided… See the full description on the dataset page: https://huggingface.co/datasets/nikolina-p/gutenberg_clean_en_splits.Superior-Reasoning-SFT-gpt-oss-120b-split-en
Superior-Reasoning SFT (stage1 + stage2) with <think> split and English filtering
Summary
This dataset is a processed derivative of Alibaba-Apsara/Superior-Reasoning-SFT-gpt-oss-120b (subsets stage1 and stage2, train split). It restructures each example into three fields:
input: the original input
reasoning: the content extracted from <think> ... </think> within the original output (inner text only)
output: the remainder of the original output after removing all <think>… See the full description on the dataset page: https://huggingface.co/datasets/Hugodonotexit/Superior-Reasoning-SFT-gpt-oss-120b-split-en.atomic-metrics-rm-splits
Atomic Metrics RM Task Splits
Preference-pair benchmark splits used by
Atomic Metrics. The release
contains four open-ended task families derived from public SHP, OASST1, and
OASST2 preference data.
Dataset structure
Each configuration contains 10,000 training pairs and 2,000 test pairs. Every
row has:
{
"sample_id": "source-specific stable ID",
"source_dataset": "shp | oasst1 | oasst2",
"category": "task configuration",
"split": "train | test"… See the full description on the dataset page: https://huggingface.co/datasets/tintin1027/atomic-metrics-rm-splits.Split-IFEval
Split IFEval
This dataset modifies the Instruction-Following Eval (IFEval) benchmark to split apart the task from the syntactic instructions in addition to fixing errors in the original dataset.
It enables the use of research methods like attention steering that require access to the instruction text.
To load the dataset, run:
from datasets import load_dataset
split_ifeval = load_dataset("ibm-research/Split-IFEval")
Dataset Structure
Each entry in the dataset… See the full description on the dataset page: https://huggingface.co/datasets/ibm-research/Split-IFEval.ashaar-with-enhanced-descriptions-baseform-final-sft-lte20-min500-splits
Ashaar Enhanced Description SFT Stratified Splits
Source dataset:
Shaer-AI/ashaar-with-enhanced-descriptions-baseform-final-sft-lte20-min500
Target dataset:
Shaer-AI/ashaar-with-enhanced-descriptions-baseform-final-sft-lte20-min500-splits
This dataset publishes deterministic train / eval / test splits with a 94 / 3 / 3 policy.
Split policy
Primary stratification key:
base_meter
form
length_bucket
Length buckets:
1-3
4-6
7-10
11-20
Small groups fall back… See the full description on the dataset page: https://huggingface.co/datasets/Shaer-AI/ashaar-with-enhanced-descriptions-baseform-final-sft-lte20-min500-splits.nvidia_instruction_following_if_split_v3
Dataset Description
This is the instruction_following split only (the chat split was intentionally excluded) from
nvidia/Nemotron-SFT-Instruction-Following-Chat-v3,
re-packaged as Parquet (sharded) instead of the original single JSONL file for faster loading and native
support in the HF datasets viewer.
No content was modified — this is a straight format conversion of the instruction_following subset.
Source dataset: nvidia/Nemotron-SFT-Instruction-Following-Chat-v3
Dataset… See the full description on the dataset page: https://huggingface.co/datasets/tuandunghcmut/nvidia_instruction_following_if_split_v3.nvidia_instruction_following_if_split_v3_non_thinking
Dataset Description
Non-thinking (no chain-of-thought) variant of
tuandunghcmut/nvidia_instruction_following_if_split_v3,
which is itself the instruction_following split of
nvidia/Nemotron-SFT-Instruction-Following-Chat-v3.
The reasoning_content field has been fully removed from every message (not just nulled) — each message
now only has role and content. This is intended for training/evaluation setups that do not use
chain-of-thought / reasoning traces.
Source dataset:… See the full description on the dataset page: https://huggingface.co/datasets/tuandunghcmut/nvidia_instruction_following_if_split_v3_non_thinking.qwen3.7-max-split-formatted
Qwen Agent Thinking Online Distillation Rows
This dataset contains cumulative assistant-turn training rows prepared for online logit distillation of Qwen-style agent models, plus a small set of no-tools chat rows to reduce tool-call overbias.
Each row is a rendered-chat-ready conversation prefix ending at a target assistant turn. The trainer uses all prior messages as context and applies loss only to the final assistant span.
Dataset Details
Source trace repo:… See the full description on the dataset page: https://huggingface.co/datasets/armand0e/qwen3.7-max-split-formatted.rukh-puzzles-split
chorcat/rukh-puzzles-split
Lichess puzzles with rating deviation <= 100 and at least 100 plays, banded by difficulty (1000-1500, 1500-2000, 2000+) and split into test and train by a seeded hash of the puzzle id, each with the moves of the game it came from, for tactical evaluation and fine-tuning.
Part of Rukh, a chess language model built from scratch
as a course on generative and agentic AI. Every derived dataset ships with the exact filters and
counts of its manifest.json, so… See the full description on the dataset page: https://huggingface.co/datasets/chorcat/rukh-puzzles-split.nemotron-post-training-samples-splits
Nemotron Post-Training Samples with Train/Val/Test Splits
This dataset contains structured train/validation/test splits from the nvidia/Llama-Nemotron-Post-Training-Dataset, with both tagged and untagged versions for different training scenarios.
Attribution
This work is derived from the Llama-Nemotron-Post-Training-Dataset-v1.1 by NVIDIA Corporation, licensed under CC BY 4.0.
Original Dataset: nvidia/Llama-Nemotron-Post-Training-Dataset
Original Authors: NVIDIA… See the full description on the dataset page: https://huggingface.co/datasets/brandolorian/nemotron-post-training-samples-splits.verl-code-corpus-track-a-file-split
archit11/verl-code-corpus-track-a-file-split
Repository-specific code corpus extracted from the verl project and split by file for training/evaluation.
What is in this dataset
Source corpus: data/code_corpus_verl
Total files: 214
Train files: 172
Validation files: 21
Test files: 21
File type filter: .py
Split mode: file (file-level holdout)
Each row has:
file_name: flattened source file name
text: full file contents
Training context
This dataset was used… See the full description on the dataset page: https://huggingface.co/datasets/archit11/verl-code-corpus-track-a-file-split.wikipedia-id-splits
Dataset: Wikipedia Indonesian (Partial Splits)
Motivasi
Mengembangkan dan mempersiapkan dataset ini untuk fine-tuning bukanlah hal yang mudah, terutama dengan keterbatasan resource yang saya alami. Meski saya sudah berlangganan Colab Pro+ yang menjanjikan akses ke GPU berperforma tinggi (seperti A100 atau H100) dan resource lebih besar, ada beberapa tantangan signifikan yang muncul:
Pembatasan Disk Space Colab VM: Saya sering menghadapi masalah "No space left on device"… See the full description on the dataset page: https://huggingface.co/datasets/nxvay/wikipedia-id-splits.alpaca-train-validation-test-split
Dataset Card for Alpaca
I have just performed train, test and validation split on the original dataset. Repository to reproduce this will be shared here soon. I am including the orignal Dataset card as follows.
Dataset Summary
Alpaca is a dataset of 52,000 instructions and demonstrations generated by OpenAI's text-davinci-003 engine. This instruction data can be used to conduct instruction-tuning for language models and make the language model follow instruction better.… See the full description on the dataset page: https://huggingface.co/datasets/disham993/alpaca-train-validation-test-split.agmind-rag-splitter-ru-data
RU Context-Aware Document Split
Датасет (teacher-distillation) для обучения русского context-aware сплиттера документов для RAG. Каждый пример учит модель где резать документ на самодостаточные смысловые чанки, держа таблицы и код целыми.
Использован для модели AGmind/agmind-rag-splitter-ru. Код генерации и обучения: github.com/botAGI/AGmind-ML.
Формат (Alpaca JSONL)
{
"instruction": "Раздели документ на смысловые части для системы поиска (RAG)...",
"input":… See the full description on the dataset page: https://huggingface.co/datasets/AGmind/agmind-rag-splitter-ru-data.splitsThis dataset contains all four versions of the Splits! dataset, detailed in https://arxiv.org/pdf/2504.04640. See github repo https://github.com/eyloncaplan/splits.
If our dataset is useful for you, please cite us:
@misc{caplan2025splitsflexibledatasetevaluating,
title={Splits! A Flexible Dataset for Evaluating a Model's Demographic Social Inference},
author={Eylon Caplan and Tania Chakraborty and Dan Goldwasser},
year={2025},
eprint={2504.04640}… See the full description on the dataset page: https://huggingface.co/datasets/ecaplan/splits.tool-n1-sft-unique-splits
Tool-N1 SFT Unique with Train/Eval Splits
This dataset contains supervised fine-tuning (SFT) data for training models on multi-hop tool usage and reasoning, with built-in train/evaluation splits.
Usage
from datasets import load_dataset
# Load the dataset with splits
dataset = load_dataset("Anna4242/tool-n1-sft-unique-splits")
# Access splits
train_data = dataset["train"] # 6,487 examples
eval_data = dataset["eval"] # 1,622 examples
# Example usage
for example in… See the full description on the dataset page: https://huggingface.co/datasets/Anna4242/tool-n1-sft-unique-splits.submission14717_fictionalqa_training_splits
Training splits view of the FictionalQA dataset
The FictionalQA dataset
Repository: omitted
Paper: omitted
Dataset Description
This dataset is a derivative of the main dataset. Please see that dataset's README for a detailed description of the assets.
The dataset splits (configs) provided here are the exact ones materialized and used in the experiments for the associated paper. The primary purpose of this dataset repository is for transparency and to help… See the full description on the dataset page: https://huggingface.co/datasets/anonymous-aardvark/submission14717_fictionalqa_training_splits.Amharic_corpus_split
Amharic Corpus — 4 x 5k Splits
A randomly shuffled subset of Reubencf/Amharic_corpus,
divided into four equal splits of 5,000 rows each (20,000 rows total).
Splits: split_1, split_2, split_3, split_4 (5,000 rows each)
Format: JSON Lines, one {"text": "..."} per line.
Sampling: random without replacement (seed 42); the four splits are mutually exclusive.
from datasets import load_dataset
ds = load_dataset("Reubencf/Amharic_corpus_split")
print(ds) # split_1..split_4, 5000 rows… See the full description on the dataset page: https://huggingface.co/datasets/Reubencf/Amharic_corpus_split.sommelier-xlam-single-call-splits-fr
sommelier-xlam-single-call-splits-fr
French paired variant of the single call tool calling rows selected by the Sommelier reference pipeline from Salesforce/xlam-function-calling-60k. Only the user query is translated. Tool schemas and gold answers are byte identical to the English source rows, so the two languages measure the same task with the same scoring.
How it was built
The Sommelier data translate tool (source) translated the exact 17,000 rows the reference… See the full description on the dataset page: https://huggingface.co/datasets/abdelstark/sommelier-xlam-single-call-splits-fr.ccisd-teks-alignment-split
[!WARNING]
Deprecated - use ccisd-teks-alignment instead.
This dataset is superseded: the two contain the same 428 rows with the same 12 columns; this copy only adds a train/validation/test partition, which you can reproduce in one line. Nothing here is unique to it.
It stays online so existing references keep resolving, but it will not be updated.
New work should point at robworks-software/ccisd-teks-alignment.
CCISD TEKS Alignment (pre-split)
The same 428 TEKS-to-course… See the full description on the dataset page: https://huggingface.co/datasets/robworks-software/ccisd-teks-alignment-split.sommelier-xlam-single-call-splits-he-hymt-sanitized
Sommelier xLAM single-call Hebrew paired rows (Hy-MT2, sanitized release)
This CC-BY-4.0 dataset is derived from
Salesforce/xlam-function-calling-60k.
Sommelier filters the source corpus to single-tool-call examples, deterministically
splits it, and machine-translates only each natural-language query into Hebrew.
The exact training snapshot kept tool schemas and gold answers byte-identical to
the English root. For public release, 15 GitHub-PAT-shaped substrings inherited
from… See the full description on the dataset page: https://huggingface.co/datasets/abdelstark/sommelier-xlam-single-call-splits-he-hymt-sanitized.mini_gutenberg_splits
Dataset Card for Mini Project Gutenberg Dataset
This dataset is a mini subset of the dataset nikolina-p/gutenberg_clean_en, created for learning, testing streaming datasets, and quick downloading and manipulation.
It is made from the first 24 books, which are randomly split into 39 shards, mirroring the structure of the original dataset.
The text of the books is randomly split into small chunks, allowing users to experiment with dataset operations on a smaller scale.
This dataset is… See the full description on the dataset page: https://huggingface.co/datasets/nikolina-p/mini_gutenberg_splits.
