datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
muri-it-language-split
MURI-IT: Multilingual Instruction Tuning Dataset for 200 Languages via Multilingual Reverse Instructions
MURI-IT is a large-scale multilingual instruction tuning dataset containing 2.2 million instruction-output pairs across 200 languages. It is designed to address the challenges of instruction tuning in low-resource languages with Multilingual Reverse Instructions (MURI), which ensures that the output is human-written, high-quality, and authentic to the cultural and linguistic… See the full description on the dataset page: https://huggingface.co/datasets/akoksal/muri-it-language-split.mala-monolingual-split
MaLA Corpus: Massive Language Adaptation Corpus
This version contains train and validation splits.
Dataset Summary
The MaLA Corpus (Massive Language Adaptation) is a comprehensive, multilingual dataset designed to support the continual pre-training of large language models. It covers 939 languages and consists of over 74 billion tokens, making it one of the largest datasets of its kind. With a focus on improving the representation of low-resource languages, the… See the full description on the dataset page: https://huggingface.co/datasets/MaLA-LM/mala-monolingual-split.toricgt-curated-splits
ToricGT Curated Graph Reasoning Splits
Curated working dataset repository for ToricGT.
The upload contains only curated split Parquet files and metadata generated locally.
Raw upstream downloads are not uploaded. Each row preserves source dataset, license, split, hashes, and graph JSON fields for audit.
Hebrew/Jewish-text records are sourced from Sefaria and UniMorph Hebrew sources.
Files
train.parquet
validation.parquet
test.parquet
all.parquet if… See the full description on the dataset page: https://huggingface.co/datasets/AmelieSchreiber/toricgt-curated-splits.sft_processed_large_split
sft_processed_large — profile-disjoint split
This is the train / val / test split of Xuhui/sft_processed_large, the
OdysSim midtraining corpus (21.4M interactions across 63 datasets).
Split structure
split
rows
how it's built
train
21.20M
what's left after val + test are carved out
val
28K
per-dataset random sample, in-distribution; for checkpoint selection
test
128K
profile-disjoint where the dataset's profile space supports it; for OOD generalization… See the full description on the dataset page: https://huggingface.co/datasets/Xuhui/sft_processed_large_split.UltraData-Math-L3-Textbook-Exercise-Synthetic-split
UltraData-Math L3 Textbook Exercise Synthetic Split
Source dataset: openbmb/UltraData-Math
Source config: UltraData-Math-L3-Textbook-Exercise-Synthetic
Each row contains:
uid
question
answer
The original content field was split using the literal markers
The exercise: and The solution:.
fictionalqa_training_splits
Training splits view of the FictionalQA dataset
The FictionalQA dataset
Repository: https://github.com/jwkirchenbauer/fictionalqa
Paper: https://arxiv.org/abs/2506.05639
Dataset Description
This dataset is a derivative of the main dataset hf.co/datasets/jwkirchenbauer/fictionalqa. Please see that dataset's README for a detailed description of the assets.
The dataset splits (configs) provided here are the exact ones materialized and used in the experiments for… See the full description on the dataset page: https://huggingface.co/datasets/jwkirchenbauer/fictionalqa_training_splits.gutenberg_clean_en_splits
Dataset Card for Project Gutenberg Cleaned with splits (English Only) Dataset
This dataset is a cleaned English-language subset of the Project Gutenberg Dataset manu/project_gutenberg, originally containing ~70,000 digitized books.
The original dataset includes multiple languages, duplicate entries, and boilerplate content, all of which were removed for practicality and cleaner downstream use.
This dataset containg 38.026 books.
Dataset Splits
The dataset is divided… See the full description on the dataset page: https://huggingface.co/datasets/nikolina-p/gutenberg_clean_en_splits.ashaar-with-enhanced-descriptions-baseform-final-sft-lte20-min500-splits
Ashaar Enhanced Description SFT Stratified Splits
Source dataset:
Shaer-AI/ashaar-with-enhanced-descriptions-baseform-final-sft-lte20-min500
Target dataset:
Shaer-AI/ashaar-with-enhanced-descriptions-baseform-final-sft-lte20-min500-splits
This dataset publishes deterministic train / eval / test splits with a 94 / 3 / 3 policy.
Split policy
Primary stratification key:
base_meter
form
length_bucket
Length buckets:
1-3
4-6
7-10
11-20
Small groups fall back… See the full description on the dataset page: https://huggingface.co/datasets/Shaer-AI/ashaar-with-enhanced-descriptions-baseform-final-sft-lte20-min500-splits.nvidia_instruction_following_if_split_v3
Dataset Description
This is the instruction_following split only (the chat split was intentionally excluded) from
nvidia/Nemotron-SFT-Instruction-Following-Chat-v3,
re-packaged as Parquet (sharded) instead of the original single JSONL file for faster loading and native
support in the HF datasets viewer.
No content was modified — this is a straight format conversion of the instruction_following subset.
Source dataset: nvidia/Nemotron-SFT-Instruction-Following-Chat-v3
Dataset… See the full description on the dataset page: https://huggingface.co/datasets/tuandunghcmut/nvidia_instruction_following_if_split_v3.nvidia_instruction_following_if_split_v3_non_thinking
Dataset Description
Non-thinking (no chain-of-thought) variant of
tuandunghcmut/nvidia_instruction_following_if_split_v3,
which is itself the instruction_following split of
nvidia/Nemotron-SFT-Instruction-Following-Chat-v3.
The reasoning_content field has been fully removed from every message (not just nulled) — each message
now only has role and content. This is intended for training/evaluation setups that do not use
chain-of-thought / reasoning traces.
Source dataset:… See the full description on the dataset page: https://huggingface.co/datasets/tuandunghcmut/nvidia_instruction_following_if_split_v3_non_thinking.rukh-puzzles-split
chorcat/rukh-puzzles-split
Lichess puzzles with rating deviation <= 100 and at least 100 plays, banded by difficulty (1000-1500, 1500-2000, 2000+) and split into test and train by a seeded hash of the puzzle id, each with the moves of the game it came from, for tactical evaluation and fine-tuning.
Part of Rukh, a chess language model built from scratch
as a course on generative and agentic AI. Every derived dataset ships with the exact filters and
counts of its manifest.json, so… See the full description on the dataset page: https://huggingface.co/datasets/chorcat/rukh-puzzles-split.nemotron-post-training-samples-splits
Nemotron Post-Training Samples with Train/Val/Test Splits
This dataset contains structured train/validation/test splits from the nvidia/Llama-Nemotron-Post-Training-Dataset, with both tagged and untagged versions for different training scenarios.
Attribution
This work is derived from the Llama-Nemotron-Post-Training-Dataset-v1.1 by NVIDIA Corporation, licensed under CC BY 4.0.
Original Dataset: nvidia/Llama-Nemotron-Post-Training-Dataset
Original Authors: NVIDIA… See the full description on the dataset page: https://huggingface.co/datasets/brandolorian/nemotron-post-training-samples-splits.verl-code-corpus-track-a-file-split
archit11/verl-code-corpus-track-a-file-split
Repository-specific code corpus extracted from the verl project and split by file for training/evaluation.
What is in this dataset
Source corpus: data/code_corpus_verl
Total files: 214
Train files: 172
Validation files: 21
Test files: 21
File type filter: .py
Split mode: file (file-level holdout)
Each row has:
file_name: flattened source file name
text: full file contents
Training context
This dataset was used… See the full description on the dataset page: https://huggingface.co/datasets/archit11/verl-code-corpus-track-a-file-split.wikipedia-id-splits
Dataset: Wikipedia Indonesian (Partial Splits)
Motivasi
Mengembangkan dan mempersiapkan dataset ini untuk fine-tuning bukanlah hal yang mudah, terutama dengan keterbatasan resource yang saya alami. Meski saya sudah berlangganan Colab Pro+ yang menjanjikan akses ke GPU berperforma tinggi (seperti A100 atau H100) dan resource lebih besar, ada beberapa tantangan signifikan yang muncul:
Pembatasan Disk Space Colab VM: Saya sering menghadapi masalah "No space left on device"… See the full description on the dataset page: https://huggingface.co/datasets/nxvay/wikipedia-id-splits.alpaca-train-validation-test-split
Dataset Card for Alpaca
I have just performed train, test and validation split on the original dataset. Repository to reproduce this will be shared here soon. I am including the orignal Dataset card as follows.
Dataset Summary
Alpaca is a dataset of 52,000 instructions and demonstrations generated by OpenAI's text-davinci-003 engine. This instruction data can be used to conduct instruction-tuning for language models and make the language model follow instruction better.… See the full description on the dataset page: https://huggingface.co/datasets/disham993/alpaca-train-validation-test-split.tool-n1-sft-unique-splits
Tool-N1 SFT Unique with Train/Eval Splits
This dataset contains supervised fine-tuning (SFT) data for training models on multi-hop tool usage and reasoning, with built-in train/evaluation splits.
Usage
from datasets import load_dataset
# Load the dataset with splits
dataset = load_dataset("Anna4242/tool-n1-sft-unique-splits")
# Access splits
train_data = dataset["train"] # 6,487 examples
eval_data = dataset["eval"] # 1,622 examples
# Example usage
for example in… See the full description on the dataset page: https://huggingface.co/datasets/Anna4242/tool-n1-sft-unique-splits.submission14717_fictionalqa_training_splits
Training splits view of the FictionalQA dataset
The FictionalQA dataset
Repository: omitted
Paper: omitted
Dataset Description
This dataset is a derivative of the main dataset. Please see that dataset's README for a detailed description of the assets.
The dataset splits (configs) provided here are the exact ones materialized and used in the experiments for the associated paper. The primary purpose of this dataset repository is for transparency and to help… See the full description on the dataset page: https://huggingface.co/datasets/anonymous-aardvark/submission14717_fictionalqa_training_splits.ccisd-teks-alignment-split
[!WARNING]
Deprecated - use ccisd-teks-alignment instead.
This dataset is superseded: the two contain the same 428 rows with the same 12 columns; this copy only adds a train/validation/test partition, which you can reproduce in one line. Nothing here is unique to it.
It stays online so existing references keep resolving, but it will not be updated.
New work should point at robworks-software/ccisd-teks-alignment.
CCISD TEKS Alignment (pre-split)
The same 428 TEKS-to-course… See the full description on the dataset page: https://huggingface.co/datasets/robworks-software/ccisd-teks-alignment-split.mini_gutenberg_splits
Dataset Card for Mini Project Gutenberg Dataset
This dataset is a mini subset of the dataset nikolina-p/gutenberg_clean_en, created for learning, testing streaming datasets, and quick downloading and manipulation.
It is made from the first 24 books, which are randomly split into 39 shards, mirroring the structure of the original dataset.
The text of the books is randomly split into small chunks, allowing users to experiment with dataset operations on a smaller scale.
This dataset is… See the full description on the dataset page: https://huggingface.co/datasets/nikolina-p/mini_gutenberg_splits.alpaca-train-validation-test-split-50
Dataset Card for Alpaca
I have just performed train, test and validation split on the original dataset. Repository to reproduce this will be shared here soon. I am including the orignal Dataset card as follows.
Dataset Summary
Alpaca is a dataset of 52,000 instructions and demonstrations generated by OpenAI's text-davinci-003 engine. This instruction data can be used to conduct instruction-tuning for language models and make the language model follow instruction… See the full description on the dataset page: https://huggingface.co/datasets/nelsonmaligro/alpaca-train-validation-test-split-50.dz-lahja-dataset-50k-splithigh_educability_training_split
high_educability_training_split
Textos em português selecionados para treinamento: originais de Carolina e Wikipédia classificados nas classes 3 ou 4 pelo educability-norberto-mini-4class-v1, mais as reformulações publicadas vinculadas aos originais elegíveis.
Carregamento
from datasets import load_dataset
ds = load_dataset(
"br-llm-data/high_educability_training_split",
split="train",
streaming=True,
)
registro = next(iter(ds))
Conteúdo… See the full description on the dataset page: https://huggingface.co/datasets/br-llm-data/high_educability_training_split.gutenberg_clean_tokenized_en_splits
Overview
This dataset is a tokenized version of the cleaned English-language subset of the Project Gutenberg Dataset manu/project_gutenberg. It contains full-text books in English, free of boilerplate content and duplicates, and includes a pre-tokenized version of each book's content using the GPT-2 tokenizer (tiktoken.get_encoding("gpt2")).
This dataset is identical to nikolina-p/gutenberg_clean_tokenized_en except for the split configuration.
Cleaning and Preprocessing… See the full description on the dataset page: https://huggingface.co/datasets/nikolina-p/gutenberg_clean_tokenized_en_splits.agustin-guarani-llm-splits
Guarani LLM Splits
This repository contains Parquet splits used for Guarani LLM adaptation experiments.
Files
train.parquet: main training split
synthetic.parquet: synthetic training data
val_id.parquet: in-domain validation split
val_ood.parquet: out-of-domain validation split
test_id.parquet: in-domain test split
test_ood.parquet: out-of-domain test split
Loading
from datasets import load_dataset
repo_id = "agustin-lucas/guarani-llm-splits"… See the full description on the dataset page: https://huggingface.co/datasets/guaran-ia/agustin-guarani-llm-splits.
