datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
github-code-2025-language-split
📜 Source Data & Attribution
This dataset is a processed derivative of nick007x/github-code-2025.
Origination
The original data was aggregated by nick007x from public GitHub repositories. We have retained the original content, file paths, and metadata while restructuring the format for easier consumption by language-specific models.
Processing Steps
To create this dataset, we performed the following processing on the source data:
Language… See the full description on the dataset page: https://huggingface.co/datasets/hasankursun/github-code-2025-language-split.aya_collection_language_split
This is a re-upload of the aya_collection, and only differs in the structure of upload. While the original aya_collection is structured by folders split according to dataset name, this dataset is split by language. We recommend you use this version of the dataset if you are only interested in downloading all of the Aya collection for a single or smaller set of languages.
Dataset Summary
The Aya Collection is a massive multilingual collection consisting of 513 million instances of… See the full description on the dataset page: https://huggingface.co/datasets/CohereLabs/aya_collection_language_split.muri-it-language-split
MURI-IT: Multilingual Instruction Tuning Dataset for 200 Languages via Multilingual Reverse Instructions
MURI-IT is a large-scale multilingual instruction tuning dataset containing 2.2 million instruction-output pairs across 200 languages. It is designed to address the challenges of instruction tuning in low-resource languages with Multilingual Reverse Instructions (MURI), which ensures that the output is human-written, high-quality, and authentic to the cultural and linguistic… See the full description on the dataset page: https://huggingface.co/datasets/akoksal/muri-it-language-split.fineweb-2-sentence-splitFineweb 2 split into sentences. Instances per languages were sampled by us to balance the data w.r.t. Fineweb-edu.
To split the text into sentences we used the sat3-l model from the wtpsplit library.
We fix a sentence threshold of 0.02 and a maximum sentence length of 256.
If you use this dataset, you should cite:
@misc{penedo2025fineweb2pipelinescale,
title={FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language},
author={Guilherme Penedo and… See the full description on the dataset page: https://huggingface.co/datasets/mimir-lcm/fineweb-2-sentence-split.c4-pt-randMore35M-part04-deduplicated-128000-no-digit-split-mask-train-15003771-lines
Dataset Card for "c4-pt-randMore35M-part04-deduplicated-128000-no-digit-split-mask-train-15003771-lines"
More Information needed
mala-monolingual-split
MaLA Corpus: Massive Language Adaptation Corpus
This version contains train and validation splits.
Dataset Summary
The MaLA Corpus (Massive Language Adaptation) is a comprehensive, multilingual dataset designed to support the continual pre-training of large language models. It covers 939 languages and consists of over 74 billion tokens, making it one of the largest datasets of its kind. With a focus on improving the representation of low-resource languages, the… See the full description on the dataset page: https://huggingface.co/datasets/MaLA-LM/mala-monolingual-split.split_OpenOrca_1M-GPT4-Augmentedtoricgt-curated-splits
ToricGT Curated Graph Reasoning Splits
Curated working dataset repository for ToricGT.
The upload contains only curated split Parquet files and metadata generated locally.
Raw upstream downloads are not uploaded. Each row preserves source dataset, license, split, hashes, and graph JSON fields for audit.
Hebrew/Jewish-text records are sourced from Sefaria and UniMorph Hebrew sources.
Files
train.parquet
validation.parquet
test.parquet
all.parquet if… See the full description on the dataset page: https://huggingface.co/datasets/AmelieSchreiber/toricgt-curated-splits.ai2thor-perspective-qa-20k-balanced-splits-with-objai2thor-perspective-qa-20k-raw-splitsshelf_bin_long_splitThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.0",
"robot_type": "easo",
"total_episodes": 2268,
"total_frames": 744700,
"total_tasks": 2,
"total_videos": 0,
"total_chunks": 3,
"chunks_size": 1000,
"fps": 50,
"splits": {
"train": "0:2268"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/willx0909/shelf_bin_long_split.miriad-4.4M-split
MIRIAD 4.4M, split
MIRIAD reformatted for training retrieval
models: train, eval and test splits, and two subsets depending on what you want the model to
retrieve.
subset
columns
use it to retrieve
default
question, passage_text
the source passage a question was generated from (averaging 941 tokens)
question-answer
question, answer
the generated answer to a question (much shorter)
split
rows
train
4,467,542
eval
10,000
test
10,000
[!TIP]… See the full description on the dataset page: https://huggingface.co/datasets/tomaarsen/miriad-4.4M-split.anyedit-splitOpenBible_Swahili_book_splitai2thor-perspective-qa-100k-balanced-training-v1-splitswikipedia-22-12-concat-split
Dataset Card for "wikipedia-22-12-concat-split"
More Information needed
chinese-fineweb-edu-v2_splitted_1_filtered_wo_engmvtec_all_objects_splitr-asts-splitted-tokenizedcommon_voice_10_1_th_clean_split_0_old
Dataset Card for "common_voice_10_1_th_clean_split_0"
More Information needed
open-images-v7-subset-splitscommon_voice_10_1_th_clean_split_1
Dataset Card for "common_voice_10_1_th_clean_split_1_fix_spacial_char"
More Information needed
WMT-month-splitscrossref_metadata_2025_split
Dataset Overview
This dataset contains bibliographic metadata from the public Crossref snapshot released in 2025. It provides core fields for scholarly documents, including DOI, title, abstract, authorship, publication month and year, and URLs. The entire public dump (~196.94 GB) was filtered and extracted into a parquet format for efficient loading and querying.
Total size: 196.94 GB (parquet files)
Number of records: 34,308,730
Use this dataset for large-scale text mining… See the full description on the dataset page: https://huggingface.co/datasets/bluuebunny/crossref_metadata_2025_split.common_voice_13_0_bn_multi_splitcommon_voice_10_1_th_clean_split_0
Dataset Card for "common_voice_10_1_th_clean_split_0_fix_spacial_char"
More Information needed
nabirds_custom_split_preprocessedsft_processed_large_split
sft_processed_large — profile-disjoint split
This is the train / val / test split of Xuhui/sft_processed_large, the
OdysSim midtraining corpus (21.4M interactions across 63 datasets).
Split structure
split
rows
how it's built
train
21.20M
what's left after val + test are carved out
val
28K
per-dataset random sample, in-distribution; for checkpoint selection
test
128K
profile-disjoint where the dataset's profile space supports it; for OOD generalization… See the full description on the dataset page: https://huggingface.co/datasets/Xuhui/sft_processed_large_split.SQaLe_2_Splitfineweb-edu-350BT-sentence-splitFineweb-edu 350BT subset split into sentences.
To split the text into sentences we used the sat3-l model from the wtpsplit library.
We fix a sentence threshold of 0.02 and a maximum sentence length of 256.
If you use this dataset, you should cite:
@misc{lozhkov2024fineweb-edu,
author = { Lozhkov, Anton and Ben Allal, Loubna and von Werra, Leandro and Wolf, Thomas },
title = { FineWeb-Edu: the Finest Collection of Educational Content },
year = 2024… See the full description on the dataset page: https://huggingface.co/datasets/mimir-lcm/fineweb-edu-350BT-sentence-split.
