datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
github-code-2025-language-split
📜 Source Data & Attribution
This dataset is a processed derivative of nick007x/github-code-2025.
Origination
The original data was aggregated by nick007x from public GitHub repositories. We have retained the original content, file paths, and metadata while restructuring the format for easier consumption by language-specific models.
Processing Steps
To create this dataset, we performed the following processing on the source data:
Language… See the full description on the dataset page: https://huggingface.co/datasets/hasankursun/github-code-2025-language-split.aya_collection_language_split
This is a re-upload of the aya_collection, and only differs in the structure of upload. While the original aya_collection is structured by folders split according to dataset name, this dataset is split by language. We recommend you use this version of the dataset if you are only interested in downloading all of the Aya collection for a single or smaller set of languages.
Dataset Summary
The Aya Collection is a massive multilingual collection consisting of 513 million instances of… See the full description on the dataset page: https://huggingface.co/datasets/CohereLabs/aya_collection_language_split.the-pile-splitted
Dataset description
The pile is an 800GB dataset of english text
designed by EleutherAI to train large-scale language models. The original version of
the dataset can be found here.
The dataset is divided into 22 smaller high-quality datasets. For more information
each of them, please refer to the datasheet for the pile.
However, the current version of the dataset, available on the Hub, is not splitted accordingly.
We had to solve this problem in order to improve the user… See the full description on the dataset page: https://huggingface.co/datasets/ArmelR/the-pile-splitted.muri-it-language-split
MURI-IT: Multilingual Instruction Tuning Dataset for 200 Languages via Multilingual Reverse Instructions
MURI-IT is a large-scale multilingual instruction tuning dataset containing 2.2 million instruction-output pairs across 200 languages. It is designed to address the challenges of instruction tuning in low-resource languages with Multilingual Reverse Instructions (MURI), which ensures that the output is human-written, high-quality, and authentic to the cultural and linguistic… See the full description on the dataset page: https://huggingface.co/datasets/akoksal/muri-it-language-split.fineweb-2-sentence-splitFineweb 2 split into sentences. Instances per languages were sampled by us to balance the data w.r.t. Fineweb-edu.
To split the text into sentences we used the sat3-l model from the wtpsplit library.
We fix a sentence threshold of 0.02 and a maximum sentence length of 256.
If you use this dataset, you should cite:
@misc{penedo2025fineweb2pipelinescale,
title={FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language},
author={Guilherme Penedo and… See the full description on the dataset page: https://huggingface.co/datasets/mimir-lcm/fineweb-2-sentence-split.c4-pt-randMore35M-part04-deduplicated-128000-no-digit-split-mask-train-15003771-lines
Dataset Card for "c4-pt-randMore35M-part04-deduplicated-128000-no-digit-split-mask-train-15003771-lines"
More Information needed
mala-monolingual-split
MaLA Corpus: Massive Language Adaptation Corpus
This version contains train and validation splits.
Dataset Summary
The MaLA Corpus (Massive Language Adaptation) is a comprehensive, multilingual dataset designed to support the continual pre-training of large language models. It covers 939 languages and consists of over 74 billion tokens, making it one of the largest datasets of its kind. With a focus on improving the representation of low-resource languages, the… See the full description on the dataset page: https://huggingface.co/datasets/MaLA-LM/mala-monolingual-split.FalaBracarense_splitsdataset website: projectofalabracarense
Licence
CC - BY - NC - ND
Restrictions: Academic - Non Commercial Use, Attribution, No Derivatives
split_OpenOrca_1M-GPT4-Augmentedtoricgt-curated-splits
ToricGT Curated Graph Reasoning Splits
Curated working dataset repository for ToricGT.
The upload contains only curated split Parquet files and metadata generated locally.
Raw upstream downloads are not uploaded. Each row preserves source dataset, license, split, hashes, and graph JSON fields for audit.
Hebrew/Jewish-text records are sourced from Sefaria and UniMorph Hebrew sources.
Files
train.parquet
validation.parquet
test.parquet
all.parquet if… See the full description on the dataset page: https://huggingface.co/datasets/AmelieSchreiber/toricgt-curated-splits.ai2thor-perspective-qa-20k-balanced-splits-with-objai2thor-perspective-qa-20k-raw-splitsminipile-splitSmaller version of armelr's dataset, which is a
a split version of the general text dataset, "The Pile".
Features of this version
This version has a both a train + test set
Is easily downloadable in ~2.3GB
Can choose text split
If you want a small mixes set, look at minipile instead.
miriad-4.4M-split
MIRIAD 4.4M, split
MIRIAD reformatted for training retrieval
models: train, eval and test splits, and two subsets depending on what you want the model to
retrieve.
subset
columns
use it to retrieve
default
question, passage_text
the source passage a question was generated from (averaging 941 tokens)
question-answer
question, answer
the generated answer to a question (much shorter)
split
rows
train
4,467,542
eval
10,000
test
10,000
[!TIP]… See the full description on the dataset page: https://huggingface.co/datasets/tomaarsen/miriad-4.4M-split.anyedit-splitOpenBible_Swahili_book_splitai2thor-perspective-qa-100k-balanced-training-v1-splitsRSCC-RSEdit-Test-Split
RSCC-RSEdit-Test-Split
This directory contains the test split for RSCC-RSEdit dataset.
Directory Structure
RSCC-RSEdit-Test-Split/
├── images/ # Original images (676 PNG files)
├── masks/ # Original grayscale masks (338 PNG files)
│ └── [mask files with pixel values 0,1,2,3,4]
├── masks_colorful/ # Colorful RGBA visualization masks (338 PNG files)
│ └── [same filenames as masks/, but in RGBA format with colors]
├──… See the full description on the dataset page: https://huggingface.co/datasets/BiliSakura/RSCC-RSEdit-Test-Split.SPACCC_Sentence-Splitter
The Sentence Splitter (SS) for Clinical Cases Written in Spanish
Introduction
This repository contains the sentence splitting model trained using the SPACCC_SPLIT corpus (https://github.com/PlanTL-SANIDAD/SPACCC_SPLIT). The model was trained using the 90% of the corpus (900 clinical cases) and tested against the 10% (100 clinical cases). This model is a great resource to split sentences in biomedical documents, specially clinical cases written in Spanish. This model… See the full description on the dataset page: https://huggingface.co/datasets/Biomedical-TeMU/SPACCC_Sentence-Splitter.CCCPT-splited_preprocessed_max1024sz_sentenceswikipedia-22-12-concat-split
Dataset Card for "wikipedia-22-12-concat-split"
More Information needed
Chinese_Children_Image_Captioning_Dataset_Split0
CODP-1200:Children Oral Description of Picture(Chinese-Child-Captions)
CODP-1200: An AIGC based benchmark for assisting in child language acquisition
数据集介绍
目前已知最大的儿童图像描述数据集,children image captioning
共有1200张图片
每张图片对应五个中文描述,每两张图片为一组
描述文字600*5=3000
如果使用CODP-1200数据集,请引用以下文章
@article{LENG2024102627,
title = {CODP-1200: An AIGC based benchmark for assisting in child language acquisition},
journal = {Displays},
volume = {82},
pages = {102627},
year =… See the full description on the dataset page: https://huggingface.co/datasets/svjack/Chinese_Children_Image_Captioning_Dataset_Split0.chinese-fineweb-edu-v2_splitted_1_filtered_wo_engtofu_custom_split_ESUmvtec_all_objects_splitr-asts-splitted-tokenizedcommon_voice_10_1_th_clean_split_0_old
Dataset Card for "common_voice_10_1_th_clean_split_0"
More Information needed
common_voice_10_1_th_clean_split_1
Dataset Card for "common_voice_10_1_th_clean_split_1_fix_spacial_char"
More Information needed
WMT-month-splitscrossref_metadata_2025_split
Dataset Overview
This dataset contains bibliographic metadata from the public Crossref snapshot released in 2025. It provides core fields for scholarly documents, including DOI, title, abstract, authorship, publication month and year, and URLs. The entire public dump (~196.94 GB) was filtered and extracted into a parquet format for efficient loading and querying.
Total size: 196.94 GB (parquet files)
Number of records: 34,308,730
Use this dataset for large-scale text mining… See the full description on the dataset page: https://huggingface.co/datasets/bluuebunny/crossref_metadata_2025_split.
