datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Open-Sora-Plan-v1.1.0
Annotation
We resized the dataset to 1080p for easier uploading. Therefore, the original annotation file might not match the video names. Please refer to this https://github.com/PKU-YuanGroup/Open-Sora-Plan/issues/312#issuecomment-2197312973
Pexels
Pexels consists of multiple folders, but each folder exceeds the size limit for Huggingface uploads. Therefore, we divided each folder into 5 parts. You need to merge the 5 parts of each folder first, and then extract each… See the full description on the dataset page: https://huggingface.co/datasets/LanguageBind/Open-Sora-Plan-v1.1.0.github-code-2025-language-split
📜 Source Data & Attribution
This dataset is a processed derivative of nick007x/github-code-2025.
Origination
The original data was aggregated by nick007x from public GitHub repositories. We have retained the original content, file paths, and metadata while restructuring the format for easier consumption by language-specific models.
Processing Steps
To create this dataset, we performed the following processing on the source data:
Language… See the full description on the dataset page: https://huggingface.co/datasets/hasankursun/github-code-2025-language-split.aya_collection_language_split
This is a re-upload of the aya_collection, and only differs in the structure of upload. While the original aya_collection is structured by folders split according to dataset name, this dataset is split by language. We recommend you use this version of the dataset if you are only interested in downloading all of the Aya collection for a single or smaller set of languages.
Dataset Summary
The Aya Collection is a massive multilingual collection consisting of 513 million instances of… See the full description on the dataset page: https://huggingface.co/datasets/CohereLabs/aya_collection_language_split.muri-it-language-split
MURI-IT: Multilingual Instruction Tuning Dataset for 200 Languages via Multilingual Reverse Instructions
MURI-IT is a large-scale multilingual instruction tuning dataset containing 2.2 million instruction-output pairs across 200 languages. It is designed to address the challenges of instruction tuning in low-resource languages with Multilingual Reverse Instructions (MURI), which ensures that the output is human-written, high-quality, and authentic to the cultural and linguistic… See the full description on the dataset page: https://huggingface.co/datasets/akoksal/muri-it-language-split.Language-Grounded_Sparse_Encoder_Training
Language-Grounded Sparse Encoder (LanSE) — Training Data
This repository hosts the AI-generated images and human annotation datasets accompanying the paper:
Human-like Content Analysis for Generative AI with Language-Grounded Sparse Encoders
Yiming Tang, Arash Lagzian, Srinivas Anumasa, Qiran Zou, Yingtao Zhu, Ye Zhang, Trang Nguyen, Yih-Chung Tham, Ehsan Adeli, Ching-Yu Cheng, Yilun Du, Dianbo Liu
National University of Singapore · Tsinghua University · Stanford University ·… See the full description on the dataset page: https://huggingface.co/datasets/DesmondYMTang2024/Language-Grounded_Sparse_Encoder_Training.language
Dataset Card for "livebench/language"
LiveBench is a benchmark for LLMs designed with test set contamination and objective evaluation in mind. It has the following properties:
LiveBench is designed to limit potential contamination by releasing new questions monthly, as well as having questions based on recently-released datasets, arXiv papers, news articles, and IMDb movie synopses.
Each question has verifiable, objective ground-truth answers, allowing hard questions to be scored… See the full description on the dataset page: https://huggingface.co/datasets/livebench/language.language-identification
Dataset Card for Language Identification dataset
Dataset Summary
The Language Identification dataset is a collection of 90k samples consisting of text passages and corresponding language label.
This dataset was created by collecting data from 3 sources: Multilingual Amazon Reviews Corpus, XNLI, and STSb Multi MT.
Supported Tasks and Leaderboards
The dataset can be used to train a model for language identification, which is a multi-class text classification… See the full description on the dataset page: https://huggingface.co/datasets/papluca/language-identification.UniWorld-V1
The Geneval-style dataset is sourced from BLIP3o-60k.
This dataset is presented in the paper: UniWorld: High-Resolution Semantic Encoders for Unified Visual Understanding and Generation
More details can be found in UniWorld-V1
Data preparation
Download the data from LanguageBind/UniWorld-V1. The dataset consists of two parts: source images and annotation JSON files.
Prepare a data.txt file in the following format:
The first column is the root path to the image.
The second… See the full description on the dataset page: https://huggingface.co/datasets/LanguageBind/UniWorld-V1.Language-DetectionLanguage_Detection
Language_Detection - Multilingual Text Classification Dataset
This dataset is a collection of multilingual text samples designed for training and predicting languages in Artificial Intelligence (AI), Machine Learning (ML), Deep Learning (DL), and Data Science (DS) applications. It contains labeled data that associates text samples with their respective languages, enabling language detection and classification tasks.
Dataset Overview
The dataset consists of two columns:… See the full description on the dataset page: https://huggingface.co/datasets/sakthivinash/Language_Detection.low_resource_languages_pretrain_data5Open-Sora-Plan-v1.0.0
Open-Sora-Dataset
Welcome to the Open-Sora-DataSet project! As part of the Open-Sora-Plan project, we specifically talk about the collection and processing of data sets. To build a high-quality video dataset for the open-source world, we started this project. 💪
We warmly welcome you to join us! Let's contribute to the open-source world together! Thank you for your support and contribution.
If you like our project, please give us a star ⭐ on GitHub for latest update.… See the full description on the dataset page: https://huggingface.co/datasets/LanguageBind/Open-Sora-Plan-v1.0.0.low_resource_languages_pretrain_data2ethiopian-languages-speech-dataset
Leyu Ethiopian Languages Speech Dataset
Audio recordings paired with corresponding text transcripts, collected on the Leyu Data Collection Platform — an open-source platform for crowdsourced speech data collection — for the Leyu Platform Competition, covering 4 languages: Amharic, Afaan Oromo, Sidama, Tigrinya.
Dataset Summary
Languages: Amharic (am), Afaan Oromo (om), Sidama (sid), Tigrinya (ti)
Total examples: 2750
License: CC-BY-4.0
Task categories: Automatic… See the full description on the dataset page: https://huggingface.co/datasets/Benji-fish/ethiopian-languages-speech-dataset.language_table_inpainting
language_table robot-removal inpainting dataset
This dataset contains robot-removal inpainting results for language_table. Each episode provides:
inpainting.mp4: the robot visually removed via inpainting
mask.mp4: the robot mask video used for inpainting
original_episode.mp4: the original (unmodified) episode video
language_instructions_{split}_all.txt: tab-separated mapping from episode_id to instruction
Relation to OXE-AugE
This release is produced as part of… See the full description on the dataset page: https://huggingface.co/datasets/oxe-auge/language_table_inpainting.low_resource_languages_pretrain_data8common_language_preprocessed
Dataset Card for "common_language_preprocessed"
More Information needed
language-detection
Dataset Card for "language-detection"
More Information needed
low_resource_languages_pretrainEurovoc_2025_by_language
🇪🇺 🏷️ EuroVoc dataset (by language)
This is the EuropeanParliament/Eurovoc_2025 dataset, but split up by language, not by period.
The original is split up into periods (1996-03 through 2025-11), with documents in different languages mixed together.
For ease of training this dataset splits the data by language instead, with documents in different periods put together.
License
This dataset is redistributed under the original European Union Public License 1.2. When… See the full description on the dataset page: https://huggingface.co/datasets/AIStudioDelta/Eurovoc_2025_by_language.mixed_multilingual_commonvoice_all_languages_100kBuild from mozilla commonvoice 13 using the script commited in this repo.
Used to teach a model to ignore languages that are not french
europarl_for_language_detection_10kraw-text-corpus
📝 Zomi Raw Text Corpus (Community-Contributed)
The Zomi Raw Text Corpus is an open, community-driven collection of unprocessed Zomi-language text.It serves as the foundational dataset for building the full Zomi NLP ecosystem, including tokenizers, language models, ASR/TTS systems, and downstream tasks.
This dataset is intentionally raw — no normalization, deduplication, or cleaning is applied.Cleaned and task-specific datasets will be released separately.
📦 Dataset… See the full description on the dataset page: https://huggingface.co/datasets/zomi-language-corpora/raw-text-corpus.samromur_childrenThe Samrómur Children corpus contains more than 137000 validated speech-recordings uttered by Icelandic children.BiasShadesInterested in contributing? Speak a language not represented here? Disagree with an annotation? Please submit feedback in the Community tab!
Dataset Card for BiasShades
Note: This dataset may NOT be used as training data in any form (pre-training, fine-tuning, post-training, etc.) without express permission from creators.
Dataset Details
Version: 1.0
License: SHADES 1 Montreal Data License
Dataset Description
728 stereotypes and associated… See the full description on the dataset page: https://huggingface.co/datasets/LanguageShades/BiasShades.low_resource_languages_pretrain_data4nepalitext-language-model-dataset
Dataset Card for "nepalitext-language-model-dataset"
Dataset Summary
"NepaliText" language modeling dataset is a collection of over 13 million Nepali text sequences (phrases/sentences/paragraphs) extracted by combining the datasets: OSCAR , cc100 and a set of scraped Nepali articles on Wikipedia.
Supported Tasks and Leaderboards
This dataset is intended to pre-train language models and word representations on Nepali Language.
Languages
The data is… See the full description on the dataset page: https://huggingface.co/datasets/Sakonii/nepalitext-language-model-dataset.language_detection_trainneuromorphic-event-language-bridge
Neuromorphic Event-Language Bridge
Rights & intended use: legacy public research corpus / portfolio
artifact. Hosted frontier-model outputs are research-only inputs under
project policy (synthetic-factory#161):
intended_use: research_only, project_training_policy: blocked. Not
training data for any model-weight update. Machine-readable record:
rights.json.
Release status: The raw, uncurated event-language payload is now
published under data/raw/. It is available for… See the full description on the dataset page: https://huggingface.co/datasets/rmems/neuromorphic-event-language-bridge.Marco-Bench-MIF-languages
