datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
muri-it-language-split
MURI-IT: Multilingual Instruction Tuning Dataset for 200 Languages via Multilingual Reverse Instructions
MURI-IT is a large-scale multilingual instruction tuning dataset containing 2.2 million instruction-output pairs across 200 languages. It is designed to address the challenges of instruction tuning in low-resource languages with Multilingual Reverse Instructions (MURI), which ensures that the output is human-written, high-quality, and authentic to the cultural and linguistic… See the full description on the dataset page: https://huggingface.co/datasets/akoksal/muri-it-language-split.Eurovoc_2025_by_language
🇪🇺 🏷️ EuroVoc dataset (by language)
This is the EuropeanParliament/Eurovoc_2025 dataset, but split up by language, not by period.
The original is split up into periods (1996-03 through 2025-11), with documents in different languages mixed together.
For ease of training this dataset splits the data by language instead, with documents in different periods put together.
License
This dataset is redistributed under the original European Union Public License 1.2. When… See the full description on the dataset page: https://huggingface.co/datasets/AIStudioDelta/Eurovoc_2025_by_language.raw-text-corpus
📝 Zomi Raw Text Corpus (Community-Contributed)
The Zomi Raw Text Corpus is an open, community-driven collection of unprocessed Zomi-language text.It serves as the foundational dataset for building the full Zomi NLP ecosystem, including tokenizers, language models, ASR/TTS systems, and downstream tasks.
This dataset is intentionally raw — no normalization, deduplication, or cleaning is applied.Cleaned and task-specific datasets will be released separately.
📦 Dataset… See the full description on the dataset page: https://huggingface.co/datasets/zomi-language-corpora/raw-text-corpus.BiasShadesInterested in contributing? Speak a language not represented here? Disagree with an annotation? Please submit feedback in the Community tab!
Dataset Card for BiasShades
Note: This dataset may NOT be used as training data in any form (pre-training, fine-tuning, post-training, etc.) without express permission from creators.
Dataset Details
Version: 1.0
License: SHADES 1 Montreal Data License
Dataset Description
728 stereotypes and associated… See the full description on the dataset page: https://huggingface.co/datasets/LanguageShades/BiasShades.nepalitext-language-model-dataset
Dataset Card for "nepalitext-language-model-dataset"
Dataset Summary
"NepaliText" language modeling dataset is a collection of over 13 million Nepali text sequences (phrases/sentences/paragraphs) extracted by combining the datasets: OSCAR , cc100 and a set of scraped Nepali articles on Wikipedia.
Supported Tasks and Leaderboards
This dataset is intended to pre-train language models and word representations on Nepali Language.
Languages
The data is… See the full description on the dataset page: https://huggingface.co/datasets/Sakonii/nepalitext-language-model-dataset.Kashmiri-language-text-datasetThis is a Dataset for Kashmiri language containing Kashmiri words their phonemes , along with their meaning in ENGlISH and HINDI
and an example sentence in english . The phoenmes are written phonemes present in wordphonemes-meaning.csv
language-decoded-data
Language Decoded | Multilingual Code Dataset
Experiment and proposed paper title: Language Decoded: Exploring the Impact of Native Code on Multilingual Models
Note (2026-05-18): Current Phase 3 configs use the short condition-* namespace and include 103k, 20k, and 5k sizes for Conditions 1--2. Phase 2 configs remain available under the phase-2-the-stack-v1-* namespace for reproducibility.
Multilingual Python code datasets for the Language Decoded project (part of Cohere's… See the full description on the dataset page: https://huggingface.co/datasets/legesher/language-decoded-data.wikipedia-language-snippets-filtered
Wikipedia Snippets (Filtered)
Filtered sentence snippets in Wikipedia, by taking the first 60% of an article after filtering for stubs. Minor Latin groups are additionally filtered again for English leakage.
Sentences are mostly filtered out for non matching scripts, such as Arabic in a Cyrllic language.
Files
Each file is in this format for languages in ISO 639 2-letter codes:
train/en/en.parquet
train/es/es.parquet
From wikimedia/wikipedia
Licensing… See the full description on the dataset page: https://huggingface.co/datasets/polyglot-tagger/wikipedia-language-snippets-filtered.natural-language-to-mongosh
Natural Language to MongoDB Shell (mongosh) Benchmark
Benchmark dataset for performing natural language (NL) to MongoDB Shell (mongosh) code generation.
There is an emerging desire from users for NL query generation.
This benchmarks examines how LLMs generate MongoDB queries and provides proactive guidance for making systems that map NL to MongoDB queries.
Repository Contents
This repository contains:
Benchmark dataset (flat CSV file, Braintrust evaluation… See the full description on the dataset page: https://huggingface.co/datasets/mongodb-eai/natural-language-to-mongosh.yuxiaowang-prompts-2025
Yuxiaowang Semantic Dataset · Hugging Face Version
🧠 English Summary
Yuxiaowang · Semantic Dataset for Japanese Language Schools (Chinese)
This project provides structured semantic definitions and prompt examples for the domain of Japanese language schools in China.It aims to serve as a grounding corpus for large language models (LLMs) to understand terms like "语校", "语校网", and related concepts.
Source platform: https://www.yuxiaowang.comAll prompts and term… See the full description on the dataset page: https://huggingface.co/datasets/languagehub-ai/yuxiaowang-prompts-2025.task1577_amazon_reviews_multi_japanese_language_classification
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1577_amazon_reviews_multi_japanese_language_classification
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1577_amazon_reviews_multi_japanese_language_classification.Natural_Language_to_Ffmpeg_Commands
Natural Language to FFmpeg Dataset
Disclaimer: This dataset was synthetically generated using a large language model and is intended for research purposes only. The dataset may contain inaccuracies, errors, or inconsistencies. Users should exercise caution and verify the correctness of the data before using it in any application.
This dataset contains 1000+ pairs of English natural language instructions and corresponding FFmpeg commands.
The dataset is designed for tasks… See the full description on the dataset page: https://huggingface.co/datasets/burak29/Natural_Language_to_Ffmpeg_Commands.swahili-language-exposure
swahili-language-exposure
Dataset Summary
swahili-language-exposure is a large-scale Swahili (Kiswahili) corpus designed for language exposure and continued pretraining of language models.
Unlike instruction-tuning datasets, this dataset focuses on exposing models to natural Swahili usage across conversations, explanations, narratives, technical discussions, and mixed-domain text. The goal is to improve fluency, vocabulary coverage, syntax, and cultural grounding in… See the full description on the dataset page: https://huggingface.co/datasets/nileagi/swahili-language-exposure.task427_hindienglish_corpora_hi-en_language_identification
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task427_hindienglish_corpora_hi-en_language_identification
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task427_hindienglish_corpora_hi-en_language_identification.uzbek-language-dataset
Uzbek Language Dataset Collection
Bu repository o'zbek tili uchun eng keng ko'lamli va keng qamrovli dataset to'plami hisoblanadi. Dataset turli manbalardan to'plangan va NLP modellari, til modellari va boshqa AI ilovalar uchun mo'ljallangan.
📊 Dataset Overview
Bu dataset to'plami 4ta asosiy qism va qo'shimcha merge qilish asboblaridan iborat:
🎯 Dataset Qismlari
Dataset
Hajmi
Maqsad
Source
community-oscar-uzbek
1.1GB
OSCAR Community data
Common… See the full description on the dataset page: https://huggingface.co/datasets/xkas2001/uzbek-language-dataset.HuggingFaceFW-finetranslations-100-languages-sample
Finetranslations 100 Language Sample Dataset
Subset of HuggingFaceFW/finetranslations with the top 100 languages by number of documents.
Configurations
all: 100 languages combined (100k rows), shuffled
100 individual language configs: 1000 rows each
Columns
Original columns + language (source language indicator which is the name of the config)
Usage
from datasets import load_dataset
dataset =… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/HuggingFaceFW-finetranslations-100-languages-sample.nepalitext-language-model-dataset
Dataset Card for "nepalitext-language-model-dataset"
Dataset Summary
"NepaliText" language modeling dataset is a collection of over 13 million Nepali text sequences (phrases/sentences/paragraphs) extracted by combining the datasets: OSCAR , cc100 and a set of scraped Nepali articles on Wikipedia.
Supported Tasks and Leaderboards
This dataset is intended to pre-train language models and word representations on Nepali Language.
Languages
The data is… See the full description on the dataset page: https://huggingface.co/datasets/Arpuuu/nepalitext-language-model-dataset.swahili-language-exposure-v2
Swahili Language Exposure
Large-scale Swahili corpus for continued pretraining and language exposure.
Maintained by NileAGI.
multimodal-vision-language-video-models-2026
👁️ Multimodal Vision-Language & Video Foundation Models Dataset (2026 Edition)
A structured research dataset featuring 1,000 domain-verified research papers and code repositories focused on Multimodal Vision-Language Models (VLM), Video Foundation Models, Diffusion Transformers (DiT), Visual Grounding, and World Simulators.
Built with Universal Scientific Engine V15.1 Gold, providing 47 schema attributes with verified repository attribution, modality capability matrix, vision… See the full description on the dataset page: https://huggingface.co/datasets/beatsprom/multimodal-vision-language-video-models-2026.Tumbuka_language
About this dataset
This dataset mainly focuses on Tumbuka Language, found in Northern Malawi and Zambia.
Usecases
mainly focuses on datasets that are to be used for fine-tuning already existing AI Models, so that they are able to understand the Tumbuka Bantu Language (Malawi & Zambia & Tanzania).
Formats
The datasets are in different formats, and sometimes you will notice that the same dataset, have been uploaded with several file formats like .txt… See the full description on the dataset page: https://huggingface.co/datasets/Mwanzau/Tumbuka_language.sa-languages-corpus
South African Languages Text Corpus
Plain-text corpus covering all 11 official South African languages, for language modeling.
every-language-dataset-v3
Every Language Dataset V3
Next-generation synthetic multilingual dataset.
Size
Total: 25,000,000
Train: 24,000,000
Validation: 500,000
Test: 500,000
Diversity
Human-language catalog:
171 language codes.
Programming languages:
50.
Task families:
conversation
question answering
reasoning
logic
arithmetic
translation
summarization
explanation
code generation
code explanation
debugging
Format
Parquet + ZSTD.
Generation… See the full description on the dataset page: https://huggingface.co/datasets/Lelonthecodeur/every-language-dataset-v3.bluesky-10m-posts-15-languages
Dataset Card: Bluesky 10M Multilingual
📊 Overview
Total Posts: 10,099,990
Languages: 15 (en, tr, es, pt, de, fr, ja, it, nl, pl, ru, ko, zh, ar, hi)
Collection Period: August 9-12, 2026
Source: Bluesky Jetstream API (public firehose)
Format: JSONL
Size: ~3 GB
🌍 Language Distribution
Language
Code
Posts
%
English
en
6,843,995
67.8%
Japanese
ja
1,547,179
15.3%
German
de
373,626
3.7%
Portuguese
pt
331,093
3.3%
Spanish
es
325,865… See the full description on the dataset page: https://huggingface.co/datasets/itsmebatuhan/bluesky-10m-posts-15-languages.git-natural-language-commands
Git Natural Language Commands
A dataset mapping English natural-language instructions to their corresponding git commands, intended for training and evaluating models that translate user intent into safe, correct shell commands.
Disclaimer: This dataset was generated using Large Language Models (LLMs). The examples have not been manually verified against real-world usage and may contain errors, inconsistencies, or non-canonical phrasings. Use with appropriate caution.… See the full description on the dataset page: https://huggingface.co/datasets/burak29/git-natural-language-commands.CSDN-C_Language-2013_2023CSDN - C 语言社区 2013 ~ 2023.10.2 的问答数据,未包含图片,仅有文本内容。
共 29K+ 条,数据已经经过初步清洗和脱敏,去除了所有 0 回复的贴子 & 机器人回复的贴子。为了方便不同使用目的,按照回复盖楼的格式对数据进行了组织,一个样例(展开后)如下:
{
"question": "刚学C语言,为什么这个代码运行不了呢",
"poster": "user-0",
"comments": [
{
"cid": "2",
"user": "user-2",
"content": "intunsigned intlong longunsigned long long统统容纳不下29的阶乘,早就溢出了。",
"referer": "user-0"
},
{
"cid": "3",
"user": "user-3"… See the full description on the dataset page: https://huggingface.co/datasets/Mxode/CSDN-C_Language-2013_2023.Tunisian_Language_Dataset
Dataset Card for Tunisian Text Compilation
This dataset is a curated compilation of various Tunisian datasets, aimed at gathering as much Tunisian text data as possible in one place. It combines multiple sources of Tunisian language data, providing a rich resource for research, development of NLP models, and linguistic studies on Tunisian text.
Dataset Details
Dataset Description
This dataset aggregates several publicly available datasets that contain Tunisian… See the full description on the dataset page: https://huggingface.co/datasets/AzizBelaweid/Tunisian_Language_Dataset.language-decoded-community
Language Decoded — Community Code
Natively-authored multilingual code for the Language Decoded project (part of Cohere's Tiny Aya Expedition). This dataset contains code written by developers in non-English programming languages and code with significant CJK content — not mechanically transpiled or LLM-translated from English.
Experiment and proposed paper title: Language Decoded: Exploring the Impact of Native Code on Multilingual Models
This data serves as the corpus for… See the full description on the dataset page: https://huggingface.co/datasets/legesher/language-decoded-community.gsm8k-translated
Multilingual GSM8K Translations
This dataset contains machine-translated versions of GSM8K in these languages:
French (fr)
German (de)
Hindi (hi)
Dataset Structure
For each language, we provide the original GSM8K train and test splits:
train: 7,473 samples
test: 1,319 samples
Each sample consists of a question and an answer.
The question describes a grade-school-level math word problem that requires multi-step mathematical reasoning. The answer contains a… See the full description on the dataset page: https://huggingface.co/datasets/math-across-languages/gsm8k-translated.Chinese-StackOverflow-QA-C_Language
中文 StackOverflow C 语言问答数据集
💻 Github Repo
基本信息
本数据集提供了两个子集:
translated:原数据集 Mxode/StackOverflow-QA-C-Language-40k 的中文翻译版本,数量约 40K。
synthetic **(Default)**:在原数据集 Mxode/StackOverflow-QA-C-Language-40k 的基础上,重新扩充、合成的问答数据集,数量约 200K。
数据格式
请注意:两个子集的数据格式并不完全相同。
translated 子集:
{
"id": << 12位nanoid >>,
"question_en": << 用户提问(英文) >>,
"question_zh": << 用户提问(中文) >>,
"answer_en": << 用户回答(英文) >>,
"answer_zh": << 用户回答(中文) >>,
}
synthetic 子集:
{
"id": <<… See the full description on the dataset page: https://huggingface.co/datasets/Mxode/Chinese-StackOverflow-QA-C_Language.time-series-language-alignment
TS-Insights Dataset
Dataset Description
TS-Insights is the official dataset for the paper "Insight Miner: A Time Series Analysis Dataset for Cross-Domain Alignment with Natural Language". This work is done by Project Mineral from Google X in 2023.
It is the first large-scale general-domain dataset designed to align time-series data with natural language descriptions. The dataset supports the training of Large Multimodal Models (LMMs) to understand time series as a new… See the full description on the dataset page: https://huggingface.co/datasets/zhykoties/time-series-language-alignment.
