CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01akoksal /muri-it-language-split MURI-IT: Multilingual Instruction Tuning Dataset for 200 Languages via Multilingual Reverse Instructions MURI-IT is a large-scale multilingual instruction tuning dataset containing 2.2 million instruction-output pairs across 200 languages. It is designed to address the challenges of instruction tuning in low-resource languages with Multilingual Reverse Instructions (MURI), which ensures that the output is human-written, high-quality, and authentic to the cultural and linguistic… See the full description on the dataset page: https://huggingface.co/datasets/akoksal/muri-it-language-split.texttext-generation1M<n<10M6 likes11k downloads2y agoHugging Face02AIStudioDelta /Eurovoc_2025_by_language 🇪🇺 🏷️ EuroVoc dataset (by language) This is the EuropeanParliament/Eurovoc_2025 dataset, but split up by language, not by period. The original is split up into periods (1996-03 through 2025-11), with documents in different languages mixed together. For ease of training this dataset splits the data by language instead, with documents in different periods put together. License This dataset is redistributed under the original European Union Public License 1.2. When… See the full description on the dataset page: https://huggingface.co/datasets/AIStudioDelta/Eurovoc_2025_by_language.texttext-generation1M<n<10M1 likes688 downloads10mo agoHugging Face03zomi-language-corpora /raw-text-corpus 📝 Zomi Raw Text Corpus (Community-Contributed) The Zomi Raw Text Corpus is an open, community-driven collection of unprocessed Zomi-language text.It serves as the foundational dataset for building the full Zomi NLP ecosystem, including tokenizers, language models, ASR/TTS systems, and downstream tasks. This dataset is intentionally raw — no normalization, deduplication, or cleaning is applied.Cleaned and task-specific datasets will be released separately. 📦 Dataset… See the full description on the dataset page: https://huggingface.co/datasets/zomi-language-corpora/raw-text-corpus.texttext-generationn<1K0 likes659 downloads5mo agoHugging Face04LanguageShades /BiasShadesgatedInterested in contributing? Speak a language not represented here? Disagree with an annotation? Please submit feedback in the Community tab! Dataset Card for BiasShades Note: This dataset may NOT be used as training data in any form (pre-training, fine-tuning, post-training, etc.) without express permission from creators. Dataset Details Version: 1.0 License: SHADES 1 Montreal Data License Dataset Description 728 stereotypes and associated… See the full description on the dataset page: https://huggingface.co/datasets/LanguageShades/BiasShades.imagetext-classificationn<1K26 likes605 downloads3mo agoHugging Face05Sakonii /nepalitext-language-model-dataset Dataset Card for "nepalitext-language-model-dataset" Dataset Summary "NepaliText" language modeling dataset is a collection of over 13 million Nepali text sequences (phrases/sentences/paragraphs) extracted by combining the datasets: OSCAR , cc100 and a set of scraped Nepali articles on Wikipedia. Supported Tasks and Leaderboards This dataset is intended to pre-train language models and word representations on Nepali Language. Languages The data is… See the full description on the dataset page: https://huggingface.co/datasets/Sakonii/nepalitext-language-model-dataset.texttext-generation10M<n<100M8 likes479 downloads1y agoHugging Face06legesher /language-decoded-data Language Decoded | Multilingual Code Dataset Experiment and proposed paper title: Language Decoded: Exploring the Impact of Native Code on Multilingual Models Note (2026-05-18): Current Phase 3 configs use the short condition-* namespace and include 103k, 20k, and 5k sizes for Conditions 1--2. Phase 2 configs remain available under the phase-2-the-stack-v1-* namespace for reproducibility. Multilingual Python code datasets for the Language Decoded project (part of Cohere's… See the full description on the dataset page: https://huggingface.co/datasets/legesher/language-decoded-data.tabulartext-generation100K<n<1M1 likes383 downloads2mo agoHugging Face07polyglot-tagger /wikipedia-language-snippets-filtered Wikipedia Snippets (Filtered) Filtered sentence snippets in Wikipedia, by taking the first 60% of an article after filtering for stubs. Minor Latin groups are additionally filtered again for English leakage. Sentences are mostly filtered out for non matching scripts, such as Arabic in a Cyrllic language. Files Each file is in this format for languages in ISO 639 2-letter codes: train/en/en.parquet train/es/es.parquet From wikimedia/wikipedia Licensing… See the full description on the dataset page: https://huggingface.co/datasets/polyglot-tagger/wikipedia-language-snippets-filtered.texttext-generation10M<n<100M0 likes295 downloads5mo agoHugging Face08burak29 /Natural_Language_to_Ffmpeg_Commands Natural Language to FFmpeg Dataset Disclaimer: This dataset was synthetically generated using a large language model and is intended for research purposes only. The dataset may contain inaccuracies, errors, or inconsistencies. Users should exercise caution and verify the correctness of the data before using it in any application. This dataset contains 1000+ pairs of English natural language instructions and corresponding FFmpeg commands. The dataset is designed for tasks… See the full description on the dataset page: https://huggingface.co/datasets/burak29/Natural_Language_to_Ffmpeg_Commands.texttext-generation1K<n<10K1 likes160 downloads13d agoHugging Face09Lots-of-LoRAs /task1577_amazon_reviews_multi_japanese_language_classification Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1577_amazon_reviews_multi_japanese_language_classification Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1577_amazon_reviews_multi_japanese_language_classification.texttext-generationn<1K0 likes154 downloads2y agoHugging Face10nileagi /swahili-language-exposure swahili-language-exposure Dataset Summary swahili-language-exposure is a large-scale Swahili (Kiswahili) corpus designed for language exposure and continued pretraining of language models. Unlike instruction-tuning datasets, this dataset focuses on exposing models to natural Swahili usage across conversations, explanations, narratives, technical discussions, and mixed-domain text. The goal is to improve fluency, vocabulary coverage, syntax, and cultural grounding in… See the full description on the dataset page: https://huggingface.co/datasets/nileagi/swahili-language-exposure.texttext-generation1M<n<10M0 likes149 downloads8mo agoHugging Face11nileagi /swahili-language-exposure-v2 Swahili Language Exposure Large-scale Swahili corpus for continued pretraining and language exposure. Maintained by NileAGI. texttext-generation1M<n<10M3 likes133 downloads3mo agoHugging Face12Lots-of-LoRAs /task427_hindienglish_corpora_hi-en_language_identification Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task427_hindienglish_corpora_hi-en_language_identification Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task427_hindienglish_corpora_hi-en_language_identification.texttext-generation1K<n<10K0 likes132 downloads2y agoHugging Face13xkas2001 /uzbek-language-dataset Uzbek Language Dataset Collection Bu repository o'zbek tili uchun eng keng ko'lamli va keng qamrovli dataset to'plami hisoblanadi. Dataset turli manbalardan to'plangan va NLP modellari, til modellari va boshqa AI ilovalar uchun mo'ljallangan. 📊 Dataset Overview Bu dataset to'plami 4ta asosiy qism va qo'shimcha merge qilish asboblaridan iborat: 🎯 Dataset Qismlari Dataset Hajmi Maqsad Source community-oscar-uzbek 1.1GB OSCAR Community data Common… See the full description on the dataset page: https://huggingface.co/datasets/xkas2001/uzbek-language-dataset.texttext-generation2 likes129 downloads1y agoHugging Face14Arpuuu /nepalitext-language-model-dataset Dataset Card for "nepalitext-language-model-dataset" Dataset Summary "NepaliText" language modeling dataset is a collection of over 13 million Nepali text sequences (phrases/sentences/paragraphs) extracted by combining the datasets: OSCAR , cc100 and a set of scraped Nepali articles on Wikipedia. Supported Tasks and Leaderboards This dataset is intended to pre-train language models and word representations on Nepali Language. Languages The data is… See the full description on the dataset page: https://huggingface.co/datasets/Arpuuu/nepalitext-language-model-dataset.texttext-generation10M<n<100M0 likes113 downloads6mo agoHugging Face15Mwanzau /Tumbuka_language About this dataset This dataset mainly focuses on Tumbuka Language, found in Northern Malawi and Zambia. Usecases mainly focuses on datasets that are to be used for fine-tuning already existing AI Models, so that they are able to understand the Tumbuka Bantu Language (Malawi & Zambia & Tanzania). Formats The datasets are in different formats, and sometimes you will notice that the same dataset, have been uploaded with several file formats like .txt… See the full description on the dataset page: https://huggingface.co/datasets/Mwanzau/Tumbuka_language.texttext-generation100K<n<1M0 likes109 downloads2mo agoHugging Face16beatsprom /multimodal-vision-language-video-models-2026 👁️ Multimodal Vision-Language & Video Foundation Models Dataset (2026 Edition) A structured research dataset featuring 1,000 domain-verified research papers and code repositories focused on Multimodal Vision-Language Models (VLM), Video Foundation Models, Diffusion Transformers (DiT), Visual Grounding, and World Simulators. Built with Universal Scientific Engine V15.1 Gold, providing 47 schema attributes with verified repository attribution, modality capability matrix, vision… See the full description on the dataset page: https://huggingface.co/datasets/beatsprom/multimodal-vision-language-video-models-2026.tabularfeature-extractionn<1K3 likes107 downloads1mo agoHugging Face17Lelonthecodeur /every-language-dataset-v3 Every Language Dataset V3 Next-generation synthetic multilingual dataset. Size Total: 25,000,000 Train: 24,000,000 Validation: 500,000 Test: 500,000 Diversity Human-language catalog: 171 language codes. Programming languages: 50. Task families: conversation question answering reasoning logic arithmetic translation summarization explanation code generation code explanation debugging Format Parquet + ZSTD. Generation… See the full description on the dataset page: https://huggingface.co/datasets/Lelonthecodeur/every-language-dataset-v3.texttext-generation10M<n<100M2 likes91 downloads6d agoHugging Face18burak29 /git-natural-language-commands Git Natural Language Commands A dataset mapping English natural-language instructions to their corresponding git commands, intended for training and evaluating models that translate user intent into safe, correct shell commands. Disclaimer: This dataset was generated using Large Language Models (LLMs). The examples have not been manually verified against real-world usage and may contain errors, inconsistencies, or non-canonical phrasings. Use with appropriate caution.… See the full description on the dataset page: https://huggingface.co/datasets/burak29/git-natural-language-commands.texttext-generation1K<n<10K1 likes80 downloads11d agoHugging Face19Mxode /CSDN-C_Language-2013_2023CSDN - C 语言社区 2013 ~ 2023.10.2 的问答数据,未包含图片,仅有文本内容。 共 29K+ 条,数据已经经过初步清洗和脱敏,去除了所有 0 回复的贴子 & 机器人回复的贴子。为了方便不同使用目的,按照回复盖楼的格式对数据进行了组织,一个样例(展开后)如下: { "question": "刚学C语言,为什么这个代码运行不了呢", "poster": "user-0", "comments": [ { "cid": "2", "user": "user-2", "content": "intunsigned intlong longunsigned long long统统容纳不下29的阶乘,早就溢出了。", "referer": "user-0" }, { "cid": "3", "user": "user-3"… See the full description on the dataset page: https://huggingface.co/datasets/Mxode/CSDN-C_Language-2013_2023.texttext-generation10K<n<100K4 likes74 downloads1y agoHugging Face20itsmebatuhan /bluesky-10m-posts-15-languages Dataset Card: Bluesky 10M Multilingual 📊 Overview Total Posts: 10,099,990 Languages: 15 (en, tr, es, pt, de, fr, ja, it, nl, pl, ru, ko, zh, ar, hi) Collection Period: August 9-12, 2026 Source: Bluesky Jetstream API (public firehose) Format: JSONL Size: ~3 GB 🌍 Language Distribution Language Code Posts % English en 6,843,995 67.8% Japanese ja 1,547,179 15.3% German de 373,626 3.7% Portuguese pt 331,093 3.3% Spanish es 325,865… See the full description on the dataset page: https://huggingface.co/datasets/itsmebatuhan/bluesky-10m-posts-15-languages.texttext-classification10M<n<100M0 likes65 downloads1mo agoHugging Face21agentlans /HuggingFaceFW-finetranslations-100-languages-sample Finetranslations 100 Language Sample Dataset Subset of HuggingFaceFW/finetranslations with the top 100 languages by number of documents. Configurations all: 100 languages combined (100k rows), shuffled 100 individual language configs: 1000 rows each Columns Original columns + language (source language indicator which is the name of the config) Usage from datasets import load_dataset dataset =… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/HuggingFaceFW-finetranslations-100-languages-sample.tabulartranslation100K<n<1M0 likes64 downloads7mo agoHugging Face22math-across-languages /gsm8k-translated Multilingual GSM8K Translations This dataset contains machine-translated versions of GSM8K in these languages: French (fr) German (de) Hindi (hi) Dataset Structure For each language, we provide the original GSM8K train and test splits: train: 7,473 samples test: 1,319 samples Each sample consists of a question and an answer. The question describes a grade-school-level math word problem that requires multi-step mathematical reasoning. The answer contains a… See the full description on the dataset page: https://huggingface.co/datasets/math-across-languages/gsm8k-translated.textquestion-answering10K<n<100K0 likes59 downloads3mo agoHugging Face23legesher /language-decoded-community Language Decoded — Community Code Natively-authored multilingual code for the Language Decoded project (part of Cohere's Tiny Aya Expedition). This dataset contains code written by developers in non-English programming languages and code with significant CJK content — not mechanically transpiled or LLM-translated from English. Experiment and proposed paper title: Language Decoded: Exploring the Impact of Native Code on Multilingual Models This data serves as the corpus for… See the full description on the dataset page: https://huggingface.co/datasets/legesher/language-decoded-community.tabulartext-generation1K<n<10K0 likes56 downloads2mo agoHugging Face24AzizBelaweid /Tunisian_Language_Dataset Dataset Card for Tunisian Text Compilation This dataset is a curated compilation of various Tunisian datasets, aimed at gathering as much Tunisian text data as possible in one place. It combines multiple sources of Tunisian language data, providing a rich resource for research, development of NLP models, and linguistic studies on Tunisian text. Dataset Details Dataset Description This dataset aggregates several publicly available datasets that contain Tunisian… See the full description on the dataset page: https://huggingface.co/datasets/AzizBelaweid/Tunisian_Language_Dataset.texttext-generation100K<n<1M7 likes55 downloads2y agoHugging Face25Mxode /Chinese-StackOverflow-QA-C_Language 中文 StackOverflow C 语言问答数据集 💻 Github Repo 基本信息 本数据集提供了两个子集: translated:原数据集 Mxode/StackOverflow-QA-C-Language-40k 的中文翻译版本,数量约 40K。 synthetic **(Default)**:在原数据集 Mxode/StackOverflow-QA-C-Language-40k 的基础上,重新扩充、合成的问答数据集,数量约 200K。 数据格式 请注意:两个子集的数据格式并不完全相同。 translated 子集: { "id": << 12位nanoid >>, "question_en": << 用户提问(英文) >>, "question_zh": << 用户提问(中文) >>, "answer_en": << 用户回答(英文) >>, "answer_zh": << 用户回答(中文) >>, } synthetic 子集: { "id": <<… See the full description on the dataset page: https://huggingface.co/datasets/Mxode/Chinese-StackOverflow-QA-C_Language.texttext-generation100K<n<1M1 likes47 downloads1y agoHugging Face26ScoutieAutoML /scoutieDataset_russian_language_grammar_and_rules_vectorized Description in English: A dataset collected from 30 Russian-language Telegram channels on the topic of learning the Russian language. This dataset contains grammar, syntax, spelling and punctuation rules. The dataset was collected and marked automatically using the Scoutie data collection and marking service.Try Scoutie and collect the same or another dataset using the link. Dataset fields: taskId - task identifier in the Scouti service. text - main text. url -… See the full description on the dataset page: https://huggingface.co/datasets/ScoutieAutoML/scoutieDataset_russian_language_grammar_and_rules_vectorized.tabulartext-classification10K<n<100K2 likes46 downloads2y agoHugging Face27spectralbranding /meaningfulness-cross-language-rendering Cross-Language Rendering for Meaning vs Meaningfulness (Paper B 2026ap) HF dataset DOI: 10.57967/hf/8971 Companion paper concept DOI: 10.5281/zenodo.20409701 Companion GitHub mirror: https://github.com/spectralbranding/meaningfulness-papers/tree/main/meaning-meaningfulness-empirical Dataset Summary This dataset contains the multi-language rendering and extraction artifacts demonstrating Proposition P4 (rendering-equivalence under spine-preservation) from Zharnikov… See the full description on the dataset page: https://huggingface.co/datasets/spectralbranding/meaningfulness-cross-language-rendering.tabulartext-generationn<1K0 likes43 downloads2mo agoHugging Face28MichiganNLP /language-energy-divide 🌍⚡ The Language–Energy Divide Per-language energy measurements & prompts for multilingual LLM inference 📢 News Aug 2026 — Our paper has been accepted to EMNLP 2026 (Main Conference)! 🎉 This dataset accompanies the paper "The Language–Energy Divide: Measuring Energy Costs of Multilingual LLM Inference." It releases the per-language energy measurements and the prompts used in the study, so researchers can build on our numbers without… See the full description on the dataset page: https://huggingface.co/datasets/MichiganNLP/language-energy-divide.tabularquestion-answeringn<1K0 likes42 downloads1mo agoHugging Face29anrilombard /sa-languages South African Languages Dataset Dataset Overview Language Training Documents Training GPT2 Tokens Avg Tokens/Doc Max Tokens Test Documents Test GPT2 Tokens Test Avg Tokens/Doc Test Max Tokens isiZulu 116,693 192,622,799 1,650.68 335,530 687 1,080,961 1,573.45 15,691 Sesotho 83,329 144,337,938 1,732.15 98,542 841 1,393,086 1,656.4614,071 isiXhosa 99,567 141,484,241 1,421.00 113,710 788 1,161,296 1,473.73 17,220 isiNdebele 21,922 17,533,799 799.83 42,701… See the full description on the dataset page: https://huggingface.co/datasets/anrilombard/sa-languages.texttext-generation100K<n<1M1 likes36 downloads2y agoHugging Face30yasalma /tt-en-language-corpusThe Tatar-English Parallel Corpus is a collection of parallel sentences in Tatar and English. It is designed to facilitate research and development in machine translation, natural language processing (NLP), and multilingual studies. This dataset contains aligned Tatar-English sentence pairs that can be used for training and evaluating machine translation models, cross-lingual tasks, and linguistic analysis of the Tatar language. Languages Tatar: ISO 639-1 code tt English: ISO 639-1… See the full description on the dataset page: https://huggingface.co/datasets/yasalma/tt-en-language-corpus.texttext-generation1K<n<10K1 likes35 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.