CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01akoksal /muri-it-language-split MURI-IT: Multilingual Instruction Tuning Dataset for 200 Languages via Multilingual Reverse Instructions MURI-IT is a large-scale multilingual instruction tuning dataset containing 2.2 million instruction-output pairs across 200 languages. It is designed to address the challenges of instruction tuning in low-resource languages with Multilingual Reverse Instructions (MURI), which ensures that the output is human-written, high-quality, and authentic to the cultural and linguistic… See the full description on the dataset page: https://huggingface.co/datasets/akoksal/muri-it-language-split.texttext-generation1M<n<10M6 likes11k downloads2y agoHugging Face02AIStudioDelta /Eurovoc_2025_by_language 🇪🇺 🏷️ EuroVoc dataset (by language) This is the EuropeanParliament/Eurovoc_2025 dataset, but split up by language, not by period. The original is split up into periods (1996-03 through 2025-11), with documents in different languages mixed together. For ease of training this dataset splits the data by language instead, with documents in different periods put together. License This dataset is redistributed under the original European Union Public License 1.2. When… See the full description on the dataset page: https://huggingface.co/datasets/AIStudioDelta/Eurovoc_2025_by_language.texttext-generation1M<n<10M1 likes688 downloads10mo agoHugging Face03zomi-language-corpora /raw-text-corpus 📝 Zomi Raw Text Corpus (Community-Contributed) The Zomi Raw Text Corpus is an open, community-driven collection of unprocessed Zomi-language text.It serves as the foundational dataset for building the full Zomi NLP ecosystem, including tokenizers, language models, ASR/TTS systems, and downstream tasks. This dataset is intentionally raw — no normalization, deduplication, or cleaning is applied.Cleaned and task-specific datasets will be released separately. 📦 Dataset… See the full description on the dataset page: https://huggingface.co/datasets/zomi-language-corpora/raw-text-corpus.texttext-generationn<1K0 likes647 downloads4mo agoHugging Face04LanguageShades /BiasShadesgatedInterested in contributing? Speak a language not represented here? Disagree with an annotation? Please submit feedback in the Community tab! Dataset Card for BiasShades Note: This dataset may NOT be used as training data in any form (pre-training, fine-tuning, post-training, etc.) without express permission from creators. Dataset Details Version: 1.0 License: SHADES 1 Montreal Data License Dataset Description 728 stereotypes and associated… See the full description on the dataset page: https://huggingface.co/datasets/LanguageShades/BiasShades.imagetext-classificationn<1K26 likes577 downloads3mo agoHugging Face05Sakonii /nepalitext-language-model-dataset Dataset Card for "nepalitext-language-model-dataset" Dataset Summary "NepaliText" language modeling dataset is a collection of over 13 million Nepali text sequences (phrases/sentences/paragraphs) extracted by combining the datasets: OSCAR , cc100 and a set of scraped Nepali articles on Wikipedia. Supported Tasks and Leaderboards This dataset is intended to pre-train language models and word representations on Nepali Language. Languages The data is… See the full description on the dataset page: https://huggingface.co/datasets/Sakonii/nepalitext-language-model-dataset.texttext-generation10M<n<100M8 likes479 downloads1y agoHugging Face06injilashah /Kashmiri-language-text-datasetThis is a Dataset for Kashmiri language containing Kashmiri words their phonemes , along with their meaning in ENGlISH and HINDI and an example sentence in english . The phoenmes are written phonemes present in wordphonemes-meaning.csv translation0 likes401 downloads2y agoHugging Face07legesher /language-decoded-data Language Decoded | Multilingual Code Dataset Experiment and proposed paper title: Language Decoded: Exploring the Impact of Native Code on Multilingual Models Note (2026-05-18): Current Phase 3 configs use the short condition-* namespace and include 103k, 20k, and 5k sizes for Conditions 1--2. Phase 2 configs remain available under the phase-2-the-stack-v1-* namespace for reproducibility. Multilingual Python code datasets for the Language Decoded project (part of Cohere's… See the full description on the dataset page: https://huggingface.co/datasets/legesher/language-decoded-data.tabulartext-generation100K<n<1M1 likes380 downloads2mo agoHugging Face08polyglot-tagger /wikipedia-language-snippets-filtered Wikipedia Snippets (Filtered) Filtered sentence snippets in Wikipedia, by taking the first 60% of an article after filtering for stubs. Minor Latin groups are additionally filtered again for English leakage. Sentences are mostly filtered out for non matching scripts, such as Arabic in a Cyrllic language. Files Each file is in this format for languages in ISO 639 2-letter codes: train/en/en.parquet train/es/es.parquet From wikimedia/wikipedia Licensing… See the full description on the dataset page: https://huggingface.co/datasets/polyglot-tagger/wikipedia-language-snippets-filtered.texttext-generation10M<n<100M0 likes291 downloads5mo agoHugging Face09mongodb-eai /natural-language-to-mongosh Natural Language to MongoDB Shell (mongosh) Benchmark Benchmark dataset for performing natural language (NL) to MongoDB Shell (mongosh) code generation. There is an emerging desire from users for NL query generation. This benchmarks examines how LLMs generate MongoDB queries and provides proactive guidance for making systems that map NL to MongoDB queries. Repository Contents This repository contains: Benchmark dataset (flat CSV file, Braintrust evaluation… See the full description on the dataset page: https://huggingface.co/datasets/mongodb-eai/natural-language-to-mongosh.text-generationn<1K4 likes214 downloads1y agoHugging Face10languagehub-ai /yuxiaowang-prompts-2025 Yuxiaowang Semantic Dataset · Hugging Face Version 🧠 English Summary Yuxiaowang · Semantic Dataset for Japanese Language Schools (Chinese) This project provides structured semantic definitions and prompt examples for the domain of Japanese language schools in China.It aims to serve as a grounding corpus for large language models (LLMs) to understand terms like "语校", "语校网", and related concepts. Source platform: https://www.yuxiaowang.comAll prompts and term… See the full description on the dataset page: https://huggingface.co/datasets/languagehub-ai/yuxiaowang-prompts-2025.text-generationn<1K0 likes206 downloads8mo agoHugging Face11Lots-of-LoRAs /task1577_amazon_reviews_multi_japanese_language_classification Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1577_amazon_reviews_multi_japanese_language_classification Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1577_amazon_reviews_multi_japanese_language_classification.texttext-generationn<1K0 likes163 downloads2y agoHugging Face12burak29 /Natural_Language_to_Ffmpeg_Commands Natural Language to FFmpeg Dataset Disclaimer: This dataset was synthetically generated using a large language model and is intended for research purposes only. The dataset may contain inaccuracies, errors, or inconsistencies. Users should exercise caution and verify the correctness of the data before using it in any application. This dataset contains 1000+ pairs of English natural language instructions and corresponding FFmpeg commands. The dataset is designed for tasks… See the full description on the dataset page: https://huggingface.co/datasets/burak29/Natural_Language_to_Ffmpeg_Commands.texttext-generation1K<n<10K1 likes158 downloads11d agoHugging Face13nileagi /swahili-language-exposure swahili-language-exposure Dataset Summary swahili-language-exposure is a large-scale Swahili (Kiswahili) corpus designed for language exposure and continued pretraining of language models. Unlike instruction-tuning datasets, this dataset focuses on exposing models to natural Swahili usage across conversations, explanations, narratives, technical discussions, and mixed-domain text. The goal is to improve fluency, vocabulary coverage, syntax, and cultural grounding in… See the full description on the dataset page: https://huggingface.co/datasets/nileagi/swahili-language-exposure.texttext-generation1M<n<10M0 likes146 downloads8mo agoHugging Face14Lots-of-LoRAs /task427_hindienglish_corpora_hi-en_language_identification Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task427_hindienglish_corpora_hi-en_language_identification Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task427_hindienglish_corpora_hi-en_language_identification.texttext-generation1K<n<10K0 likes138 downloads2y agoHugging Face15xkas2001 /uzbek-language-dataset Uzbek Language Dataset Collection Bu repository o'zbek tili uchun eng keng ko'lamli va keng qamrovli dataset to'plami hisoblanadi. Dataset turli manbalardan to'plangan va NLP modellari, til modellari va boshqa AI ilovalar uchun mo'ljallangan. 📊 Dataset Overview Bu dataset to'plami 4ta asosiy qism va qo'shimcha merge qilish asboblaridan iborat: 🎯 Dataset Qismlari Dataset Hajmi Maqsad Source community-oscar-uzbek 1.1GB OSCAR Community data Common… See the full description on the dataset page: https://huggingface.co/datasets/xkas2001/uzbek-language-dataset.texttext-generation2 likes134 downloads1y agoHugging Face16agentlans /HuggingFaceFW-finetranslations-100-languages-sample Finetranslations 100 Language Sample Dataset Subset of HuggingFaceFW/finetranslations with the top 100 languages by number of documents. Configurations all: 100 languages combined (100k rows), shuffled 100 individual language configs: 1000 rows each Columns Original columns + language (source language indicator which is the name of the config) Usage from datasets import load_dataset dataset =… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/HuggingFaceFW-finetranslations-100-languages-sample.tabulartranslation100K<n<1M0 likes121 downloads7mo agoHugging Face17Arpuuu /nepalitext-language-model-dataset Dataset Card for "nepalitext-language-model-dataset" Dataset Summary "NepaliText" language modeling dataset is a collection of over 13 million Nepali text sequences (phrases/sentences/paragraphs) extracted by combining the datasets: OSCAR , cc100 and a set of scraped Nepali articles on Wikipedia. Supported Tasks and Leaderboards This dataset is intended to pre-train language models and word representations on Nepali Language. Languages The data is… See the full description on the dataset page: https://huggingface.co/datasets/Arpuuu/nepalitext-language-model-dataset.texttext-generation10M<n<100M0 likes115 downloads6mo agoHugging Face18nileagi /swahili-language-exposure-v2 Swahili Language Exposure Large-scale Swahili corpus for continued pretraining and language exposure. Maintained by NileAGI. texttext-generation1M<n<10M3 likes109 downloads3mo agoHugging Face19beatsprom /multimodal-vision-language-video-models-2026 👁️ Multimodal Vision-Language & Video Foundation Models Dataset (2026 Edition) A structured research dataset featuring 1,000 domain-verified research papers and code repositories focused on Multimodal Vision-Language Models (VLM), Video Foundation Models, Diffusion Transformers (DiT), Visual Grounding, and World Simulators. Built with Universal Scientific Engine V15.1 Gold, providing 47 schema attributes with verified repository attribution, modality capability matrix, vision… See the full description on the dataset page: https://huggingface.co/datasets/beatsprom/multimodal-vision-language-video-models-2026.tabularfeature-extractionn<1K3 likes109 downloads1mo agoHugging Face20Mwanzau /Tumbuka_language About this dataset This dataset mainly focuses on Tumbuka Language, found in Northern Malawi and Zambia. Usecases mainly focuses on datasets that are to be used for fine-tuning already existing AI Models, so that they are able to understand the Tumbuka Bantu Language (Malawi & Zambia & Tanzania). Formats The datasets are in different formats, and sometimes you will notice that the same dataset, have been uploaded with several file formats like .txt… See the full description on the dataset page: https://huggingface.co/datasets/Mwanzau/Tumbuka_language.texttext-generation100K<n<1M0 likes108 downloads2mo agoHugging Face21Mehgoss /sa-languages-corpus South African Languages Text Corpus Plain-text corpus covering all 11 official South African languages, for language modeling. text-generation0 likes89 downloads9d agoHugging Face22Lelonthecodeur /every-language-dataset-v3 Every Language Dataset V3 Next-generation synthetic multilingual dataset. Size Total: 25,000,000 Train: 24,000,000 Validation: 500,000 Test: 500,000 Diversity Human-language catalog: 171 language codes. Programming languages: 50. Task families: conversation question answering reasoning logic arithmetic translation summarization explanation code generation code explanation debugging Format Parquet + ZSTD. Generation… See the full description on the dataset page: https://huggingface.co/datasets/Lelonthecodeur/every-language-dataset-v3.texttext-generation10M<n<100M2 likes86 downloads5d agoHugging Face23itsmebatuhan /bluesky-10m-posts-15-languages Dataset Card: Bluesky 10M Multilingual 📊 Overview Total Posts: 10,099,990 Languages: 15 (en, tr, es, pt, de, fr, ja, it, nl, pl, ru, ko, zh, ar, hi) Collection Period: August 9-12, 2026 Source: Bluesky Jetstream API (public firehose) Format: JSONL Size: ~3 GB 🌍 Language Distribution Language Code Posts % English en 6,843,995 67.8% Japanese ja 1,547,179 15.3% German de 373,626 3.7% Portuguese pt 331,093 3.3% Spanish es 325,865… See the full description on the dataset page: https://huggingface.co/datasets/itsmebatuhan/bluesky-10m-posts-15-languages.texttext-classification10M<n<100M0 likes80 downloads1mo agoHugging Face24burak29 /git-natural-language-commands Git Natural Language Commands A dataset mapping English natural-language instructions to their corresponding git commands, intended for training and evaluating models that translate user intent into safe, correct shell commands. Disclaimer: This dataset was generated using Large Language Models (LLMs). The examples have not been manually verified against real-world usage and may contain errors, inconsistencies, or non-canonical phrasings. Use with appropriate caution.… See the full description on the dataset page: https://huggingface.co/datasets/burak29/git-natural-language-commands.texttext-generation1K<n<10K1 likes77 downloads9d agoHugging Face25Mxode /CSDN-C_Language-2013_2023CSDN - C 语言社区 2013 ~ 2023.10.2 的问答数据,未包含图片,仅有文本内容。 共 29K+ 条,数据已经经过初步清洗和脱敏,去除了所有 0 回复的贴子 & 机器人回复的贴子。为了方便不同使用目的,按照回复盖楼的格式对数据进行了组织,一个样例(展开后)如下: { "question": "刚学C语言,为什么这个代码运行不了呢", "poster": "user-0", "comments": [ { "cid": "2", "user": "user-2", "content": "intunsigned intlong longunsigned long long统统容纳不下29的阶乘,早就溢出了。", "referer": "user-0" }, { "cid": "3", "user": "user-3"… See the full description on the dataset page: https://huggingface.co/datasets/Mxode/CSDN-C_Language-2013_2023.texttext-generation10K<n<100K4 likes73 downloads1y agoHugging Face26AzizBelaweid /Tunisian_Language_Dataset Dataset Card for Tunisian Text Compilation This dataset is a curated compilation of various Tunisian datasets, aimed at gathering as much Tunisian text data as possible in one place. It combines multiple sources of Tunisian language data, providing a rich resource for research, development of NLP models, and linguistic studies on Tunisian text. Dataset Details Dataset Description This dataset aggregates several publicly available datasets that contain Tunisian… See the full description on the dataset page: https://huggingface.co/datasets/AzizBelaweid/Tunisian_Language_Dataset.texttext-generation100K<n<1M7 likes59 downloads2y agoHugging Face27legesher /language-decoded-community Language Decoded — Community Code Natively-authored multilingual code for the Language Decoded project (part of Cohere's Tiny Aya Expedition). This dataset contains code written by developers in non-English programming languages and code with significant CJK content — not mechanically transpiled or LLM-translated from English. Experiment and proposed paper title: Language Decoded: Exploring the Impact of Native Code on Multilingual Models This data serves as the corpus for… See the full description on the dataset page: https://huggingface.co/datasets/legesher/language-decoded-community.tabulartext-generation1K<n<10K0 likes58 downloads2mo agoHugging Face28math-across-languages /gsm8k-translated Multilingual GSM8K Translations This dataset contains machine-translated versions of GSM8K in these languages: French (fr) German (de) Hindi (hi) Dataset Structure For each language, we provide the original GSM8K train and test splits: train: 7,473 samples test: 1,319 samples Each sample consists of a question and an answer. The question describes a grade-school-level math word problem that requires multi-step mathematical reasoning. The answer contains a… See the full description on the dataset page: https://huggingface.co/datasets/math-across-languages/gsm8k-translated.textquestion-answering10K<n<100K0 likes56 downloads3mo agoHugging Face29Mxode /Chinese-StackOverflow-QA-C_Language 中文 StackOverflow C 语言问答数据集 💻 Github Repo 基本信息 本数据集提供了两个子集: translated:原数据集 Mxode/StackOverflow-QA-C-Language-40k 的中文翻译版本,数量约 40K。 synthetic **(Default)**:在原数据集 Mxode/StackOverflow-QA-C-Language-40k 的基础上,重新扩充、合成的问答数据集,数量约 200K。 数据格式 请注意:两个子集的数据格式并不完全相同。 translated 子集: { "id": << 12位nanoid >>, "question_en": << 用户提问(英文) >>, "question_zh": << 用户提问(中文) >>, "answer_en": << 用户回答(英文) >>, "answer_zh": << 用户回答(中文) >>, } synthetic 子集: { "id": <<… See the full description on the dataset page: https://huggingface.co/datasets/Mxode/Chinese-StackOverflow-QA-C_Language.texttext-generation100K<n<1M1 likes50 downloads1y agoHugging Face30zhykoties /time-series-language-alignment TS-Insights Dataset Dataset Description TS-Insights is the official dataset for the paper "Insight Miner: A Time Series Analysis Dataset for Cross-Domain Alignment with Natural Language". This work is done by Project Mineral from Google X in 2023. It is the first large-scale general-domain dataset designed to align time-series data with natural language descriptions. The dataset supports the training of Large Multimodal Models (LMMs) to understand time series as a new… See the full description on the dataset page: https://huggingface.co/datasets/zhykoties/time-series-language-alignment.visual-question-answering100K<n<1M3 likes49 downloads9mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.