CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01papluca /language-identification Dataset Card for Language Identification dataset Dataset Summary The Language Identification dataset is a collection of 90k samples consisting of text passages and corresponding language label. This dataset was created by collecting data from 3 sources: Multilingual Amazon Reviews Corpus, XNLI, and STSb Multi MT. Supported Tasks and Leaderboards The dataset can be used to train a model for language identification, which is a multi-class text classification… See the full description on the dataset page: https://huggingface.co/datasets/papluca/language-identification.texttext-classification10K<n<100K70 likes3.3k downloads4y agoHugging Face02candradhipa /Language-Detectiontext10K<n<100K0 likes1.3k downloads2y agoHugging Face03sakthivinash /Language_Detection Language_Detection - Multilingual Text Classification Dataset This dataset is a collection of multilingual text samples designed for training and predicting languages in Artificial Intelligence (AI), Machine Learning (ML), Deep Learning (DL), and Data Science (DS) applications. It contains labeled data that associates text samples with their respective languages, enabling language detection and classification tasks. Dataset Overview The dataset consists of two columns:… See the full description on the dataset page: https://huggingface.co/datasets/sakthivinash/Language_Detection.text10K<n<100K0 likes1.2k downloads2y agoHugging Face04simoneteglia /europarl_for_language_detection_10ktext100K<n<1M0 likes649 downloads3y agoHugging Face05LanguageShades /BiasShadesgatedInterested in contributing? Speak a language not represented here? Disagree with an annotation? Please submit feedback in the Community tab! Dataset Card for BiasShades Note: This dataset may NOT be used as training data in any form (pre-training, fine-tuning, post-training, etc.) without express permission from creators. Dataset Details Version: 1.0 License: SHADES 1 Montreal Data License Dataset Description 728 stereotypes and associated… See the full description on the dataset page: https://huggingface.co/datasets/LanguageShades/BiasShades.imagetext-classificationn<1K26 likes554 downloads3mo agoHugging Face06legesher /language-decoded-experiments Language Decoded — Experiment Tracking Central hub for training logs, configurations, evaluation results, and analysis for the Language Decoded project. The project originated as a proposal during Cohere's Tiny Aya Expedition (March 2026 hackathon) and was extended into Phase 3 for the accompanying paper. Submitted paper title (2026-05-26): Language, Decoded: Exploring the Impact of Fine-Tuning a Multilingual Model on Native-Language Code ⚠️ Phase 3 numbers — read… See the full description on the dataset page: https://huggingface.co/datasets/legesher/language-decoded-experiments.tabular10K<n<100K2 likes334 downloads2mo agoHugging Face07FrancophonIA /language_detection [!NOTE] Dataset origin: https://www.kaggle.com/datasets/basilb2s/language-detection It's a small language detection dataset. This dataset consists of text details for 17 different languages, ie, you will be able to create an NLP model for predicting 17 different language.. text10K<n<100K0 likes313 downloads1y agoHugging Face08MoazIrfan /language-detectiontext10K<n<100K1 likes303 downloads9mo agoHugging Face09nawabhussain /Kashmiri-Language-Corpus Kashmiri Textual Data Corpus Introduction This repository contains a combined dataset of Kashmiri textual data collected from various sources. The data has been sourced from different locations and may contain non-Kashmiri text (e.g., Urdu, Persian). The goal of this corpus is to provide a wide variety of Kashmiri text data for research and language processing tasks. Sources of Data 1. mzmmoazam/kashmiri_dataset (HTML Data) Source: GitHub -… See the full description on the dataset page: https://huggingface.co/datasets/nawabhussain/Kashmiri-Language-Corpus.text10K<n<100K2 likes187 downloads2y agoHugging Face10huggingface /language_codes_marianMTtextn<1K0 likes179 downloads2y agoHugging Face11Toygar /turkish-offensive-language-detection Dataset Summary This dataset is enhanced version of existing offensive language studies. Existing studies are highly imbalanced, and solving this problem is too costly. To solve this, we proposed contextual data mining method for dataset augmentation. Our method is basically prevent us from retrieving random tweets and label individually. We can directly access almost exact hate related tweets and label them directly without any further human interaction in order to solve imbalanced… See the full description on the dataset page: https://huggingface.co/datasets/Toygar/turkish-offensive-language-detection.tabulartext-classification10K<n<100K20 likes172 downloads3y agoHugging Face12burak29 /Natural_Language_to_Ffmpeg_Commands Natural Language to FFmpeg Dataset Disclaimer: This dataset was synthetically generated using a large language model and is intended for research purposes only. The dataset may contain inaccuracies, errors, or inconsistencies. Users should exercise caution and verify the correctness of the data before using it in any application. This dataset contains 1000+ pairs of English natural language instructions and corresponding FFmpeg commands. The dataset is designed for tasks… See the full description on the dataset page: https://huggingface.co/datasets/burak29/Natural_Language_to_Ffmpeg_Commands.texttext-generation1K<n<10K1 likes157 downloads11d agoHugging Face13unklefedor /language-identificationtext100K<n<1M1 likes155 downloads3y agoHugging Face14Overfit-GM /turkish-toxic-language Turkish Texts for Toxic Language Detection Dataset Description Dataset Summary This text dataset is a collection of Turkish texts that have been merged from various existing offensive language datasets found online. The dataset contains a total of 77,800 instances, each labeled as either offensive or not offensive. To ensure the dataset's completeness, we utilized multiple transformer models to augment the dataset with pseudo labels. The resulting dataset is… See the full description on the dataset page: https://huggingface.co/datasets/Overfit-GM/turkish-toxic-language.texttext-classification10K<n<100K34 likes116 downloads3y agoHugging Face15dirtycomputer /Automated_Hate_Speech_Detection_and_the_Problem_of_Offensive_Languagetabular10K<n<100K0 likes102 downloads3y agoHugging Face16Okwu /african-language-parallel-corpus African Language Parallel Corpus Human-created, human-validated parallel sentence pairs for three African languages, released openly by Okwu. Version 1.0. Dataset summary A parallel corpus of everyday-register sentence pairs for Yorùbá, Swahili, and Nigerian Pidgin, each paired with English. The core is derived from NKENNE's own language-learning curriculum — content authored and reviewed by native-speaker educators — supplemented for Swahili with public-domain… See the full description on the dataset page: https://huggingface.co/datasets/Okwu/african-language-parallel-corpus.texttranslation10K<n<100K0 likes100 downloads14h agoHugging Face17burak29 /git-natural-language-commands Git Natural Language Commands A dataset mapping English natural-language instructions to their corresponding git commands, intended for training and evaluating models that translate user intent into safe, correct shell commands. Disclaimer: This dataset was generated using Large Language Models (LLMs). The examples have not been manually verified against real-world usage and may contain errors, inconsistencies, or non-canonical phrasings. Use with appropriate caution.… See the full description on the dataset page: https://huggingface.co/datasets/burak29/git-natural-language-commands.texttext-generation1K<n<10K1 likes76 downloads9d agoHugging Face18Charif-Ayfarah /Afar-language-text-to-speech-TTS Usage This dataset is designed to support the development of Text-to-Speech (TTS) systems for the Afar language. It can be integrated into web applications, mobile apps, desktop software, or other platforms that require natural-sounding Afar voice synthesis or accurate spoken language recognition. For applications involving virtual avatars or voice personas, the following culturally appropriate voice names are recommended: Female Voices: Emeli, Hanaawi, Kareera, Laysani, Kulsuma… See the full description on the dataset page: https://huggingface.co/datasets/Charif-Ayfarah/Afar-language-text-to-speech-TTS.text-to-speech1 likes56 downloads10mo agoHugging Face19AzizBelaweid /Tunisian_Language_Dataset Dataset Card for Tunisian Text Compilation This dataset is a curated compilation of various Tunisian datasets, aimed at gathering as much Tunisian text data as possible in one place. It combines multiple sources of Tunisian language data, providing a rich resource for research, development of NLP models, and linguistic studies on Tunisian text. Dataset Details Dataset Description This dataset aggregates several publicly available datasets that contain Tunisian… See the full description on the dataset page: https://huggingface.co/datasets/AzizBelaweid/Tunisian_Language_Dataset.texttext-generation100K<n<1M7 likes54 downloads2y agoHugging Face20pranavagrawal /Language-Identificationtext1M<n<10M0 likes53 downloads2y agoHugging Face21math-across-languages /gsm8k-translated Multilingual GSM8K Translations This dataset contains machine-translated versions of GSM8K in these languages: French (fr) German (de) Hindi (hi) Dataset Structure For each language, we provide the original GSM8K train and test splits: train: 7,473 samples test: 1,319 samples Each sample consists of a question and an answer. The question describes a grade-school-level math word problem that requires multi-step mathematical reasoning. The answer contains a… See the full description on the dataset page: https://huggingface.co/datasets/math-across-languages/gsm8k-translated.textquestion-answering10K<n<100K0 likes53 downloads3mo agoHugging Face22Charif-Ayfarah /English-to-Afar-language-translation Author Created by Charif Ayfarah. Contact: afbarit@gmail.com License Licensed under CC BY 4.0. You are free to use, modify, and distribute this dataset, including for commercial purposes, as long as you give appropriate credit. textn<1K2 likes52 downloads6mo agoHugging Face23MichiganNLP /language-energy-divide 🌍⚡ The Language–Energy Divide Per-language energy measurements & prompts for multilingual LLM inference 📢 News Aug 2026 — Our paper has been accepted to EMNLP 2026 (Main Conference)! 🎉 This dataset accompanies the paper "The Language–Energy Divide: Measuring Energy Costs of Multilingual LLM Inference." It releases the per-language energy measurements and the prompts used in the study, so researchers can build on our numbers without… See the full description on the dataset page: https://huggingface.co/datasets/MichiganNLP/language-energy-divide.tabularquestion-answeringn<1K0 likes52 downloads1mo agoHugging Face24pythainlp /thai-local-language-translation-dataset Thai Local Language Translation Dataset Thai Local Language Translation Dataset is a translation dataset for translate Thai Local Language to Thai Central Language. We create the dataset from Thai Dialect Corpus (Thai dialects ASR corpus). We select train set only from Thai Dialect Corpus. The dataset support Khummuang, Korat, and Pattani. Reference Suwanbandit, A., Naowarat, B., Sangpetch, O., Chuangsuwanich, E. (2023) Thai Dialect Corpus and Transfer-based Curriculum… See the full description on the dataset page: https://huggingface.co/datasets/pythainlp/thai-local-language-translation-dataset.texttranslation10K<n<100K4 likes51 downloads2y agoHugging Face25lewoniewski /grokipedia-wikipedia-16-languages Dataset description This dataset contains a mapping between Grokipedia v0.1 article pages and the corresponding Wikipedia article titles across 16 language editions (based on Wikipedia and Wikidata dumps from 1 November 2025). Each record includes: The URL of the Grokipedia page: grokipedia_url Wikipedia titles in the following languages (if they exist): ar (Arabic), de (German), en (English), es (Spanish), fa (Persian), fr (French), it (Italian) , ja (Japanese), nl (Dutch), pl… See the full description on the dataset page: https://huggingface.co/datasets/lewoniewski/grokipedia-wikipedia-16-languages.text100K<n<1M0 likes49 downloads10mo agoHugging Face26ScoutieAutoML /scoutieDataset_russian_language_grammar_and_rules_vectorized Description in English: A dataset collected from 30 Russian-language Telegram channels on the topic of learning the Russian language. This dataset contains grammar, syntax, spelling and punctuation rules. The dataset was collected and marked automatically using the Scoutie data collection and marking service.Try Scoutie and collect the same or another dataset using the link. Dataset fields: taskId - task identifier in the Scouti service. text - main text. url -… See the full description on the dataset page: https://huggingface.co/datasets/ScoutieAutoML/scoutieDataset_russian_language_grammar_and_rules_vectorized.tabulartext-classification10K<n<100K2 likes45 downloads2y agoHugging Face27neon-mao /language-dataset texttext-classification100K<n<1M3 likes43 downloads3y agoHugging Face28mahdi-hasan-shuvo /nagri-language-dataset Nagri Language Dataset - Description Nagri is a writing system for the Sylheti language, which is spoken in Bangladesh and India. The Nagri Language Dataset is a collection of text and image data specifically focused on Syloti Nagri script. This dataset is designed for OCR (Optical Character Recognition), handwriting recognition, and language modeling tasks related to the Syloti Nagri language. Dataset Features: ✅ Text Samples: Contains a variety of words, phrases… See the full description on the dataset page: https://huggingface.co/datasets/mahdi-hasan-shuvo/nagri-language-dataset.texttext-classificationn<1K1 likes40 downloads7mo agoHugging Face29lbourdois /language_tags Description Dataset listing 27,328 languages and dialects (also includes macrolanguage names).For each language, either the ISO 639 code, the Glottolog code or both are provided. Columns English_Name: Language name in English (e.g. "French"). Native_Name: If value is not 0, corresponds to the name of the language by native speakers (e.g. "Français") which may have been found in Wikipedia's nativename field. Glottocode: The language tag in the Glottolog convention (e.g.… See the full description on the dataset page: https://huggingface.co/datasets/lbourdois/language_tags.text10K<n<100K9 likes39 downloads3y agoHugging Face30md-nishat-008 /Code-Mixed-Offensive-Language-Detection-Dataset Code-Mixed-Offensive-Language-Identification This is a dataset for the offensive language detection task. It contains 100k code mixed data. The languages are Bangla-English-Hindi. Dataset Generation: Initially, the labelling schema of OLID[^1] and SOLID[^2] serves as the seed data, from which we randomly select 100,000 data instances. The labels in this dataset are categorized as Non-Offensive and Offensive for the purpose of our task. We meticulously ensure an equal… See the full description on the dataset page: https://huggingface.co/datasets/md-nishat-008/Code-Mixed-Offensive-Language-Detection-Dataset.text100K<n<1M1 likes38 downloads3y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.