datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
language-identification
Dataset Card for Language Identification dataset
Dataset Summary
The Language Identification dataset is a collection of 90k samples consisting of text passages and corresponding language label.
This dataset was created by collecting data from 3 sources: Multilingual Amazon Reviews Corpus, XNLI, and STSb Multi MT.
Supported Tasks and Leaderboards
The dataset can be used to train a model for language identification, which is a multi-class text classification… See the full description on the dataset page: https://huggingface.co/datasets/papluca/language-identification.Language-DetectionLanguage_Detection
Language_Detection - Multilingual Text Classification Dataset
This dataset is a collection of multilingual text samples designed for training and predicting languages in Artificial Intelligence (AI), Machine Learning (ML), Deep Learning (DL), and Data Science (DS) applications. It contains labeled data that associates text samples with their respective languages, enabling language detection and classification tasks.
Dataset Overview
The dataset consists of two columns:… See the full description on the dataset page: https://huggingface.co/datasets/sakthivinash/Language_Detection.europarl_for_language_detection_10kBiasShadesInterested in contributing? Speak a language not represented here? Disagree with an annotation? Please submit feedback in the Community tab!
Dataset Card for BiasShades
Note: This dataset may NOT be used as training data in any form (pre-training, fine-tuning, post-training, etc.) without express permission from creators.
Dataset Details
Version: 1.0
License: SHADES 1 Montreal Data License
Dataset Description
728 stereotypes and associated… See the full description on the dataset page: https://huggingface.co/datasets/LanguageShades/BiasShades.language-decoded-experiments
Language Decoded — Experiment Tracking
Central hub for training logs, configurations, evaluation results, and analysis for the Language Decoded project. The project originated as a proposal during Cohere's Tiny Aya Expedition (March 2026 hackathon) and was extended into Phase 3 for the accompanying paper.
Submitted paper title (2026-05-26): Language, Decoded: Exploring the Impact of Fine-Tuning a Multilingual Model on Native-Language Code
⚠️ Phase 3 numbers — read… See the full description on the dataset page: https://huggingface.co/datasets/legesher/language-decoded-experiments.language_detection
[!NOTE]
Dataset origin: https://www.kaggle.com/datasets/basilb2s/language-detection
It's a small language detection dataset. This dataset consists of text details for 17 different languages, ie, you will be able to create an NLP model for predicting 17 different language..
language-detectionKashmiri-Language-Corpus
Kashmiri Textual Data Corpus
Introduction
This repository contains a combined dataset of Kashmiri textual data collected from various sources. The data has been sourced from different locations and may contain non-Kashmiri text (e.g., Urdu, Persian). The goal of this corpus is to provide a wide variety of Kashmiri text data for research and language processing tasks.
Sources of Data
1. mzmmoazam/kashmiri_dataset (HTML Data)
Source: GitHub -… See the full description on the dataset page: https://huggingface.co/datasets/nawabhussain/Kashmiri-Language-Corpus.language_codes_marianMTturkish-offensive-language-detection
Dataset Summary
This dataset is enhanced version of existing offensive language studies. Existing studies are highly imbalanced, and solving this problem is too costly. To solve this, we proposed contextual data mining method for dataset augmentation. Our method is basically prevent us from retrieving random tweets and label individually. We can directly access almost exact hate related tweets and label them directly without any further human interaction in order to solve imbalanced… See the full description on the dataset page: https://huggingface.co/datasets/Toygar/turkish-offensive-language-detection.Natural_Language_to_Ffmpeg_Commands
Natural Language to FFmpeg Dataset
Disclaimer: This dataset was synthetically generated using a large language model and is intended for research purposes only. The dataset may contain inaccuracies, errors, or inconsistencies. Users should exercise caution and verify the correctness of the data before using it in any application.
This dataset contains 1000+ pairs of English natural language instructions and corresponding FFmpeg commands.
The dataset is designed for tasks… See the full description on the dataset page: https://huggingface.co/datasets/burak29/Natural_Language_to_Ffmpeg_Commands.language-identificationturkish-toxic-language
Turkish Texts for Toxic Language Detection
Dataset Description
Dataset Summary
This text dataset is a collection of Turkish texts that have been merged from various existing offensive language datasets found online. The dataset contains a total of 77,800 instances, each labeled as either offensive or not offensive.
To ensure the dataset's completeness, we utilized multiple transformer models to augment the dataset with pseudo labels. The resulting dataset is… See the full description on the dataset page: https://huggingface.co/datasets/Overfit-GM/turkish-toxic-language.Automated_Hate_Speech_Detection_and_the_Problem_of_Offensive_Languageafrican-language-parallel-corpus
African Language Parallel Corpus
Human-created, human-validated parallel sentence pairs for three African languages,
released openly by Okwu. Version 1.0.
Dataset summary
A parallel corpus of everyday-register sentence pairs for Yorùbá, Swahili, and
Nigerian Pidgin, each paired with English. The core is derived from NKENNE's own
language-learning curriculum — content authored and reviewed by native-speaker educators —
supplemented for Swahili with public-domain… See the full description on the dataset page: https://huggingface.co/datasets/Okwu/african-language-parallel-corpus.git-natural-language-commands
Git Natural Language Commands
A dataset mapping English natural-language instructions to their corresponding git commands, intended for training and evaluating models that translate user intent into safe, correct shell commands.
Disclaimer: This dataset was generated using Large Language Models (LLMs). The examples have not been manually verified against real-world usage and may contain errors, inconsistencies, or non-canonical phrasings. Use with appropriate caution.… See the full description on the dataset page: https://huggingface.co/datasets/burak29/git-natural-language-commands.Afar-language-text-to-speech-TTS
Usage
This dataset is designed to support the development of Text-to-Speech (TTS) systems for the Afar language. It can be integrated into web applications, mobile apps, desktop software, or other platforms that require natural-sounding Afar voice synthesis or accurate spoken language recognition.
For applications involving virtual avatars or voice personas, the following culturally appropriate voice names are recommended:
Female Voices: Emeli, Hanaawi, Kareera, Laysani, Kulsuma… See the full description on the dataset page: https://huggingface.co/datasets/Charif-Ayfarah/Afar-language-text-to-speech-TTS.Tunisian_Language_Dataset
Dataset Card for Tunisian Text Compilation
This dataset is a curated compilation of various Tunisian datasets, aimed at gathering as much Tunisian text data as possible in one place. It combines multiple sources of Tunisian language data, providing a rich resource for research, development of NLP models, and linguistic studies on Tunisian text.
Dataset Details
Dataset Description
This dataset aggregates several publicly available datasets that contain Tunisian… See the full description on the dataset page: https://huggingface.co/datasets/AzizBelaweid/Tunisian_Language_Dataset.Language-Identificationgsm8k-translated
Multilingual GSM8K Translations
This dataset contains machine-translated versions of GSM8K in these languages:
French (fr)
German (de)
Hindi (hi)
Dataset Structure
For each language, we provide the original GSM8K train and test splits:
train: 7,473 samples
test: 1,319 samples
Each sample consists of a question and an answer.
The question describes a grade-school-level math word problem that requires multi-step mathematical reasoning. The answer contains a… See the full description on the dataset page: https://huggingface.co/datasets/math-across-languages/gsm8k-translated.English-to-Afar-language-translation
Author
Created by Charif Ayfarah.
Contact: afbarit@gmail.com
License
Licensed under CC BY 4.0. You are free to use, modify, and distribute this dataset, including for commercial purposes, as long as you give appropriate credit.
language-energy-divide
🌍⚡ The Language–Energy Divide
Per-language energy measurements & prompts for multilingual LLM inference
📢 News
Aug 2026 — Our paper has been accepted to EMNLP 2026 (Main Conference)! 🎉
This dataset accompanies the paper "The Language–Energy Divide: Measuring Energy Costs of
Multilingual LLM Inference." It releases the per-language energy measurements and the
prompts used in the study, so researchers can build on our numbers without… See the full description on the dataset page: https://huggingface.co/datasets/MichiganNLP/language-energy-divide.thai-local-language-translation-dataset
Thai Local Language Translation Dataset
Thai Local Language Translation Dataset is a translation dataset for translate Thai Local Language to Thai Central Language. We create the dataset from Thai Dialect Corpus (Thai dialects ASR corpus). We select train set only from Thai Dialect Corpus.
The dataset support Khummuang, Korat, and Pattani.
Reference
Suwanbandit, A., Naowarat, B., Sangpetch, O., Chuangsuwanich, E. (2023) Thai Dialect Corpus and Transfer-based Curriculum… See the full description on the dataset page: https://huggingface.co/datasets/pythainlp/thai-local-language-translation-dataset.grokipedia-wikipedia-16-languages
Dataset description
This dataset contains a mapping between Grokipedia v0.1 article pages and the corresponding Wikipedia article titles across 16 language editions (based on Wikipedia and Wikidata dumps from 1 November 2025). Each record includes:
The URL of the Grokipedia page: grokipedia_url
Wikipedia titles in the following languages (if they exist): ar (Arabic), de (German), en (English), es (Spanish), fa (Persian), fr (French), it (Italian) , ja (Japanese), nl (Dutch), pl… See the full description on the dataset page: https://huggingface.co/datasets/lewoniewski/grokipedia-wikipedia-16-languages.scoutieDataset_russian_language_grammar_and_rules_vectorized
Description in English:
A dataset collected from 30 Russian-language Telegram channels on the topic of learning the Russian language. This dataset contains grammar, syntax, spelling and punctuation rules.
The dataset was collected and marked automatically using the Scoutie data collection and marking service.Try Scoutie and collect the same or another dataset using the link.
Dataset fields:
taskId - task identifier in the Scouti service. text - main text. url -… See the full description on the dataset page: https://huggingface.co/datasets/ScoutieAutoML/scoutieDataset_russian_language_grammar_and_rules_vectorized.language-dataset
nagri-language-dataset
Nagri Language Dataset - Description
Nagri is a writing system for the Sylheti language, which is spoken in Bangladesh and India.
The Nagri Language Dataset is a collection of text and image data specifically focused on Syloti Nagri script. This dataset is designed for OCR (Optical Character Recognition), handwriting recognition, and language modeling tasks related to the Syloti Nagri language.
Dataset Features:
✅ Text Samples: Contains a variety of words, phrases… See the full description on the dataset page: https://huggingface.co/datasets/mahdi-hasan-shuvo/nagri-language-dataset.language_tags
Description
Dataset listing 27,328 languages and dialects (also includes macrolanguage names).For each language, either the ISO 639 code, the Glottolog code or both are provided.
Columns
English_Name: Language name in English (e.g. "French").
Native_Name: If value is not 0, corresponds to the name of the language by native speakers (e.g. "Français") which may have been found in Wikipedia's nativename field.
Glottocode: The language tag in the Glottolog convention (e.g.… See the full description on the dataset page: https://huggingface.co/datasets/lbourdois/language_tags.Code-Mixed-Offensive-Language-Detection-Dataset
Code-Mixed-Offensive-Language-Identification
This is a dataset for the offensive language detection task. It contains 100k code mixed data. The languages are Bangla-English-Hindi.
Dataset Generation:
Initially, the labelling schema of OLID[^1] and SOLID[^2] serves as the seed data, from which we randomly select 100,000 data instances. The labels in this dataset are categorized as Non-Offensive and Offensive for the purpose of our task. We meticulously ensure an equal… See the full description on the dataset page: https://huggingface.co/datasets/md-nishat-008/Code-Mixed-Offensive-Language-Detection-Dataset.
