datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
PHINCAbstract
Code-mixing is the phenomenon of using more than one language in a sentence. In the multilingual communities, it is a very frequently observed pattern of communication on social media platforms. Flexibility to use multiple languages in one text message might help to communicate efficiently with the target audience. But, the noisy user-generated code-mixed text adds to the challenge of processing and understanding natural language to a much larger extent. Machine translation from… See the full description on the dataset page: https://huggingface.co/datasets/LingoIITGN/PHINC.Tunisian-Proverbs-with-Image-Associations-A-Cultural-and-Linguistic-DatasetTunisian Proverbs with Image Associations: A Cultural and Linguistic Dataset
Description
This dataset explores the rich oral tradition of Tunisian proverbs mapped into text format, pairing each with contextual explanations, English translations both word-to-word and it's equivalent Target Language dynamic, Automated prompt and AI-generated visual interpretations.
It bridges linguistic, cultural, and visual modalities making it valuable for tasks in cross-cultural NLP, generative… See the full description on the dataset page: https://huggingface.co/datasets/HabibaAbderrahim/Tunisian-Proverbs-with-Image-Associations-A-Cultural-and-Linguistic-Dataset.lingnli-multi
Dataset Card for Dataset Name
Dataset Summary
This repository contains a collection of machine translations of LingNLI dataset
into 9 different languages (Bulgarian, Finnish, French, Greek, Italian, Korean, Lithuanian, Portuguese, Spanish). The goal is to predict textual entailment (does sentence A
imply/contradict/neither sentence B), which is a classification task (given two sentences,
predict one of three labels). It is here formatted in the same manner as the… See the full description on the dataset page: https://huggingface.co/datasets/maximoss/lingnli-multi.linguistic_calibrationThis Datasets repo contains training and evaluation datasets for the paper "Linguistic Calibration of Long-Form Generations".
Please refer to our GitHub repo at https://github.com/tatsu-lab/linguistic_calibration for more information, and check out our paper for our research findings: https://arxiv.org/abs/2404.00474
HinGEAbstract
Text generation is a highly active area of research in the computational linguistic community. The evaluation of the generated text is a challenging task and multiple theories and metrics have been proposed over the years. Unfortunately, text generation and evaluation are relatively understudied due to the scarcity of high-quality resources in code-mixed languages where the words and phrases from multiple languages are mixed in a single utterance of text and speech. To address this… See the full description on the dataset page: https://huggingface.co/datasets/LingoIITGN/HinGE.lingvanex_test_references
LTR
LTR -- Lingvanex Test References for MT Evaluation from English into a total of 30 target languages for a big variety of cases.
TEST CASES
Parameter
Description
Length
Sentences from 1 to 100 words.
Domain
Medicine (12%), Automobile (11%), Finance (8%)
Tokenizer
Jupiter is 1.000.000 km far. Ask Mr. Johnson for training
Tags
I want to eat and swim
Capitalisation (Case)
HELLO my Dear frIEND
Different languages in one text (Up to 3 languages)
I see… See the full description on the dataset page: https://huggingface.co/datasets/lingvanex/lingvanex_test_references.COMI-LINGUA
Dataset Details
COMI-LINGUA (COde-MIxing and LINGuistic Insights on Natural Hinglish Usage and Annotation) is a high-quality Hindi-English code-mixed dataset, manually annotated by three annotators. It serves as a benchmark for multilingual NLP models by covering multiple foundational tasks.
COMI-LINGUA provides annotations for several key NLP tasks:
Language Identification (LID): Token-wise classification of Hindi, English, and other linguistic units.
Initial predictions were… See the full description on the dataset page: https://huggingface.co/datasets/Tanushreeeeee/COMI-LINGUA.wikipedia-2023-11-kikongo-lingala-cohere-multilingual-v3lingnaam-cantonese-cot-qa
嶺南文化粵語思維鏈問答數據集
本數據集係由羊城晚報開源嘅 LNWHDMXSYS/lingnan-cantonese-cot-qa 改進而成。主要修改有:
將簡化字轉換成傳統漢字
依據粵文常見錯別字、粵語語氣詞規範用字
用 Google Cloud Translation v2 將官話表達翻譯成粵語
授權協議遵循源數據集嘅 cc-by-nc-4.0 許可證。
數據結構與字段說明
字段名稱
數據類型
是否必填
字段說明
id
Integer
係
樣本編號,自增主鍵
layer_name
String
係
文化層級,如 “ 表層文化/中層文化/深層文化 ” 等
domain
String
係
領域,如 “ 建築景觀/飲食文化/語言與語言學 ” 等
subcategory
String
係
子領域或子類,如 “ 嶺南建築 ” “ 傳統器物 ” “ 人生禮儀 ” 等
tag
String
係
主題標籤,更細粒度描述知識點,如 “ 騎樓 ” “ 碉樓 ” 等
subject
String
係… See the full description on the dataset page: https://huggingface.co/datasets/CanCLID/lingnaam-cantonese-cot-qa.PersonaEval
PersonaEval: A Benchmark for Role Identification in Dialogues
This dataset is released with the COLM 2025 conference paper: "PersonaEval: Are LLM Evaluators Human Enough to Judge Role-Play?".
PersonaEval is the first benchmark designed to test whether Large Language Models (LLMs) can reliably identify character roles from natural dialogue. We argue that correctly identifying who is speaking is a fundamental prerequisite for any meaningful evaluation of role-playing quality (how… See the full description on the dataset page: https://huggingface.co/datasets/lingfengzhou/PersonaEval.COMI-LINGUA
Dataset Details
COMI-LINGUA (COde-MIxing and LINGuistic Insights on Natural Hinglish Usage and Annotation) is a high-quality Hindi-English code-mixed dataset, manually annotated by three annotators. It serves as a benchmark for multilingual NLP models by covering multiple foundational tasks.
COMI-LINGUA provides annotations for several key NLP tasks:
Language Identification (LID): Token-wise classification of Hindi, English, and other linguistic units.
Initial predictions were… See the full description on the dataset page: https://huggingface.co/datasets/NikhilRajSoni/COMI-LINGUA.grading-question-triage-datasetcross_lingual_wsd_en_ro
English-Romanian Cross-Lingual Word Sense Disambiguation Dataset
This dataset accompanies the paper "Cross-Lingual Word Sense Disambiguation Remains Challenging for Large Language Models", accepted at KES 2026.
It contains a sense-aligned English-Romanian benchmark for evaluating monolingual and cross-lingual word sense disambiguation (WSD). Each row pairs an English sentence and a Romanian sentence that instantiate the same WordNet/RoWordNet synset.
The dataset is intended as… See the full description on the dataset page: https://huggingface.co/datasets/balmussebastian/cross_lingual_wsd_en_ro.Sylheti-Bangla-English-Russian-German-Parallel-Corpus-for-NLP
Sylheti-Bangla-English-Russian-German Parallel Corpus for NLP
Welcome to the first open-source multilingual parallel corpus for the Sylheti (syl) language, engineered by a native Linguistics student. This dataset bridges Sylheti with four major global high-resource languages spanning three distinct language families (Indo-Aryan, Germanic, and Slavic) to support Computational Linguistics (CL), Natural Language Processing (NLP) research, and Large Language Model (LLM) fine-tuning.… See the full description on the dataset page: https://huggingface.co/datasets/fahim-ling/Sylheti-Bangla-English-Russian-German-Parallel-Corpus-for-NLP.linguistic_sq
Physics and Math Problems Dataset
This repository contains a dataset of 3,200 enteries of different Albanian linguistics to improve Albanian queries further by introducing Albanian language rules and literature. The dataset is designed to support various NLP tasks and educational applications.
Dataset Overview
Problems: 3,200
Language: Albanian
Topics:
letërsi shqiptare: 47
poezi shqiptare: 45
proza shqiptare: 48
drama shqiptare: 46
autorë shqiptarë: 48
veprat kryesore… See the full description on the dataset page: https://huggingface.co/datasets/LTS-VVE/linguistic_sq.MMT
MMT (Multilingual and Multi‑Topic Twitter Language Identification Dataset)
MMT: A Multilingual and Multi‑Topic Indian Social Media Dataset is a large-scale language identification dataset derived from 1.7 million tweets collected from Indian Twitter/X, annotated with coarse and fine-grained language labels. It supports research on multilingual and code-mixed text in noisy, real-world social media settings.
📁 Files Included in This Release… See the full description on the dataset page: https://huggingface.co/datasets/LingoIITGN/MMT.SemMol_datasetlinguawave-competition
LinguaWave — Language Identification Competition
Pelatnas IOAI 2026 | Task 2 of 3
Identify the language of a 10-second speech clip from 8 languages. Compete to achieve the highest Macro F1-score on the test set.
Task
Input: .wav audio file (10 seconds, 16 kHz mono)Output: Language code from {id, ms, vi, th, en, zh, ar, fr}Metric: Macro F1-score
Languages
Code
Language
Region
id
Indonesian
Southeast Asia
ms
Malay
Southeast Asia
vi… See the full description on the dataset page: https://huggingface.co/datasets/fassabilf/linguawave-competition.kamba-lingala_sentence-pairs
Kamba-Lingala_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Kamba-Lingala_Sentence-Pairs
Number of Rows: 50317
Number of… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/kamba-lingala_sentence-pairs.Gurukul
Gurukul
Gurukul is an educational question-answering dataset aligned with the Indian school curriculum, building on the original Gurukul series. It contains high-quality QA pairs derived from Class-level textbooks (primarily English prose and related subjects), designed to support reading comprehension, vocabulary building, inference, and curriculum-based language understanding in educational AI applications.
Overview
Gurukul provides structured question-answer… See the full description on the dataset page: https://huggingface.co/datasets/LingoIITGN/Gurukul.PoliWAM
🗂️ Dataset Overview
PoliWAM is a large-scale corpus of WhatsApp political discussions collected during the Indian General Elections 2019. It consists of both raw and annotated data, enabling research in political discourse, misinformation, propaganda, and multilingual code-mixing.
Total Messages: 223,000
Groups: 281 public political groups
Users: ~31,000 unique users
Annotation Subset: 3,848 messages manually labeled for:
Political Orientation
Linguistic Composition… See the full description on the dataset page: https://huggingface.co/datasets/LingoIITGN/PoliWAM.bemba-lingala_sentence-pairs
Bemba-Lingala_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Bemba-Lingala_Sentence-Pairs
Number of Rows: 140378
Number of… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/bemba-lingala_sentence-pairs.lingala-swahili_sentence-pairs
Lingala-Swahili_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Lingala-Swahili_Sentence-Pairs
Number of Rows: 391907
Number… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/lingala-swahili_sentence-pairs.English-Giriama-Dataset
🗣️ English-Giriama Parallel Sentence Dataset
This dataset consists of sentence pairs in English and their corresponding translations in Giriama (Kigiryama), a Bantu language spoken primarily in coastal Kenya. It supports machine translation (MT) and other cross-lingual NLP tasks, especially in low-resource language research.
📋 Dataset Structure
Each row in the dataset contains:
English Sentence: A sentence in standard English.
Giriama Translation: The corresponding… See the full description on the dataset page: https://huggingface.co/datasets/Lingua-Connect/English-Giriama-Dataset.databricks-dolly-15k-context-3k-ragMulti-lingual_Detectiondyula-lingala_sentence-pairs
Dyula-Lingala_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Dyula-Lingala_Sentence-Pairs
Number of Rows: 57764
Number of… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/dyula-lingala_sentence-pairs.dinka-lingala_sentence-pairs
Dinka-Lingala_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Dinka-Lingala_Sentence-Pairs
Number of Rows: 22370
Number of… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/dinka-lingala_sentence-pairs.gsl-multimodal-annotation
GSL Multimodal Annotation Dataset
Multimodal annotation of 20 signs from Ghanaian Sign Language (GSL),
capturing manual and non-manual phonological features across 14 columns
including handshape, location, movement, facial expression, mouth
pattern, and head movement.
Dataset description
This dataset accompanies a pilot Linked Data representation of GSL
(DOI: 10.5281/zenodo.20961293). It documents the annotation decisions,
uncertainties, and limitations… See the full description on the dataset page: https://huggingface.co/datasets/LINGUISTEUNICE/gsl-multimodal-annotation.lingala-yoruba_sentence-pairs
Lingala-Yoruba_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Lingala-Yoruba_Sentence-Pairs
Number of Rows: 146711
Number… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/lingala-yoruba_sentence-pairs.
