datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Chinese_Multi-Emotion_Dialogue_Dataset
Chinese_Multi-Emotion_Dialogue_Dataset
📄 Description
This dataset contains 4159 Chinese dialogues annotated with 8 distinct emotion categories. The data is suitable for emotion recognition, sentiment analysis, and other NLP tasks involving Chinese text.
Data Sources:
Daily Conversations: Captured from natural, informal human conversations.
Movie Dialogues: Extracted from diverse Chinese-language movies.
AI-Generated Dialogues: Synthesized using… See the full description on the dataset page: https://huggingface.co/datasets/Johnson8187/Chinese_Multi-Emotion_Dialogue_Dataset.multidomain-kazakh-dataset
Dataset Description
Point of Contact: Sanzhar Murzakhmetov, Besultan Sagyndyk
Dataset Summary
MDBKD | Multi-Domain Bilingual Kazakh Dataset is a Kazakh-language dataset containing just over 24 883 808 unique texts from multiple domains.
Supported Tasks
'MLM/CLM': can be used to train a model for casual and masked languange modeling
Languages
The kk code for Kazakh as generally spoken in the Kazakhstan
Data Instances
For each instance… See the full description on the dataset page: https://huggingface.co/datasets/kz-transformers/multidomain-kazakh-dataset.IndustryBench
IndustryBench: Probing the Industrial Knowledge Boundaries of LLMs
💻Github | 📝Paper
IndustryBench is a multi-lingual benchmark for evaluating the industrial domain knowledge of large language models. It comprises 2,049 expert-curated QA pairs spanning 12 industrial sectors, with human-reviewed translations in Chinese, English, Russian, and Vietnamese.
Overview
Dimension
Details
Total questions
2,049
Languages
Chinese (zh), English (en), Russian (ru)… See the full description on the dataset page: https://huggingface.co/datasets/alibaba-multimodal-industrial-ai/IndustryBench.multilingual-lima
Multilingual LIMA
A multilingual extension of the LIMA instruction-tuning dataset. The original English prompt–response pairs were translated into 9 additional typologically diverse languages with google/gemini-2.0-flash-001. Each language is stored as a separate Hugging Face config.
Field
Description
prompt
User instruction (translated; en is the original).
output
Assistant response (translated; en is the original).
Languages (configs): en (original), zh, it, bn… See the full description on the dataset page: https://huggingface.co/datasets/iNLP-Lab/multilingual-lima.DeepResearch-Bench-Multilingual
DeepResearch Bench Multilingual Prompts
This dataset provides prompt-level multilingual translations for the 100 research tasks used in muset-ai/DeepResearch-Bench-Dataset.
The translations cover eight languages:
en
zh
es
it
ar
bn
ja
el
What is included
This repository focuses on the benchmark prompts only.
On the Hugging Face Hub, the Dataset Viewer is configured with one default subset named all plus nine explicit subset configurations: source_prompt, en, zh, es… See the full description on the dataset page: https://huggingface.co/datasets/JRQi/DeepResearch-Bench-Multilingual.multilingual-s1
Multilingual s1
A multilingual extension of the s1K-1.1 reasoning dataset. The original English reasoning questions and DeepSeek-R1 distilled solutions were translated into 9 additional typologically diverse languages with google/gemini-2.0-flash-001. Each language is stored as a separate Hugging Face config.
We filter the upstream simplescaling/s1K-1.1 corpus to keep only samples whose DeepSeek-R1 trajectories were marked as correctly distilled, then translate the resulting subset.… See the full description on the dataset page: https://huggingface.co/datasets/iNLP-Lab/multilingual-s1.synthetic_multilingual_llm_prompts
Image generated by DALL-E. See prompt for more details
📝🌐 Synthetic Multilingual LLM Prompts
Welcome to the "Synthetic Multilingual LLM Prompts" dataset! This comprehensive collection features 1,250 synthetic LLM prompts generated using Gretel Navigator, available in seven different languages. To ensure accuracy and diversity in prompts, and translation quality and consistency across the different languages, we employed Gretel Navigator both as a generation tool and as an… See the full description on the dataset page: https://huggingface.co/datasets/gretelai/synthetic_multilingual_llm_prompts.Multi-IaC-Eval
Multi-IaC-Eval
We present Multi-IaC-Eval is a novel benchmark dataset for evaluating LLM-based IaC generation and mutation across AWS CloudFormation, Terraform, and Cloud Development Kit (CDK) formats. The dataset consists of triplets containing initial IaC templates, natural language modification requests, and corresponding updated templates, created through a synthetic data generation pipeline with rigorous validation.
Cloudformation: 263
Terraform: 446
CDK (Python): 64
CDK… See the full description on the dataset page: https://huggingface.co/datasets/AmazonScience/Multi-IaC-Eval.afrofinchain-multilingual-web3
AfroFinChain — Multilingual Web3 & Blockchain Dataset
Multilingual Web3 & blockchain dataset in Yoruba, Hausa, Igbo, and Nigerian Pidgin with 1,451 terminology entries and 1,451 conversational Q&A pairs. Designed for LLM fine-tuning, financial literacy, and conversational AI in low-resource African languages. Uses culturally grounded analogies (e.g., ajo, adashi, isusu) to make DeFi concepts actually understandable.
Built with Adaptive Data by Adaption as part of the Adaption… See the full description on the dataset page: https://huggingface.co/datasets/FirstBML1/afrofinchain-multilingual-web3.multilingual-safety
Multilingual Safety Instructions
A multilingual extension of the safety-only instruction–refusal pairs released with the Safety-Tuned LLaMAs project. The original 1,000 harmful-prompt / refusal-response pairs (English) were translated into 11 additional typologically diverse languages with google/gemini-2.0-flash-001. Each language is stored as a separate Hugging Face config.
Field
Description
prompt
Harmful user instruction (translated; en is the original).
output
Safe… See the full description on the dataset page: https://huggingface.co/datasets/iNLP-Lab/multilingual-safety.ponys-multilingual-ai-character-consistency-benchmark
Ponys Multilingual AI Character Consistency Benchmark
This repository contains a preregistered test instrument, not collected product
results and not an independent product ranking.
140 fixed test cases across seven locales
four dimensions: persona, register, relationship state, and visual identity
three planned clean-session runs per case
result state: not_collected
publisher: Ponys.ai Research (official first-party research)
official source: https://ponys.ai/
research feeds:… See the full description on the dataset page: https://huggingface.co/datasets/wujoe132/ponys-multilingual-ai-character-consistency-benchmark.multilevel-legal-reasoning
Legal Reasoning Dataset with Multilevel Human and Model-Annotated Explanations
Prepared by Mst Rafia Islam, Umong Sain, Azmine Toushik Wasi
Prepared as a part of Reasoning Datasets Competition by Bespoke Labs, Hugging Face, and Together.ai.
🧭 Purpose and Scope
The Legal Reasoning Dataset aims to support the evaluation and training of legal reasoning systems, particularly in multilingual or jurisdiction-agnostic contexts. It focuses on international acts and treaties… See the full description on the dataset page: https://huggingface.co/datasets/ciol-research/multilevel-legal-reasoning.MULTICOMMULTICOM V1.1
This repository hosts the MULTICOM dataset, a novel benchmark for evaluating the multilingual commonsense generation abilities of Large Language Models (LLMs), as presented in the paper Do LLMs exhibit the same commonsense capabilities across languages?.
The dataset extends the COCOTEROS dataset to four languages: English, Spanish, Dutch, and Valencian. The task involves generating a commonsensical sentence that includes a given triplet of words.
wikipedia-2023-11-kikongo-lingala-cohere-multilingual-v3multi-tafseer-quran-rag
Quran Tafseer RAG Dataset
A structured Arabic dataset of Quranic tafseer collected from eight classical and modern tafseer books.The dataset contains verse-aligned tafseer passages designed for Retrieval-Augmented Generation (RAG) systems and Arabic NLP research.
Each record links a Quran verse with its corresponding tafseer explanation from one of the tafseer books and includes rich metadata such as surah information, tafseer source, and embedding-ready text.
The dataset was… See the full description on the dataset page: https://huggingface.co/datasets/omaressam1111/multi-tafseer-quran-rag.Multi-Doc-Multi-QA-ChineseDeprecated, please use Multi-Doc-QA-Chinese instead.
文档和问答对都来自 Multi-Doc-QA-Chinese,通过随机抽取和组合形成多轮问答形式。
推荐直接使用原始数据集Multi-Doc-QA-Chinese自己生成指令微调数据,可以控制参考文档和问答的数量
经过随机组合,每条数据形成了 20-60个参考文档 + 10个问答对的形式
chat格式为chatml
illicit-general-multi-turn
Illicit General Multi-Turn Conversations
Multi-turn adversarial conversations that successfully elicited harmful illicit content from AI models. This sample dataset contains 5 conversations (52 turns) covering chemical weapons, cyber threats, and other safety-critical domains.
Dataset Statistics
Metric
Value
Conversations
5
Total Turns
52
Avg Turns/Conv
10.4
Harm Categories
3
Harm Categories
Category
Turns
Description
Chemical… See the full description on the dataset page: https://huggingface.co/datasets/GoJulyAI/illicit-general-multi-turn.IFAAB-MULTI-LLM-2026
Dataset Card for IFAAB-MULTI-LLM-2026
This dataset card serves as a comprehensive datasheet for the kmhj1306/IFAAB-MULTI-LLM-2026 dataset repository. It maps demographic persona features to localized automated financial planning prompts and responses, specifically curated to evaluate LLM behavior within the Indian socio-economic context.
Dataset Details
Dataset Description
This dataset consists of 222,138 rows of tabular text data designed to… See the full description on the dataset page: https://huggingface.co/datasets/kmhj1306/IFAAB-MULTI-LLM-2026.FRACTURED-SORRY-Bench-Automated-Multishot-Jailbreak
FRACTURED-SORRY-Bench: Framework for Revealing Attacks in Conversational Turns Undermining Refusal Efficacy and Defenses over SORRY-Bench (Automated Multi-shot Jailbreaks)
Dataset Card for FRACTURED-SORRY-Bench Dataset
🌐Website
📑Paper
📚Dataset
💻Github
FRACTURED-SORRY-Bench is a framework for evaluating the safety of Large Language Models (LLMs) against multi-turn conversational attacks. Building upon the SORRY-Bench dataset, we propose a simple… See the full description on the dataset page: https://huggingface.co/datasets/AmanPriyanshu/FRACTURED-SORRY-Bench-Automated-Multishot-Jailbreak.a2z-multidomain-glossary
A–Z Multi-Domain Glossary Dataset
This dataset is a creative collection of A-to-Z terminology across a wide range of high-level domains including Agriculture, Technology, Environment, Artificial Intelligence, Zoology, and more.Each entry includes:
domain
letter (A–Z)
word
description (short)
📊 Structure
Column
Description
domain
The high-level category (e.g. Technology, Agriculture)
letter
The alphabetical letter from A to Z
word
The concept/keyword… See the full description on the dataset page: https://huggingface.co/datasets/tejasashinde/a2z-multidomain-glossary.mirror
MIRROR Dataset
MIRROR is a synthetic vision–language dataset for multimodal cognitive reframing under client resistance.
Paper: 🪞 MIRROR: Multimodal Cognitive Reframing Therapy for Rolling with Resistance
The dataset includes:
Client profile metadata (CACTUS idx, CelebA idx)
Dialogue written in a screenplay format, including stage directions that describe facial expressions
⚠️ Images themselves are not included to comply with the CelebA license.
However, we provide the full image… See the full description on the dataset page: https://huggingface.co/datasets/multimodal-reframing/mirror.authorship-style-transfer-multilangual
Parallel neutral / author-style fine-tuning dataset
Tabular parallel text built from matched neutral (“standard”) and author-style sources. Each row is one chunk of several consecutive non-empty lines, paired so that the same semantic content appears in both columns.
Dataset statistics
Samples (CSV rows)
4,868
Hub size bucket
1K<n<10K (matches sample count)
Primary file
fine_tune_dataset.csv (UTF-8)
The metadata field size_categories refers to number… See the full description on the dataset page: https://huggingface.co/datasets/AhmedZaky1/authorship-style-transfer-multilangual.Chinese_Multi-Emotion_Dialogue_Dataset
Chinese_Multi-Emotion_Dialogue_Dataset
📄 Description
This dataset contains 4159 Chinese dialogues annotated with 8 distinct emotion categories. The data is suitable for emotion recognition, sentiment analysis, and other NLP tasks involving Chinese text.
Data Sources:
Daily Conversations: Captured from natural, informal human conversations.
Movie Dialogues: Extracted from diverse Chinese-language movies.
AI-Generated Dialogues: Synthesized using advanced… See the full description on the dataset page: https://huggingface.co/datasets/osuih/Chinese_Multi-Emotion_Dialogue_Dataset.Multilingual_modelillicit-bio-multi-turn
Illicit Bio Multi-Turn Conversations
Multi-turn adversarial conversations that successfully elicited harmful bio-safety content from AI models. This sample dataset contains 5 conversations (57 turns) covering bioweapons and related threats.
Dataset Statistics
Metric
Value
Conversations
5
Total Turns
57
Avg Turns/Conv
11.4
Harm Categories
3
Harm Categories
Category
Turns
Description
Bioweapons
34
Information about biological… See the full description on the dataset page: https://huggingface.co/datasets/GoJulyAI/illicit-bio-multi-turn.dore
You must agree to the license and terms of use before using the dataset in this repo.
DORE: Definition MOdelling in PoRtuguEse
This repository introduces DORE, a comprehensive corpus of over 100,000 definitions from Portuguese dictionaries. Alongside DORE, we also introduce the models used to perform Portuguese DM. The release of DORE aims to fill in the gap of resources for Automatic Definition Generation, or Definition Modelling (DM), in Portuguese. DORE is the first dataset… See the full description on the dataset page: https://huggingface.co/datasets/multidefmod/dore.psychology-multi-turn
Psychology Multi-Turn Conversations
Multi-turn adversarial conversations that successfully elicited harmful psychological content from AI models. This sample dataset contains 5 conversations (54 turns) covering anthropomorphism, psychosis, self-harm, etc.
Dataset Statistics
Metric
Value
Conversations
5
Total Turns
54
Avg Turns/Conv
10.8
Harm Categories
3
Harm Categories
Category
Turns
Description
Anthropomorphism
28… See the full description on the dataset page: https://huggingface.co/datasets/GoJulyAI/psychology-multi-turn.Roomly-Student-Bios-Multimodal
Roomly: Multimodal Roommate Matching Dataset
🎯 Problem Statement
Finding a roommate is often reduced to dry filters like "budget" and "location". Roomly aims to revolutionize this by focusing on personality, lifestyle, and visual preferences. This dataset provides synthetic student profiles and their ideal room environments.
📊 Exploratory Data Analysis (EDA)
1. User Persona Distribution
Our dataset contains a balanced mix of different student… See the full description on the dataset page: https://huggingface.co/datasets/Orib24/Roomly-Student-Bios-Multimodal.
