datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Traditional-Chinese-Medicine-Multiple_choice_question
Discription
This dataset is sourced from the website of the Ministry of Examination, R.O.C (Taiwan) and contains past exam questions from the national Traditional Chinese Medicine examinations in Taiwan. The exam comprises six subjects. This dataset specifically includes questions from two subjects, including the History of Traditional Chinese Medicine, Basic Theories of Traditional Chinese Medicine, Neijing, Nanjing, Traditional Chinese Medicine Prescription Studies, and… See the full description on the dataset page: https://huggingface.co/datasets/Liavan/Traditional-Chinese-Medicine-Multiple_choice_question.hle-multichoiceHumanity Last Exam dataset with extra incorrect answers generated with Qwen3-4B
IndustryBench
IndustryBench: Probing the Industrial Knowledge Boundaries of LLMs
💻Github | 📝Paper
IndustryBench is a multi-lingual benchmark for evaluating the industrial domain knowledge of large language models. It comprises 2,049 expert-curated QA pairs spanning 12 industrial sectors, with human-reviewed translations in Chinese, English, Russian, and Vietnamese.
Overview
Dimension
Details
Total questions
2,049
Languages
Chinese (zh), English (en), Russian (ru)… See the full description on the dataset page: https://huggingface.co/datasets/alibaba-multimodal-industrial-ai/IndustryBench.multilingual-elder-safety-msgs
multilingual-elder-safety-msgs
A hand-authored, multilingual elder fraud-recognition and safety coaching dataset. 467 curated scam/safe scenarios in Chinese and English, with platform-generated coaching responses localized across 5 languages: Chinese, English, Vietnamese, Khmer (Cambodian), and Lao. Expanded to 1,029 rows through Adaption Labs platform reasoning traces and multilingual adaptation.
Built for communities where filial piety, authority deference, and fear of… See the full description on the dataset page: https://huggingface.co/datasets/vanila434/multilingual-elder-safety-msgs.DeepResearch-Bench-Multilingual
DeepResearch Bench Multilingual Prompts
This dataset provides prompt-level multilingual translations for the 100 research tasks used in muset-ai/DeepResearch-Bench-Dataset.
The translations cover eight languages:
en
zh
es
it
ar
bn
ja
el
What is included
This repository focuses on the benchmark prompts only.
On the Hugging Face Hub, the Dataset Viewer is configured with one default subset named all plus nine explicit subset configurations: source_prompt, en, zh, es… See the full description on the dataset page: https://huggingface.co/datasets/JRQi/DeepResearch-Bench-Multilingual.Multi-Opthalingua
Cite
Accepted to AAAI 2025 (https://openreview.net/group?id=AAAI.org/2025/Conference#tab-recent-activity)
Multi-OphthaLingua: A Multilingual Benchmark for Assessing and Debiasing LLM Ophthalmological QA in LMICs:
@misc{restrepo2024multiophthalinguamultilingualbenchmarkassessing,
title={Multi-OphthaLingua: A Multilingual Benchmark for Assessing and Debiasing LLM Ophthalmological QA in LMICs},
author={David Restrepo and Chenwei Wu and Zhengxu Tang and Zitao Shuai and Thao… See the full description on the dataset page: https://huggingface.co/datasets/AAAIBenchmark/Multi-Opthalingua.synthetic_multilingual_llm_prompts
Image generated by DALL-E. See prompt for more details
📝🌐 Synthetic Multilingual LLM Prompts
Welcome to the "Synthetic Multilingual LLM Prompts" dataset! This comprehensive collection features 1,250 synthetic LLM prompts generated using Gretel Navigator, available in seven different languages. To ensure accuracy and diversity in prompts, and translation quality and consistency across the different languages, we employed Gretel Navigator both as a generation tool and as an… See the full description on the dataset page: https://huggingface.co/datasets/gretelai/synthetic_multilingual_llm_prompts.afrofinchain-multilingual-web3
AfroFinChain — Multilingual Web3 & Blockchain Dataset
Multilingual Web3 & blockchain dataset in Yoruba, Hausa, Igbo, and Nigerian Pidgin with 1,451 terminology entries and 1,451 conversational Q&A pairs. Designed for LLM fine-tuning, financial literacy, and conversational AI in low-resource African languages. Uses culturally grounded analogies (e.g., ajo, adashi, isusu) to make DeFi concepts actually understandable.
Built with Adaptive Data by Adaption as part of the Adaption… See the full description on the dataset page: https://huggingface.co/datasets/FirstBML1/afrofinchain-multilingual-web3.MultiQ
Dataset Card for MultiQ
This is the dataset corresponding to the paper "Evaluating the Elementary Multilingual Capabilities of Large Language Models with MultiQ".
It is a silver standard benchmark that can be used to evaluate the basic multilingual capabilities of LLMs. It contains 200 open ended questions automatically
translated into 137 typologically diverse languages.
Curated by: Carolin Holtermann, Paul Röttger, Timm Dill, Anne Lauscher
Language(s) (NLP): 137 diverse… See the full description on the dataset page: https://huggingface.co/datasets/caro-holt/MultiQ.BMGQ-MultiHop-Sample
🧩 BMGQ (Sample Release) – Bottom-up Multi-hop Question Generation Dataset
A Sampled Subset of BMGQ: Complex, Retrieval-Resistant, Multi-hop Reasoning Questions
👥 Authors
Bingsen Qiu, Zijian Liu, Xiao Liu, Bingjie Wang, Feier Zhang, Yixuan Qin, Chunyan Li, Haoshen Yang, Zeren Gao
📘 Dataset Summary
BMGQ is a dataset of complex, hard-to-search, multi-hop reasoning questions automatically generated using our proposed framework:
BMGQ: A Bottom-up… See the full description on the dataset page: https://huggingface.co/datasets/Fayer/BMGQ-MultiHop-Sample.multilevel-legal-reasoning
Legal Reasoning Dataset with Multilevel Human and Model-Annotated Explanations
Prepared by Mst Rafia Islam, Umong Sain, Azmine Toushik Wasi
Prepared as a part of Reasoning Datasets Competition by Bespoke Labs, Hugging Face, and Together.ai.
🧭 Purpose and Scope
The Legal Reasoning Dataset aims to support the evaluation and training of legal reasoning systems, particularly in multilingual or jurisdiction-agnostic contexts. It focuses on international acts and treaties… See the full description on the dataset page: https://huggingface.co/datasets/ciol-research/multilevel-legal-reasoning.multi-figqa
Dataset Card for multi-figqa
Dataset Summary
A multilingual dataset of human-written creative figurative expressions in many languages (mostly metaphors and similes). The English version (with the same format) can be found here
Languages
Languages included are Hindi, Indonesian, Javanese, Kannada, Sundanese, Swahili, and Yoruba. The language codes are respectively hi, id, kn, su, sw, and yo.
Dataset Structure
Data Instances
{… See the full description on the dataset page: https://huggingface.co/datasets/cmu-lti/multi-figqa.majid-multitask-dataset
Majid Multi-task Dataset | مجموعه داده چندوظیفهای مجید
English Description 🇬🇧
This dataset provides Persian–English text for multi-task NLP: text classification and question answering, including subtasks such as sentiment analysis and toxicity detection.
Dataset Details
Curated by: Maid121232 (Majid)
Languages: Persian (fa), English (en)
License: Apache-2.0
Size: Small (< 10k samples)
Tasks: Text Classification, Question Answering
Dataset Structure
Files: train.csv… See the full description on the dataset page: https://huggingface.co/datasets/Maid121232/majid-multitask-dataset.genetic-counselor-multiple-choiceA collection of multiple-choice questions intended for students preparing for the
American Board of Genetic Counseling (ABGC) Certification Examination.
Also see the genetic-counselor-freeform-questions evaluation set.
A genetic counselor must be prepared to answer questions about inheritance of traits,
medical statistics, testing, empathetic and ethical conversations with patients,
and observing symptoms.
For evaluation only
The goal of this dataset is to evaluate LLMs and… See the full description on the dataset page: https://huggingface.co/datasets/monsoon-nlp/genetic-counselor-multiple-choice.abc-multiple-choice
abc-multiple-choice Dataset
abc-multiple-choice は、競技クイズの大会「abc」で使用された4択問題を元に作成された、多肢選択式の質問応答データセットです。
データセットの詳細については、下記の発表資料を参照してください。
鈴木正敏. 4択クイズを題材にした多肢選択式日本語質問応答データセットの構築. 言語処理学会第30回年次大会 (NLP2024) 併設ワークショップ 日本語言語資源の構築と利用性の向上 (JLR2024), 2024. [PDF]
下記の GitHub リポジトリで、本データセットを用いた評価実験のスクリプトを管理しています。
https://github.com/cl-tohoku/abc-multiple-choice
ライセンス
本データセットのクイズ問題の著作権は abc/EQIDEN 実行委員会 に帰属します。
本データセットは研究目的での利用許諾を得ているものです。商用目的での利用は不可とします。
Multilingual_CulturalBench
Multilingual CulturalBench (Translated Subset)
This dataset is a multilingual extension of the CulturalBench dataset ("Easy" subset). It contains 787 samples from the original benchmark, translated into five additional languages using Gemini-2.5-Flash.
Dataset Description
The original CulturalBench is designed to assess the cultural capabilities of Large Language Models (LLMs). This version extends the "Easy" subset (multiple-choice questions) by providing translations… See the full description on the dataset page: https://huggingface.co/datasets/Lossfunk/Multilingual_CulturalBench.multi-tafseer-quran-rag
Quran Tafseer RAG Dataset
A structured Arabic dataset of Quranic tafseer collected from eight classical and modern tafseer books.The dataset contains verse-aligned tafseer passages designed for Retrieval-Augmented Generation (RAG) systems and Arabic NLP research.
Each record links a Quran verse with its corresponding tafseer explanation from one of the tafseer books and includes rich metadata such as surah information, tafseer source, and embedding-ready text.
The dataset was… See the full description on the dataset page: https://huggingface.co/datasets/omaressam1111/multi-tafseer-quran-rag.JourneyBench_Multi_Image_VQA
Dataset Card for Dataset Name
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Dataset Details
Dataset Description
Curated by: [More Information Needed]
Funded by [optional]: [More Information Needed]
Shared by [optional]: [More Information Needed]
Language(s) (NLP): [More Information Needed]
License: [More Information Needed]
Dataset Sources [optional]
Repository: [More… See the full description on the dataset page: https://huggingface.co/datasets/JourneyBench/JourneyBench_Multi_Image_VQA.FRACTURED-SORRY-Bench-Automated-Multishot-Jailbreak
FRACTURED-SORRY-Bench: Framework for Revealing Attacks in Conversational Turns Undermining Refusal Efficacy and Defenses over SORRY-Bench (Automated Multi-shot Jailbreaks)
Dataset Card for FRACTURED-SORRY-Bench Dataset
🌐Website
📑Paper
📚Dataset
💻Github
FRACTURED-SORRY-Bench is a framework for evaluating the safety of Large Language Models (LLMs) against multi-turn conversational attacks. Building upon the SORRY-Bench dataset, we propose a simple… See the full description on the dataset page: https://huggingface.co/datasets/AmanPriyanshu/FRACTURED-SORRY-Bench-Automated-Multishot-Jailbreak.a2z-multidomain-glossary
A–Z Multi-Domain Glossary Dataset
This dataset is a creative collection of A-to-Z terminology across a wide range of high-level domains including Agriculture, Technology, Environment, Artificial Intelligence, Zoology, and more.Each entry includes:
domain
letter (A–Z)
word
description (short)
📊 Structure
Column
Description
domain
The high-level category (e.g. Technology, Agriculture)
letter
The alphabetical letter from A to Z
word
The concept/keyword… See the full description on the dataset page: https://huggingface.co/datasets/tejasashinde/a2z-multidomain-glossary.Multilingual_modelMultiState-DMV-Licensing-Practice-Set
Introduction
This dataset contains structured practice questions and answers derived from multiple DMV sample materials. It is designed to support training and evaluation of models for multiple-choice question answering, rule-based classification, and natural language understanding tasks in the driving regulations domain.
Key Features
Region-Specific: Focused on multiple state driving laws and traffic rules
Use Case: DMV permit/license test practice modeling… See the full description on the dataset page: https://huggingface.co/datasets/nprak26/MultiState-DMV-Licensing-Practice-Set.multilingual-llm-evaluation
Multilingual LLM Evaluation
A small evaluation dataset for comparing language models across English, Hindi, and Spanish.
Columns
language: language code (en, hi, or es)
question: question provided to the model
expected_answer: reference answer used for scoring
Intended use
This dataset can be used to compare model accuracy, language adherence, and response speed across languages.
Limitations
This is a small demonstration dataset and… See the full description on the dataset page: https://huggingface.co/datasets/userhuggingface4321/multilingual-llm-evaluation.alpine1.1-multireq-instructions-seedThis dataset is a refined version of Alpine 1.0. It was created by generating tasks using various LLMs, wrapping them in special elements {Instruction Start} ... {Instruction End}, and saving them in a text file. We then processed this file with a Python script that used regex to extract the tasks into a CSV. Afterward, we cleaned the dataset by removing near-duplicates, vague prompts, and ambiguous entries.
python clean.py -i prompts.csv -o cleaned.csv -p "prompt" -t 0.92 -l 30
This dataset… See the full description on the dataset page: https://huggingface.co/datasets/marcuscedricridia/alpine1.1-multireq-instructions-seed.alpine-1.0-multireq-instructionsthis dataset was made by generating the prompts using a mix of llms and answered by the gemini api smallest models. the purpose of this dataset is to improve the instruction-following of our models.
this is the multi-request subset of the Alpine dataset.
will be denoised and cleaned for efficient training. some responses are wrong and unhelpful.
