datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
tydiqa-goldpTyDi QA is a question answering dataset covering 11 typologically diverse languages with 204K question-answer pairs.
The languages of TyDi QA are diverse with regard to their typology -- the set of linguistic features that each language
expresses -- such that we expect models performing well on this set to generalize across a large number of the languages
in the world. It contains language phenomena that would not be found in English-only corpora. To provide a realistic
information-seeking task and avoid priming effects, questions are written by people who want to know the answer, but
don’t know the answer yet, (unlike SQuAD and its descendents) and the data is collected directly in each language without
the use of translation (unlike MLQA and XQuAD).model-written-evalsThis new dataset is designed to solve this great NLP task and is crafted with a lot of care.qa-expert-multi-hop-qa-V1.0
Dataset Card for QA-Expert-multi-hop-qa-V1.0
This dataset aims to provide multi-domain training data for the task: Question Answering, with a focus on Multi-hop Question Answering.
In total, this dataset contains 25.5k for training and 3.19k for evaluation.
You can take a look at the model we trained on this data: https://huggingface.co/khaimaitien/qa-expert-7B-V1.0
The dataset is mostly generated using the OpenAPI model (gpt-3.5-turbo-instruct). Please read more information about… See the full description on the dataset page: https://huggingface.co/datasets/khaimaitien/qa-expert-multi-hop-qa-V1.0.hunter-llm-sft-v1
HunterLLM SFT v1
Instruction + preference dataset for training a bug-bounty / offensive-security
assistant that thinks like a red-teamer, prioritizes attacker primitives and
reachable impact, and writes triager-friendly reports — strictly scoped to
authorized engagements (bug bounty programs, contracted pentests, isolated
labs).
Files
File
Rows
Schema
sft_train.jsonl
~33.1k
{instruction, input, output, tags, meta}
dpo_pairs.jsonl
~33.1k
{prompt, chosen… See the full description on the dataset page: https://huggingface.co/datasets/jabir-khan/hunter-llm-sft-v1.MSU-Benchmark
MSU-Bench: Towards Speaker-Centric Understanding in Conversational Multi-Speaker Scenarios
Interspeech 2026 · ASLP@NPU (Northwestern Polytechnical University), in collaboration with Li Auto.
Zhaokai Sun*, Shuai Wang*, Zhennan Lin*, Chengyou Wang, Dehui Gao, Yuang Cao, Chunjiang He, Pan Zhou, Lei Xie**
Audio, Speech and Language Processing Group (ASLP@NPU), School of Software, Northwestern Polytechnical University, China
School of Intelligent Science and Technology, Nanjing… See the full description on the dataset page: https://huggingface.co/datasets/Khalilah-Shields/MSU-Benchmark.AARA_Azerbaijani_LLM_Benchmark
AARA: Azerbaijani Advanced Reasoning Assessment
This dataset is the Azerbaijani-translated version of the emre/TARA_Turkish_LLM_Benchmark.
khayyam-challengetydiqa-primaryTyDi QA is a question answering dataset covering 11 typologically diverse languages with 204K question-answer pairs.
The languages of TyDi QA are diverse with regard to their typology -- the set of linguistic features that each language
expresses -- such that we expect models performing well on this set to generalize across a large number of the languages
in the world. It contains language phenomena that would not be found in English-only corpora. To provide a realistic
information-seeking task and avoid priming effects, questions are written by people who want to know the answer, but
don’t know the answer yet, (unlike SQuAD and its descendents) and the data is collected directly in each language without
the use of translation (unlike MLQA and XQuAD).CRAG-3sentences-chunks-v3
CRAG-3sentences-chunks-v3
Dataset Summary
Dataset ini berisi potongan (chunk) 3 kalimat untuk tiap passage hasil pascapengambilan (CRAG)
yang telah diolah dari khalidrizki/postretrieve-raw-dataset-v2. Tujuan utamanya adalah evaluasi RAG berbahasa Indonesia pada tugas QA. Dataset ini mirip dengan CRAG-3sentences-chunks-v2, tetapi dataset tersebut telah menghapus sekuens karakter yang menyusun sitasi (tanda kurung buka, angka, dan tanda kurung tutup). Di lain sisi, dataset… See the full description on the dataset page: https://huggingface.co/datasets/khalidrizki/CRAG-3sentences-chunks-v3.MMLU-Pro
MMLU-Pro Dataset
MMLU-Pro dataset is a more robust and challenging massive multi-task understanding dataset tailored to more rigorously benchmark large language models' capabilities. This dataset contains 12K complex questions across various disciplines.
|Github | 🏆Leaderboard | 📖Paper |
🚀 What's New
[2026.03.11] Added more cutting-edge frontier models to the leaderboard, including the Claude-4.6 series, Seed2.0 series, Qwen3.5 series, and Gemini-3.1-Pro, among… See the full description on the dataset page: https://huggingface.co/datasets/khaiise/MMLU-Pro.Atri-QA-Dataset-Vi-100K
Atri-QA-Dataset-Vi-100K 🌸
Dataset hội thoại tiếng Việt chất lượng cao cho việc huấn luyện Conversational AI
📋 Mục lục
Tổng quan
Thống kê Dataset
Cấu trúc dữ liệu
Phân loại nội dung
Quy trình tạo dữ liệu
Chất lượng dữ liệu
Cách sử dụng
Use Cases
[Hạn chế](#-hạn chế)
Ethical Considerations
Citation
🌟 Tổng quan
Atri-QA-Dataset-Vi-100K là một dataset hội thoại tiếng Việt chất lượng cao gồm 100,000 cặp câu hỏi-trả lời được thiết kế đặc biệt để huấn… See the full description on the dataset page: https://huggingface.co/datasets/khanhromvn/Atri-QA-Dataset-Vi-100K.my-distiset-1ecaaa6c
Dataset Card for my-distiset-1ecaaa6c
This dataset has been created with distilabel.
Dataset Summary
This dataset contains a pipeline.yaml which can be used to reproduce the pipeline that generated it in distilabel using the distilabel CLI:
distilabel pipeline run --config "https://huggingface.co/datasets/Khashayarrah/my-distiset-1ecaaa6c/raw/main/pipeline.yaml"
or explore the configuration:
distilabel pipeline info --config… See the full description on the dataset page: https://huggingface.co/datasets/Khashayarrah/my-distiset-1ecaaa6c.rag-mini-wikipediaIn this huggingface discussion you can share what you used the dataset for.
Derives from https://www.kaggle.com/datasets/rtatman/questionanswer-dataset?resource=download we generated our own subset using generate.py.
swin-faqskhasi-instruction-response-v1
Dataset Card for Khasi Instruction-Response Dataset v1
Dataset Summary
Khasi Instruction-Response Dataset v1 is a curated collection of prompts and responses in the Khasi language, designed to support the fine-tuning of instruction-following language models. It includes tasks such as Q&A, translation, summarization, and culturally-grounded dialogue. The dataset reflects indigenous knowledge systems, cultural expressions, and educational content unique to the Khasi context… See the full description on the dataset page: https://huggingface.co/datasets/toiar/khasi-instruction-response-v1.
