CoolFace
15 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01khalidalt /tydiqa-goldpTyDi QA is a question answering dataset covering 11 typologically diverse languages with 204K question-answer pairs. The languages of TyDi QA are diverse with regard to their typology -- the set of linguistic features that each language expresses -- such that we expect models performing well on this set to generalize across a large number of the languages in the world. It contains language phenomena that would not be found in English-only corpora. To provide a realistic information-seeking task and avoid priming effects, questions are written by people who want to know the answer, but don’t know the answer yet, (unlike SQuAD and its descendents) and the data is collected directly in each language without the use of translation (unlike MLQA and XQuAD).question-answering13 likes1.4k downloads2y agoHugging Face02khalidalt /model-written-evalsThis new dataset is designed to solve this great NLP task and is crafted with a lot of care.multiple-choice100K<n<1M2 likes334 downloads3y agoHugging Face03khaimaitien /qa-expert-multi-hop-qa-V1.0 Dataset Card for QA-Expert-multi-hop-qa-V1.0 This dataset aims to provide multi-domain training data for the task: Question Answering, with a focus on Multi-hop Question Answering. In total, this dataset contains 25.5k for training and 3.19k for evaluation. You can take a look at the model we trained on this data: https://huggingface.co/khaimaitien/qa-expert-7B-V1.0 The dataset is mostly generated using the OpenAPI model (gpt-3.5-turbo-instruct). Please read more information about… See the full description on the dataset page: https://huggingface.co/datasets/khaimaitien/qa-expert-multi-hop-qa-V1.0.textquestion-answering10K<n<100K8 likes117 downloads3y agoHugging Face04jabir-khan /hunter-llm-sft-v1 HunterLLM SFT v1 Instruction + preference dataset for training a bug-bounty / offensive-security assistant that thinks like a red-teamer, prioritizes attacker primitives and reachable impact, and writes triager-friendly reports — strictly scoped to authorized engagements (bug bounty programs, contracted pentests, isolated labs). Files File Rows Schema sft_train.jsonl ~33.1k {instruction, input, output, tags, meta} dpo_pairs.jsonl ~33.1k {prompt, chosen… See the full description on the dataset page: https://huggingface.co/datasets/jabir-khan/hunter-llm-sft-v1.texttext-generation10K<n<100K0 likes106 downloads4mo agoHugging Face05Khalilah-Shields /MSU-Benchmark MSU-Bench: Towards Speaker-Centric Understanding in Conversational Multi-Speaker Scenarios Interspeech 2026 · ASLP@NPU (Northwestern Polytechnical University), in collaboration with Li Auto. Zhaokai Sun*, Shuai Wang*, Zhennan Lin*, Chengyou Wang, Dehui Gao, Yuang Cao, Chunjiang He, Pan Zhou, Lei Xie** Audio, Speech and Language Processing Group (ASLP@NPU), School of Software, Northwestern Polytechnical University, China School of Intelligent Science and Technology, Nanjing… See the full description on the dataset page: https://huggingface.co/datasets/Khalilah-Shields/MSU-Benchmark.audioaudio-classification1K<n<10K0 likes81 downloads2mo agoHugging Face06khazarai /AARA_Azerbaijani_LLM_Benchmark AARA: Azerbaijani Advanced Reasoning Assessment This dataset is the Azerbaijani-translated version of the emre/TARA_Turkish_LLM_Benchmark. textquestion-answeringn<1K1 likes43 downloads6mo agoHugging Face07raia-center /khayyam-challengegatedtabularquestion-answering10K<n<100K13 likes38 downloads2mo agoHugging Face08khalidalt /tydiqa-primaryTyDi QA is a question answering dataset covering 11 typologically diverse languages with 204K question-answer pairs. The languages of TyDi QA are diverse with regard to their typology -- the set of linguistic features that each language expresses -- such that we expect models performing well on this set to generalize across a large number of the languages in the world. It contains language phenomena that would not be found in English-only corpora. To provide a realistic information-seeking task and avoid priming effects, questions are written by people who want to know the answer, but don’t know the answer yet, (unlike SQuAD and its descendents) and the data is collected directly in each language without the use of translation (unlike MLQA and XQuAD).question-answering0 likes25 downloads4y agoHugging Face09khalidrizki /CRAG-3sentences-chunks-v3 CRAG-3sentences-chunks-v3 Dataset Summary Dataset ini berisi potongan (chunk) 3 kalimat untuk tiap passage hasil pascapengambilan (CRAG) yang telah diolah dari khalidrizki/postretrieve-raw-dataset-v2. Tujuan utamanya adalah evaluasi RAG berbahasa Indonesia pada tugas QA. Dataset ini mirip dengan CRAG-3sentences-chunks-v2, tetapi dataset tersebut telah menghapus sekuens karakter yang menyusun sitasi (tanda kurung buka, angka, dan tanda kurung tutup). Di lain sisi, dataset… See the full description on the dataset page: https://huggingface.co/datasets/khalidrizki/CRAG-3sentences-chunks-v3.textquestion-answering1K<n<10K0 likes23 downloads1y agoHugging Face10khaiise /MMLU-Pro MMLU-Pro Dataset MMLU-Pro dataset is a more robust and challenging massive multi-task understanding dataset tailored to more rigorously benchmark large language models' capabilities. This dataset contains 12K complex questions across various disciplines. |Github | 🏆Leaderboard | 📖Paper | 🚀 What's New [2026.03.11] Added more cutting-edge frontier models to the leaderboard, including the Claude-4.6 series, Seed2.0 series, Qwen3.5 series, and Gemini-3.1-Pro, among… See the full description on the dataset page: https://huggingface.co/datasets/khaiise/MMLU-Pro.tabularquestion-answering10K<n<100K0 likes23 downloads5mo agoHugging Face11khanhromvn /Atri-QA-Dataset-Vi-100K Atri-QA-Dataset-Vi-100K 🌸 Dataset hội thoại tiếng Việt chất lượng cao cho việc huấn luyện Conversational AI 📋 Mục lục Tổng quan Thống kê Dataset Cấu trúc dữ liệu Phân loại nội dung Quy trình tạo dữ liệu Chất lượng dữ liệu Cách sử dụng Use Cases [Hạn chế](#-hạn chế) Ethical Considerations Citation 🌟 Tổng quan Atri-QA-Dataset-Vi-100K là một dataset hội thoại tiếng Việt chất lượng cao gồm 100,000 cặp câu hỏi-trả lời được thiết kế đặc biệt để huấn… See the full description on the dataset page: https://huggingface.co/datasets/khanhromvn/Atri-QA-Dataset-Vi-100K.text-generation100K<n<1M1 likes19 downloads11mo agoHugging Face12Khashayarrah /my-distiset-1ecaaa6c Dataset Card for my-distiset-1ecaaa6c This dataset has been created with distilabel. Dataset Summary This dataset contains a pipeline.yaml which can be used to reproduce the pipeline that generated it in distilabel using the distilabel CLI: distilabel pipeline run --config "https://huggingface.co/datasets/Khashayarrah/my-distiset-1ecaaa6c/raw/main/pipeline.yaml" or explore the configuration: distilabel pipeline info --config… See the full description on the dataset page: https://huggingface.co/datasets/Khashayarrah/my-distiset-1ecaaa6c.texttext-generationn<1K0 likes15 downloads2y agoHugging Face13Khatwanigaurav /rag-mini-wikipediaIn this huggingface discussion you can share what you used the dataset for. Derives from https://www.kaggle.com/datasets/rtatman/questionanswer-dataset?resource=download we generated our own subset using generate.py. textquestion-answering1K<n<10K0 likes12 downloads6mo agoHugging Face14khaxtran /swin-faqstextquestion-answeringn<1K1 likes5 downloads3y agoHugging Face15toiar /khasi-instruction-response-v1gated Dataset Card for Khasi Instruction-Response Dataset v1 Dataset Summary Khasi Instruction-Response Dataset v1 is a curated collection of prompts and responses in the Khasi language, designed to support the fine-tuning of instruction-following language models. It includes tasks such as Q&A, translation, summarization, and culturally-grounded dialogue. The dataset reflects indigenous knowledge systems, cultural expressions, and educational content unique to the Khasi context… See the full description on the dataset page: https://huggingface.co/datasets/toiar/khasi-instruction-response-v1.texttranslation1K<n<10K0 likes1 downloads9mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.