CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01FuhaiLiAiLab /Target-QA 🎯 Target-QA: The First QA Dataset Benchmarking Target Priorization Based on DepMap 📑 Dataset Summary Target-QA is derived from the DepMap multi-omics and CRISPR screening cohorts, harmonized via BioMedGraphica.It enables multi-modal reasoning by combining numeric evidence, topological knowledge and language context for CRISPR target prioritization. This dataset supports the training and benchmarking of… See the full description on the dataset page: https://huggingface.co/datasets/FuhaiLiAiLab/Target-QA.tabularquestion-answeringn<1K0 likes255 downloads1y agoHugging Face02Targoman /TLPCgated Targoman Large Persian Corpus Dataset Summary Ever since the invention of the computer, humans have always been interested in being able to communicate with computers in human language. The research and endeavors of scientists and engineers brought us to the knowledge of human language processing and large language models. Large Language Models (LLMs) with their amazing capabilities, have caused a revolution in the artificial intelligence industry, which has caused a… See the full description on the dataset page: https://huggingface.co/datasets/Targoman/TLPC.question-answering43 likes187 downloads2y agoHugging Face03Dayanand314Krishna /cross_rulings_hts_dataset_for_tariffs CROSS Rulings HTS Dataset for Tariff Classification Maintained by Flexify.AI Inc. as part of the ATLAS trade intelligence research program. Paper: ATLAS: Benchmarking and Adapting LLMs for Global Trade via Harmonized Tariff Code Classification Project Page: https://tariffpro.flexify.ai/ This dataset is constructed from the U.S. Customs and Border Protection (CBP) Rulings Online Search System (CROSS).It contains rulings where importers sought clarification on the correct Harmonized… See the full description on the dataset page: https://huggingface.co/datasets/Dayanand314Krishna/cross_rulings_hts_dataset_for_tariffs.texttext-classification10K<n<100K0 likes170 downloads7mo agoHugging Face04tarudesu /ViHealthQA Disclaimer: The dataset may contain personal information crawled along with the contents of various sources. Please make a filter in pre-processing data before starting your research training. SPBERTQA: A Two-Stage Question Answering System Based on Sentence Transformers for Medical Texts This is the official repository for the ViHealthQA dataset from the paper SPBERTQA: A Two-Stage Question Answering System Based on Sentence Transformers for Medical Texts, which was… See the full description on the dataset page: https://huggingface.co/datasets/tarudesu/ViHealthQA.textquestion-answering10K<n<100K17 likes150 downloads3y agoHugging Face05Tarotoo /tarotoo-tarot-card-meanings Tarotoo Tarot Card Meanings A complete, structured dataset of all 78 tarot cards (22 Major Arcana + 56 Minor Arcana) in the Rider–Waite–Smith tradition. Published by Tarotoo. These are the card meanings that ground the AI-generated readings on Tarotoo.com. Dataset details Curated by: Tarotoo (tarotoo.com) Language: English License: MIT Rows: 78 (one per card) · Fields: 22 DOI (Zenodo, cite this): 10.5281/zenodo.21514483 Concept DOI (Zenodo, always resolves to the… See the full description on the dataset page: https://huggingface.co/datasets/Tarotoo/tarotoo-tarot-card-meanings.tabulartext-generationn<1K0 likes109 downloads1mo agoHugging Face06onkanat /turk-tarihi-1931-sft-dpo 🏛️ Türk Tarihi 1931 Ders Kitapları Sentetik Veri Seti (SFT / DPO / Chat) Bu veri seti, 1931 yılında Türkiye Cumhuriyeti Maarif Vekaleti (Milli Eğitim Bakanlığı) tarafından Türk Tarih Tetkik Cemiyeti'ne hazırlatılan ve Devlet Matbaası'nda basılan 4 ciltlik tarihi liseler için ders kitapları arşivinden (Tarih I: Tarihten Evvelki Zamanlar ve Eski Zamanlar, Tarih II: Orta Zamanlar, Tarih III: Yeni ve Yakın Zamanlar, Tarih IV: Türkiye Cumhuriyeti) otomatize hatlar üzerinden… See the full description on the dataset page: https://huggingface.co/datasets/onkanat/turk-tarihi-1931-sft-dpo.texttext-generation10K<n<100K0 likes94 downloads2mo agoHugging Face07emre /TARA_Turkish_LLM_Benchmark TARA: Turkish Advanced Reasoning Assessment Veri Seti *Img Credit: Open AI ChatGPT **English version is given below.** Evaluation Notebook / Değerlendirme Not Defteri Dataset Summary TARA (Turkish Advanced Reasoning Assessment), Türkçe dilindeki Büyük Dil Modellerinin (LLM'ler) gelişmiş akıl yürütme yeteneklerini çoklu alanlarda ölçmek için tasarlanmış, zorluk derecesine göre sınıflandırılmış bir benchmark veri setidir. Bu veri seti, LLM'lerin sadece bilgi… See the full description on the dataset page: https://huggingface.co/datasets/emre/TARA_Turkish_LLM_Benchmark.textquestion-answeringn<1K28 likes90 downloads1y agoHugging Face08tarekys5 /egyptian_legal_v2 Egyptian Legal QA Dataset (v2) - IRAC Formatted Overview This dataset is a high-quality, structured collection of Egyptian Legal Question & Answer pairs. It is specifically designed for fine-tuning Large Language Models (LLMs) and building Retrieval-Augmented Generation (RAG) systems specialized in Egyptian jurisprudence. hat's New in v2? Unlike the first version (egyptian_legal_v1), which was limited exclusively to Constitutional data, this second version (v2)… See the full description on the dataset page: https://huggingface.co/datasets/tarekys5/egyptian_legal_v2.textquestion-answering10K<n<100K6 likes76 downloads7mo agoHugging Face09tarekmasryo /rag-qa-logs-corpus-data 🧠📚 RAG QA Logs & Corpus (Synthetic) 🧪 Multi-table synthetic RAG telemetry for quality, hallucinations, latency, and cost A production-style, privacy-safe synthetic dataset that mimics telemetry exported from a real RAG system — from corpus → chunks → retrieval events → eval runs. ✅ Fully synthetic (no real users / orgs / PII). ⚡ Quick facts Total rows: 103,255 across 6 linked tables Labels (in eval_runs): is_correct, hallucination_flag, faithfulness_label… See the full description on the dataset page: https://huggingface.co/datasets/tarekmasryo/rag-qa-logs-corpus-data.tabularquestion-answering100K<n<1M2 likes73 downloads8mo agoHugging Face10Xerv-AI /TART 🧠 Xerv-AI/TART (Textual Answers & Reasoning Traces) Welcome to the ultimate repository for TART (Textual Answers & Reasoning Traces). This dataset is a hyper-robust, massively aggregated, and meticulously filtered corpus of ~344,000 instruction-tuning records. It is expressly engineered for Supervised Fine-Tuning (SFT) and the distillation of advanced Chain-of-Thought (CoT) reasoning capabilities into open-source Large Language Models. Designed with a heavy emphasis on… See the full description on the dataset page: https://huggingface.co/datasets/Xerv-AI/TART.texttext-generation100K<n<1M1 likes67 downloads5mo agoHugging Face11Dendory /tarotThis is a dataset of 5,770 high quality tarot cards readings produced by ChatGPT based on 3 randomly drawn cards. It can be used to train smaller models for use in a tarot application. The prompt used to produce these readings was: Give me a one paragraph tarot reading if I pull the cards CARD1, CARD2 and CARD3.\n\nReading:\n The CSV dataset contains the following columns: Card 1, Card 2, Card 3, Reading There are also 2 Python scripts included: make_dataset.py: This file was used to create… See the full description on the dataset page: https://huggingface.co/datasets/Dendory/tarot.question-answering16 likes63 downloads3y agoHugging Face12Salama1429 /tarteel-ai-QuranQA Dataset Card for the Qur'anic Reading Comprehension Dataset (QRCD) Dataset Summary The QRCD (Qur'anic Reading Comprehension Dataset) is composed of 1,093 tuples of question-passage pairs that are coupled with their extracted answers to constitute 1,337 question-passage-answer triplets. Supported Tasks and Leaderboards This task is evaluated as a ranking task. To give credit to a QA system that may retrieve an answer (not necessarily at the first rank) that… See the full description on the dataset page: https://huggingface.co/datasets/Salama1429/tarteel-ai-QuranQA.textquestion-answering1K<n<10K2 likes55 downloads2y agoHugging Face13lablup /tariff_trade_domain.synthetic_trade_qa_kr Korean Trade Domain QA Dataset A Korean-language question-answering dataset for the international trade domain, combining 21,399 QA pairs from three complementary sources: official 무역영어 1급 certification exam questions, trade terminology definitions, and lecture-derived QA pairs. Designed for fine-tuning and evaluating LLMs on Korean trade domain knowledge. Dataset Description The dataset covers vocabulary, regulations, procedures, and concepts from Korean… See the full description on the dataset page: https://huggingface.co/datasets/lablup/tariff_trade_domain.synthetic_trade_qa_kr.textquestion-answering10K<n<100K1 likes48 downloads4mo agoHugging Face14taresco /piqa_yoruba_pidgin Physical Commonsense Reasoning for Yorùbá and Nigerian Pidgin Dataset Summary This dataset was developed for the MRL 2025 Shared Task on Multilingual Physical Reasoning. For more details, see Global PIQA: Evaluating Physical Commonsense Reasoning Across 100+ Languages and Cultures. It provides a test collection for evaluating physical commonsense reasoning, that is, a model's ability to understand how objects, actions, and outcomes relate in everyday scenarios. The… See the full description on the dataset page: https://huggingface.co/datasets/taresco/piqa_yoruba_pidgin.textquestion-answeringn<1K2 likes46 downloads10mo agoHugging Face15tartuNLP /smugri-copagated SMUGRI-COPA SMUGRI-COPA is a manually translated commonsense causal reasoning benchmark for Võro and Livonian, two heavily under-resourced Finnic languages. It extends COPA and follows the XCOPA structure, enabling comparison with existing XCOPA translations. For each language, the dataset contains a 500-example test set and a 100-example validation set. The XCOPA split assignments, labels, and option order are preserved. Access and benchmark preservation… See the full description on the dataset page: https://huggingface.co/datasets/tartuNLP/smugri-copa.tabularquestion-answering1K<n<10K0 likes41 downloads8d agoHugging Face16tarteel-ai /quranqaThe absence of publicly available reusable test collections for Arabic question answering on the Holy Qur’an has impeded the possibility of fairly comparing the performance of systems in that domain. In this article, we introduce AyaTEC, a reusable test collection for verse-based question answering on the Holy Qur’an, which serves as a common experimental testbed for this task. AyaTEC includes 207 questions (with their corresponding 1,762 answers) covering 11 topic categories of the Holy Qur’an that target the information needs of both curious and skeptical users. To the best of our effort, the answers to the questions (each represented as a sequence of verses) in AyaTEC were exhaustive—that is, all qur’anic verses that directly answered the questions were exhaustively extracted and annotated. To facilitate the use of AyaTEC in evaluating the systems designed for that task, we propose several evaluation measures to support the different types of questions and the nature of verse-based answers while integrating the concept of partial matching of answers in the evaluation.question-answeringn<1K17 likes39 downloads2y agoHugging Face17taruntr /SurgWound Dataset Card for SurgWound SurgWound is the first open-source dataset for surgical wound analysis across multiple procedure types. SurgWound comprises 697 surgical wound images, each annotated by surgical experts at The Ohio State University Wexner Medical Center (OSWUMC). Each image is accompanied by high-quality labels covering six surgical wound characteristic attributes and two diagnostic outcomes attributes. SurgWound-Bench is the first multimodal benchmark for surgical wound… See the full description on the dataset page: https://huggingface.co/datasets/taruntr/SurgWound.textquestion-answering1K<n<10K0 likes38 downloads8mo agoHugging Face18tarzan19990815 /x_dataset_47 Bittensor Subnet 13 X (Twitter) Dataset Miner Data Compliance Agreement In uploading this dataset, I am agreeing to the Macrocosmos Miner Data Compliance Policy. Dataset Summary This dataset is part of the Bittensor Subnet 13 decentralized network, containing preprocessed data from X (formerly Twitter). The data is continuously updated by network miners, providing a real-time stream of tweets for various analytical and machine learning… See the full description on the dataset page: https://huggingface.co/datasets/tarzan19990815/x_dataset_47.texttext-classification10K<n<100K0 likes35 downloads1y agoHugging Face19yusufttogrul /tarihsample MEB AYT Tarih Soru Veri Seti Veri Seti Özeti (Dataset Summary) Bu veri seti, T.C. Millî Eğitim Bakanlığı (MEB) tarafından hazırlanan "Dört Dörtlük Konu Pekiştirme Testleri" içerisindeki AYT Tarih sorularından oluşturulmuştur. Özellikle Türkçe Doğal Dil İşleme (NLP) projelerinde, Soru-Cevap (Question Answering), Çoktan Seçmeli (Multiple Choice) model eğitimleri ve RAG (Retrieval-Augmented Generation) sistemleri için yüksek kaliteli, çözümlü ve akademik bir Türkçe veri… See the full description on the dataset page: https://huggingface.co/datasets/yusufttogrul/tarihsample.textquestion-answeringn<1K0 likes35 downloads7mo agoHugging Face20tartuNLP /belebele-smugri Finno-Ugric Belebele (Belebele-SMUGRI) Subset of Belebele translated to three low-resource Finno-Ugric languages: Komi, Võro, Livonian. The dataset reuses translations from SMUGRI-FLORES (first 250 sentences from FLORES devtest) for the text passages. Citation @inproceedings{purason-etal-2025-llms, title = "{LLM}s for Extremely Low-Resource {F}inno-{U}gric Languages", author = "Purason, Taido and Kuulmets, Hele-Andra and Fishel, Mark", editor =… See the full description on the dataset page: https://huggingface.co/datasets/tartuNLP/belebele-smugri.textquestion-answeringn<1K0 likes33 downloads4mo agoHugging Face21emre /El-TARA_Spanish_LLM_Benchmark El-Tara: Evaluación de Razonamiento Avanzado en Español Dataset Summary El-Tara (Evaluación de Razonamiento Avanzado en Español) is a benchmark dataset designed to assess the advanced reasoning capabilities of Large Language Models (LLMs) in Spanish. It is adapted from the original TARA (Turkish Advanced Reasoning Assessment) dataset. Similar to TARA, El-Tara aims to test higher-order cognitive skills across multiple domains, using synthetically generated questions… See the full description on the dataset page: https://huggingface.co/datasets/emre/El-TARA_Spanish_LLM_Benchmark.textquestion-answeringn<1K1 likes31 downloads1y agoHugging Face22tarzan19990815 /reddit_dataset_47 Bittensor Subnet 13 Reddit Dataset Miner Data Compliance Agreement In uploading this dataset, I am agreeing to the Macrocosmos Miner Data Compliance Policy. Dataset Summary This dataset is part of the Bittensor Subnet 13 decentralized network, containing preprocessed Reddit data. The data is continuously updated by network miners, providing a real-time stream of Reddit content for various analytical and machine learning tasks. For more… See the full description on the dataset page: https://huggingface.co/datasets/tarzan19990815/reddit_dataset_47.texttext-classification10K<n<100K0 likes26 downloads1y agoHugging Face23tartuNLP /EstCOPA Estonian Choice of Plausible Alternatives (EstCOPA) Dataset Summary EstCOPA is an extended version of XCOPA that was created with a goal to further investigate Estonian language understanding of large language models. EstCOPA provides two new versions of train, eval and test datasets in Estonian: firstly, a machine translated (En->Et) version of original English COPA (Roemmele et al., 2011) and secondly, a manually post-edited version of the same machine translated data.… See the full description on the dataset page: https://huggingface.co/datasets/tartuNLP/EstCOPA.tabularquestion-answering1K<n<10K2 likes22 downloads10mo agoHugging Face24tardelr /worldcup2026 World Cup 2026 Q&A 952 question-and-answer pairs about the 2026 FIFA World Cup, data was retrieved from Openfootbal github Dataset Summary This is a very simple adjustment from the original data adapted to Q&A format using LLMs (you should double-check results for any serious application). Dataset Structure Each line is a JSON object with a messages field: { "messages": [ {"role": "system", "content": "You are a helpful assistant with… See the full description on the dataset page: https://huggingface.co/datasets/tardelr/worldcup2026.textquestion-answeringn<1K0 likes16 downloads9h agoHugging Face25Clouds4days /tarotoo-tarot-card-meanings Tarotoo Tarot Card Meanings A complete, structured dataset of all 78 tarot cards (22 Major Arcana + 56 Minor Arcana) in the Rider–Waite–Smith tradition. Published by Tarotoo. These are the card meanings that ground the AI-generated readings on Tarotoo.com. Dataset details Curated by: Tarotoo (tarotoo.com) Language: English License: MIT Rows: 78 (one per card) · Fields: 22 DOI (Zenodo, cite this): 10.5281/zenodo.21514483 Concept DOI (Zenodo, always resolves to the… See the full description on the dataset page: https://huggingface.co/datasets/Clouds4days/tarotoo-tarot-card-meanings.tabulartext-generationn<1K0 likes13 downloads2mo agoHugging Face26zastixx /bank_test_tarun Dataset Card for bank_test_tarun This dataset has been created with distilabel. Dataset Summary This dataset contains a pipeline.yaml which can be used to reproduce the pipeline that generated it in distilabel using the distilabel CLI: distilabel pipeline run --config "https://huggingface.co/datasets/zastixx/bank_test_tarun/raw/main/pipeline.yaml" or explore the configuration: distilabel pipeline info --config… See the full description on the dataset page: https://huggingface.co/datasets/zastixx/bank_test_tarun.texttext-generationn<1K0 likes11 downloads2y agoHugging Face27tarunkmr566 /Open-RL Open-RL Dataset Summary This dataset contains self-contained, verifiable, and unambiguous STEM reasoning problems across Physics, Mathematics, Biology, and Chemistry. Each problem: Requires multi-step reasoning Involves symbolic manipulation and/or numerical computation Has a deterministic, objectively verifiable final answer The problems were evaluated against contemporary large language models. Observed pass rates indicate that the tasks are non-trivial yet… See the full description on the dataset page: https://huggingface.co/datasets/tarunkmr566/Open-RL.textquestion-answeringn<1K0 likes8 downloads6mo agoHugging Face28tarzan19990815 /x_dataset_225 Bittensor Subnet 13 X (Twitter) Dataset Miner Data Compliance Agreement In uploading this dataset, I am agreeing to the Macrocosmos Miner Data Compliance Policy. Dataset Summary This dataset is part of the Bittensor Subnet 13 decentralized network, containing preprocessed data from X (formerly Twitter). The data is continuously updated by network miners, providing a real-time stream of tweets for various analytical and machine learning… See the full description on the dataset page: https://huggingface.co/datasets/tarzan19990815/x_dataset_225.texttext-classification1K<n<10K0 likes7 downloads1y agoHugging Face29tarik645 /my-distiset-561563c9 Dataset Card for my-distiset-561563c9 This dataset has been created with distilabel. Dataset Summary This dataset contains a pipeline.yaml which can be used to reproduce the pipeline that generated it in distilabel using the distilabel CLI: distilabel pipeline run --config "https://huggingface.co/datasets/tarik645/my-distiset-561563c9/raw/main/pipeline.yaml" or explore the configuration: distilabel pipeline info --config… See the full description on the dataset page: https://huggingface.co/datasets/tarik645/my-distiset-561563c9.texttext-generationn<1K0 likes6 downloads2y agoHugging Face30Tarakanta /text-2-sql_datasettextquestion-answeringn<1K0 likes6 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.