datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Target-QA
🎯 Target-QA: The First QA Dataset Benchmarking Target Priorization Based on DepMap
📑 Dataset Summary
Target-QA is derived from the DepMap multi-omics and CRISPR screening cohorts, harmonized via BioMedGraphica.It enables multi-modal reasoning by combining numeric evidence, topological knowledge and language context for CRISPR target prioritization.
This dataset supports the training and benchmarking of… See the full description on the dataset page: https://huggingface.co/datasets/FuhaiLiAiLab/Target-QA.TLPC
Targoman Large Persian Corpus
Dataset Summary
Ever since the invention of the computer, humans have always been interested in being able to communicate with computers in human language. The research and endeavors of scientists and engineers brought us to the knowledge of human language processing and large language models. Large Language Models (LLMs) with their amazing capabilities, have caused a revolution in the artificial intelligence industry, which has caused a… See the full description on the dataset page: https://huggingface.co/datasets/Targoman/TLPC.cross_rulings_hts_dataset_for_tariffs
CROSS Rulings HTS Dataset for Tariff Classification
Maintained by Flexify.AI Inc. as part of the ATLAS trade intelligence research program.
Paper: ATLAS: Benchmarking and Adapting LLMs for Global Trade via Harmonized Tariff Code Classification
Project Page: https://tariffpro.flexify.ai/
This dataset is constructed from the U.S. Customs and Border Protection (CBP) Rulings Online Search System (CROSS).It contains rulings where importers sought clarification on the correct Harmonized… See the full description on the dataset page: https://huggingface.co/datasets/Dayanand314Krishna/cross_rulings_hts_dataset_for_tariffs.ViHealthQA
Disclaimer:
The dataset may contain personal information crawled along with the contents of various sources. Please make a filter in pre-processing data before starting your research training.
SPBERTQA: A Two-Stage Question Answering System Based on Sentence Transformers for Medical Texts
This is the official repository for the ViHealthQA dataset from the paper SPBERTQA: A Two-Stage Question Answering System Based on Sentence Transformers for Medical Texts, which was… See the full description on the dataset page: https://huggingface.co/datasets/tarudesu/ViHealthQA.tarotoo-tarot-card-meanings
Tarotoo Tarot Card Meanings
A complete, structured dataset of all 78 tarot cards (22 Major Arcana + 56 Minor Arcana) in the Rider–Waite–Smith tradition. Published by Tarotoo. These are the card meanings that ground the AI-generated readings on Tarotoo.com.
Dataset details
Curated by: Tarotoo (tarotoo.com)
Language: English
License: MIT
Rows: 78 (one per card) · Fields: 22
DOI (Zenodo, cite this): 10.5281/zenodo.21514483
Concept DOI (Zenodo, always resolves to the… See the full description on the dataset page: https://huggingface.co/datasets/Tarotoo/tarotoo-tarot-card-meanings.turk-tarihi-1931-sft-dpo
🏛️ Türk Tarihi 1931 Ders Kitapları Sentetik Veri Seti (SFT / DPO / Chat)
Bu veri seti, 1931 yılında Türkiye Cumhuriyeti Maarif Vekaleti (Milli Eğitim Bakanlığı) tarafından Türk Tarih Tetkik Cemiyeti'ne hazırlatılan ve Devlet Matbaası'nda basılan 4 ciltlik tarihi liseler için ders kitapları arşivinden (Tarih I: Tarihten Evvelki Zamanlar ve Eski Zamanlar, Tarih II: Orta Zamanlar, Tarih III: Yeni ve Yakın Zamanlar, Tarih IV: Türkiye Cumhuriyeti) otomatize hatlar üzerinden… See the full description on the dataset page: https://huggingface.co/datasets/onkanat/turk-tarihi-1931-sft-dpo.TARA_Turkish_LLM_Benchmark
TARA: Turkish Advanced Reasoning Assessment Veri Seti
*Img Credit: Open AI ChatGPT
**English version is given below.**
Evaluation Notebook / Değerlendirme Not Defteri
Dataset Summary
TARA (Turkish Advanced Reasoning Assessment), Türkçe dilindeki Büyük Dil Modellerinin (LLM'ler) gelişmiş akıl yürütme yeteneklerini çoklu alanlarda ölçmek için tasarlanmış, zorluk derecesine göre sınıflandırılmış bir benchmark veri setidir. Bu veri seti, LLM'lerin sadece bilgi… See the full description on the dataset page: https://huggingface.co/datasets/emre/TARA_Turkish_LLM_Benchmark.egyptian_legal_v2
Egyptian Legal QA Dataset (v2) - IRAC Formatted
Overview
This dataset is a high-quality, structured collection of Egyptian Legal Question & Answer pairs. It is specifically designed for fine-tuning Large Language Models (LLMs) and building Retrieval-Augmented Generation (RAG) systems specialized in Egyptian jurisprudence.
hat's New in v2?
Unlike the first version (egyptian_legal_v1), which was limited exclusively to Constitutional data, this second version (v2)… See the full description on the dataset page: https://huggingface.co/datasets/tarekys5/egyptian_legal_v2.rag-qa-logs-corpus-data
🧠📚 RAG QA Logs & Corpus (Synthetic)
🧪 Multi-table synthetic RAG telemetry for quality, hallucinations, latency, and cost
A production-style, privacy-safe synthetic dataset that mimics telemetry exported from a real RAG system — from corpus → chunks → retrieval events → eval runs.
✅ Fully synthetic (no real users / orgs / PII).
⚡ Quick facts
Total rows: 103,255 across 6 linked tables
Labels (in eval_runs): is_correct, hallucination_flag, faithfulness_label… See the full description on the dataset page: https://huggingface.co/datasets/tarekmasryo/rag-qa-logs-corpus-data.TART
🧠 Xerv-AI/TART (Textual Answers & Reasoning Traces)
Welcome to the ultimate repository for TART (Textual Answers & Reasoning Traces). This dataset is a hyper-robust, massively aggregated, and meticulously filtered corpus of ~344,000 instruction-tuning records. It is expressly engineered for Supervised Fine-Tuning (SFT) and the distillation of advanced Chain-of-Thought (CoT) reasoning capabilities into open-source Large Language Models.
Designed with a heavy emphasis on… See the full description on the dataset page: https://huggingface.co/datasets/Xerv-AI/TART.tarotThis is a dataset of 5,770 high quality tarot cards readings produced by ChatGPT based on 3 randomly drawn cards. It can be used to train smaller models for use in a tarot application.
The prompt used to produce these readings was:
Give me a one paragraph tarot reading if I pull the cards CARD1, CARD2 and CARD3.\n\nReading:\n
The CSV dataset contains the following columns: Card 1, Card 2, Card 3, Reading
There are also 2 Python scripts included:
make_dataset.py: This file was used to create… See the full description on the dataset page: https://huggingface.co/datasets/Dendory/tarot.tarteel-ai-QuranQA
Dataset Card for the Qur'anic Reading Comprehension Dataset (QRCD)
Dataset Summary
The QRCD (Qur'anic Reading Comprehension Dataset) is composed of 1,093 tuples of question-passage pairs that are
coupled with their extracted answers to constitute 1,337 question-passage-answer triplets.
Supported Tasks and Leaderboards
This task is evaluated as a ranking task.
To give credit to a QA system that may retrieve an answer (not necessarily at the first rank) that… See the full description on the dataset page: https://huggingface.co/datasets/Salama1429/tarteel-ai-QuranQA.tariff_trade_domain.synthetic_trade_qa_kr
Korean Trade Domain QA Dataset
A Korean-language question-answering dataset for the international trade domain, combining 21,399 QA pairs from three complementary sources: official 무역영어 1급 certification exam questions, trade terminology definitions, and lecture-derived QA pairs. Designed for fine-tuning and evaluating LLMs on Korean trade domain knowledge.
Dataset Description
The dataset covers vocabulary, regulations, procedures, and concepts from Korean… See the full description on the dataset page: https://huggingface.co/datasets/lablup/tariff_trade_domain.synthetic_trade_qa_kr.piqa_yoruba_pidgin
Physical Commonsense Reasoning for Yorùbá and Nigerian Pidgin
Dataset Summary
This dataset was developed for the MRL 2025 Shared Task on Multilingual Physical Reasoning. For more details, see Global PIQA: Evaluating Physical Commonsense Reasoning Across 100+ Languages and Cultures.
It provides a test collection for evaluating physical commonsense reasoning, that is, a model's ability to understand how objects, actions, and outcomes relate in everyday scenarios.
The… See the full description on the dataset page: https://huggingface.co/datasets/taresco/piqa_yoruba_pidgin.smugri-copa
SMUGRI-COPA
SMUGRI-COPA is a manually translated commonsense causal reasoning benchmark for Võro and Livonian, two heavily under-resourced Finnic languages. It extends COPA and follows the XCOPA structure, enabling comparison with existing XCOPA translations.
For each language, the dataset contains a 500-example test set and a 100-example validation set. The XCOPA split assignments, labels, and option order are preserved.
Access and benchmark preservation… See the full description on the dataset page: https://huggingface.co/datasets/tartuNLP/smugri-copa.quranqaThe absence of publicly available reusable test collections for Arabic question answering on the Holy Qur’an has impeded the possibility of fairly comparing the performance of systems in that domain. In this article, we introduce AyaTEC, a reusable test collection for verse-based question answering on the Holy Qur’an, which serves as a common experimental testbed for this task. AyaTEC includes 207 questions (with their corresponding 1,762 answers) covering 11 topic categories of the Holy Qur’an that target the information needs of both curious and skeptical users. To the best of our effort, the answers to the questions (each represented as a sequence of verses) in AyaTEC were exhaustive—that is, all qur’anic verses that directly answered the questions were exhaustively extracted and annotated. To facilitate the use of AyaTEC in evaluating the systems designed for that task, we propose several evaluation measures to support the different types of questions and the nature of verse-based answers while integrating the concept of partial matching of answers in the evaluation.SurgWound
Dataset Card for SurgWound
SurgWound is the first open-source dataset for surgical wound analysis across multiple procedure types.
SurgWound comprises 697 surgical wound images, each annotated by surgical experts at The Ohio State University Wexner Medical Center (OSWUMC).
Each image is accompanied by high-quality labels covering six surgical wound characteristic attributes and two diagnostic outcomes attributes.
SurgWound-Bench is the first multimodal benchmark for surgical wound… See the full description on the dataset page: https://huggingface.co/datasets/taruntr/SurgWound.x_dataset_47
Bittensor Subnet 13 X (Twitter) Dataset
Miner Data Compliance Agreement
In uploading this dataset, I am agreeing to the Macrocosmos Miner Data Compliance Policy.
Dataset Summary
This dataset is part of the Bittensor Subnet 13 decentralized network, containing preprocessed data from X (formerly Twitter). The data is continuously updated by network miners, providing a real-time stream of tweets for various analytical and machine learning… See the full description on the dataset page: https://huggingface.co/datasets/tarzan19990815/x_dataset_47.tarihsample
MEB AYT Tarih Soru Veri Seti
Veri Seti Özeti (Dataset Summary)
Bu veri seti, T.C. Millî Eğitim Bakanlığı (MEB) tarafından hazırlanan "Dört Dörtlük Konu Pekiştirme Testleri" içerisindeki AYT Tarih sorularından oluşturulmuştur.
Özellikle Türkçe Doğal Dil İşleme (NLP) projelerinde, Soru-Cevap (Question Answering), Çoktan Seçmeli (Multiple Choice) model eğitimleri ve RAG (Retrieval-Augmented Generation) sistemleri için yüksek kaliteli, çözümlü ve akademik bir Türkçe veri… See the full description on the dataset page: https://huggingface.co/datasets/yusufttogrul/tarihsample.belebele-smugri
Finno-Ugric Belebele (Belebele-SMUGRI)
Subset of Belebele translated to three low-resource Finno-Ugric languages: Komi, Võro, Livonian.
The dataset reuses translations from SMUGRI-FLORES (first 250 sentences from FLORES devtest) for the text passages.
Citation
@inproceedings{purason-etal-2025-llms,
title = "{LLM}s for Extremely Low-Resource {F}inno-{U}gric Languages",
author = "Purason, Taido and
Kuulmets, Hele-Andra and
Fishel, Mark",
editor =… See the full description on the dataset page: https://huggingface.co/datasets/tartuNLP/belebele-smugri.El-TARA_Spanish_LLM_Benchmark
El-Tara: Evaluación de Razonamiento Avanzado en Español
Dataset Summary
El-Tara (Evaluación de Razonamiento Avanzado en Español) is a benchmark dataset designed to assess the advanced reasoning capabilities of Large Language Models (LLMs) in Spanish. It is adapted from the original TARA (Turkish Advanced Reasoning Assessment) dataset.
Similar to TARA, El-Tara aims to test higher-order cognitive skills across multiple domains, using synthetically generated questions… See the full description on the dataset page: https://huggingface.co/datasets/emre/El-TARA_Spanish_LLM_Benchmark.reddit_dataset_47
Bittensor Subnet 13 Reddit Dataset
Miner Data Compliance Agreement
In uploading this dataset, I am agreeing to the Macrocosmos Miner Data Compliance Policy.
Dataset Summary
This dataset is part of the Bittensor Subnet 13 decentralized network, containing preprocessed Reddit data. The data is continuously updated by network miners, providing a real-time stream of Reddit content for various analytical and machine learning tasks.
For more… See the full description on the dataset page: https://huggingface.co/datasets/tarzan19990815/reddit_dataset_47.EstCOPA
Estonian Choice of Plausible Alternatives (EstCOPA)
Dataset Summary
EstCOPA is an extended version of XCOPA that was created with a goal to further investigate Estonian language understanding of large language models. EstCOPA provides two new versions of train, eval and test datasets in Estonian: firstly, a machine translated (En->Et) version of original English COPA (Roemmele et al., 2011) and secondly, a manually post-edited version of the same machine translated data.… See the full description on the dataset page: https://huggingface.co/datasets/tartuNLP/EstCOPA.worldcup2026
World Cup 2026 Q&A
952 question-and-answer pairs about the 2026 FIFA World Cup, data was retrieved from Openfootbal github
Dataset Summary
This is a very simple adjustment from the original data adapted to Q&A format using LLMs (you should double-check results for any serious application).
Dataset Structure
Each line is a JSON object with a messages field:
{
"messages": [
{"role": "system", "content": "You are a helpful assistant with… See the full description on the dataset page: https://huggingface.co/datasets/tardelr/worldcup2026.tarotoo-tarot-card-meanings
Tarotoo Tarot Card Meanings
A complete, structured dataset of all 78 tarot cards (22 Major Arcana + 56 Minor Arcana) in the Rider–Waite–Smith tradition. Published by Tarotoo. These are the card meanings that ground the AI-generated readings on Tarotoo.com.
Dataset details
Curated by: Tarotoo (tarotoo.com)
Language: English
License: MIT
Rows: 78 (one per card) · Fields: 22
DOI (Zenodo, cite this): 10.5281/zenodo.21514483
Concept DOI (Zenodo, always resolves to the… See the full description on the dataset page: https://huggingface.co/datasets/Clouds4days/tarotoo-tarot-card-meanings.bank_test_tarun
Dataset Card for bank_test_tarun
This dataset has been created with distilabel.
Dataset Summary
This dataset contains a pipeline.yaml which can be used to reproduce the pipeline that generated it in distilabel using the distilabel CLI:
distilabel pipeline run --config "https://huggingface.co/datasets/zastixx/bank_test_tarun/raw/main/pipeline.yaml"
or explore the configuration:
distilabel pipeline info --config… See the full description on the dataset page: https://huggingface.co/datasets/zastixx/bank_test_tarun.Open-RL
Open-RL
Dataset Summary
This dataset contains self-contained, verifiable, and unambiguous STEM reasoning problems across Physics, Mathematics, Biology, and Chemistry.
Each problem:
Requires multi-step reasoning
Involves symbolic manipulation and/or numerical computation
Has a deterministic, objectively verifiable final answer
The problems were evaluated against contemporary large language models. Observed pass rates indicate that the tasks are non-trivial yet… See the full description on the dataset page: https://huggingface.co/datasets/tarunkmr566/Open-RL.x_dataset_225
Bittensor Subnet 13 X (Twitter) Dataset
Miner Data Compliance Agreement
In uploading this dataset, I am agreeing to the Macrocosmos Miner Data Compliance Policy.
Dataset Summary
This dataset is part of the Bittensor Subnet 13 decentralized network, containing preprocessed data from X (formerly Twitter). The data is continuously updated by network miners, providing a real-time stream of tweets for various analytical and machine learning… See the full description on the dataset page: https://huggingface.co/datasets/tarzan19990815/x_dataset_225.my-distiset-561563c9
Dataset Card for my-distiset-561563c9
This dataset has been created with distilabel.
Dataset Summary
This dataset contains a pipeline.yaml which can be used to reproduce the pipeline that generated it in distilabel using the distilabel CLI:
distilabel pipeline run --config "https://huggingface.co/datasets/tarik645/my-distiset-561563c9/raw/main/pipeline.yaml"
or explore the configuration:
distilabel pipeline info --config… See the full description on the dataset page: https://huggingface.co/datasets/tarik645/my-distiset-561563c9.text-2-sql_dataset
