datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
RAGDOLL
The RAGDOLL E-Commerce Webpage Dataset
This repository contains the RAGDOLL (Retrieval-Augmented Generation Deceived Ordering via AdversariaL materiaLs) dataset as well as its LLM-automated collection pipeline.
The RAGDOLL dataset is from the paper Ranking Manipulation for Conversational Search Engines from Samuel Pfrommer, Yatong Bai, Tanmay Gautam, and Somayeh Sojoudi. For experiment code associated with this paper, please refer to this repository.
The dataset consists of 10… See the full description on the dataset page: https://huggingface.co/datasets/Bai-YT/RAGDOLL.SVBench
Dataset Card for SVBench
This dataset card aims to provide a comprehensive overview of the SVBench dataset, including its purpose, structure, and sources. For details, see our Project, Paper and GitHub repository.
Dataset Details
Dataset Description
SVBench is the first benchmark specifically designed to evaluate long-context streaming video understanding through temporal multi-turn question-answering (QA) chains. It addresses the limitations of existing video… See the full description on the dataset page: https://huggingface.co/datasets/yzy666/SVBench.LongDA
LongDA Dataset Card
Dataset Description
LongDA is a data analysis benchmark for evaluating LLM-based agents under documentation-intensive analytical workflows. It features authentic U.S. government survey data with complete, long documentation, testing LLMs' ability to navigate complex real-world datasets before performing analysis.
Dataset Summary
505 queries extracted from 30 expert-written publications
17 U.S. national surveys covering health… See the full description on the dataset page: https://huggingface.co/datasets/Yiyang-Ian-Li/LongDA.SNAP
SNAP Benchmark
Code and annotations: [https://github.com/ykotseruba/SNAP]
SNAP (stands for Shutter speed, ISO seNsitivity, and APerture) is a new benchmark consisting of images of objects taken under controlled lighting conditions and with densely sampled camera settings.
This benchmark allows testing the effects of capture bias, which includes camera settings and illumination, on performance of vision algorithms.
SNAP contains 37,558 images of 100 scenes (10 scenes per 10 object… See the full description on the dataset page: https://huggingface.co/datasets/ykotseruba/SNAP.photonic-integrated-circuit-yield
🏭 Photonic Integrated Circuit Yield Dataset
📊 125,000 synthetic (yield query, yield reasoning response) pairs covering process variation, defect density, lithography, and metrology challenges in CMOS-compatible photonic integrated circuit (PIC) manufacturing.
⚠️ Disclaimer: All entries are synthetically generated. Yield figures are computed from textbook models over sampled inputs, and citations are placeholders styled after technical sources; none reference a real… See the full description on the dataset page: https://huggingface.co/datasets/Taylor658/photonic-integrated-circuit-yield.econ_logic_qa
EconLogicQA
EconLogicQA is a benchmark designed to test the sequential reasoning skills of large language models (LLMs) in economics, business,
and supply chain management. It diverges from typical benchmarks by requiring models to understand and sequence multiple interconnected
events, capturing complex economic logics. The benchmark includes multi-event scenarios and a thorough suite of evaluations to assess
proficiency in economic contexts.
Dataset Details… See the full description on the dataset page: https://huggingface.co/datasets/yinzhu-quan/econ_logic_qa.stock_trading_QARadFig-VQA
RadFig-VQA Dataset
Overview
RadFig-VQA is a large-scale medical visual question answering dataset based on radiological figures from PubMed Central (PMC). The dataset comprises 70,550 images and 238,294 question-answer pairs - making it the largest radiology-specific VQA dataset by number of QA pairs - generated from radiological figures across diverse imaging modalities and clinical contexts, designed for comprehensive medical VQA evaluation.
Dataset… See the full description on the dataset page: https://huggingface.co/datasets/YYama0/RadFig-VQA.piqa_yoruba_pidgin
Physical Commonsense Reasoning for Yorùbá and Nigerian Pidgin
Dataset Summary
This dataset was developed for the MRL 2025 Shared Task on Multilingual Physical Reasoning. For more details, see Global PIQA: Evaluating Physical Commonsense Reasoning Across 100+ Languages and Cultures.
It provides a test collection for evaluating physical commonsense reasoning, that is, a model's ability to understand how objects, actions, and outcomes relate in everyday scenarios.
The… See the full description on the dataset page: https://huggingface.co/datasets/taresco/piqa_yoruba_pidgin.gpqa-extended_trThis dataset is the Turkish translation of GPQA dataset(extended), which is one of the most-known benchmarks to evaluate LLMs performance on graduate level science questions.
You can find the paper of GPQA below.
https://arxiv.org/abs/2311.12022
Translation Methodology:
A subset of the columns of every row translated seperately using "google/gemma-3-27b-it" which has a good capacity for both the languages English and Turkish and performed good in scientific translation.
The following prompt… See the full description on the dataset page: https://huggingface.co/datasets/ytu-ce-cosmos/gpqa-extended_tr.indian-government-schemes-2025
Indian Government Schemes Dataset 2026
Dataset Description
The most comprehensive structured dataset of Indian central and state government schemes — 4,693 schemes across all ministries and states, with machine-readable eligibility fields.
Maintained by SmartDuke Technologies · Coimbatore, Tamil Nadu, India
This dataset powers SchemeFit — India's government scheme finder for citizens and businesses.
What Makes This Different
Most existing Indian… See the full description on the dataset page: https://huggingface.co/datasets/Yokey20/indian-government-schemes-2025.MMMLU
Multilingual Massive Multitask Language Understanding (MMMLU)
The MMLU is a widely recognized benchmark of general knowledge attained by AI models. It covers a broad range of topics from 57 different categories, covering elementary-level knowledge up to advanced professional subjects like law, physics, history, and computer science.
We translated the MMLU’s test set into 14 languages using professional human translators. Relying on human translators for this evaluation increases… See the full description on the dataset page: https://huggingface.co/datasets/Yashugowda20/MMMLU.turkish-university-mevzuat
Turkey University Regulation Data Collection
This dataset provides a comprehensive collection of regulatory documents of Turkish universities obtained from mevzuat.gov.tr.
It includes full texts of regulations with detailed publication information and unique identifiers.
Overview
Data Sources: mevzuat.gov.tr website
Technologies Used: Selenium, BeautifulSoup, Python
Data Formats: CSV
CSV Data Structure
Column
Description
Üniversite
Name of the… See the full description on the dataset page: https://huggingface.co/datasets/yusufbaykaloglu/turkish-university-mevzuat.ScienceOlympiad.tsv
Dataset Card for ScienceOlympiad.tsv
ScienceOlympiad.tsv: Challenging AI with Olympiad-Level Multimodal Science Problems.
Source: https://huggingface.co/datasets/ByteDance-Seed/ScienceOlympiad
Dataset Details
Dataset Description
The ScienceOlympiad dataset is a meticulously curated benchmark designed to evaluate the scientific reasoning capabilities of state-of-the-art AI models. It features elite, competition-level problems in physics and chemistry, addressing… See the full description on the dataset page: https://huggingface.co/datasets/YuJJJJin/ScienceOlympiad.tsv.OR-Space
OR-Space
A full-lifecycle workspace benchmark for industrial optimization agents.
OR-Space evaluates whether LLM agents can do reliable operations research work
inside executable, multi-file workspaces. Each instance keeps business
requirements, parameter files, source code, solver artifacts, and evaluation
metadata as separate files, forcing the agent to recover and maintain the
optimization model through workspace interaction rather than one-shot text
generation.… See the full description on the dataset page: https://huggingface.co/datasets/YiYao7017/OR-Space.Medical-Consultation-Questions-in-Arabic
Dataset Card for Dataset Name
Dataset Details
This dataset contains 47,705 Arabic medical questions collected from the Arabic health platform Altibbi. Each question is categorized into a medical domain such as sexual health, dermatology, pediatrics, and more.
The dataset can be used for Natural Language Processing (NLP) tasks such as:
"Text classification (predicting medical categories)".
"Question answering systems in Arabic".
"Building Arabic healthcare chatbots".… See the full description on the dataset page: https://huggingface.co/datasets/Youssefx64/Medical-Consultation-Questions-in-Arabic.EmoSupportBench
EmoSupportBench
EmoSupportBench is a comprehensive dataset and benchmark for evaluating emotional support capabilities of large language models (LLMs). It provides a systematic framework to assess how well AI systems can provide empathetic, helpful, and psychologically-grounded support to users seeking emotional assistance.
🎯 Key Features
200-question bilingual evaluation set (English & Chinese) covering 8 major emotional support scenarios
Hierarchical scenario… See the full description on the dataset page: https://huggingface.co/datasets/YueyangWang/EmoSupportBench.PedMedQA
PedMedQA: Evaluating Large Language Models in Pediatrics and Adult Medicine
Overview
PedMedQA is an openly accessible pediatric-specific benchmark for evaluating the performance of large language models (LLMs) in pediatric scenarios. It is curated from the widely used MedQA benchmark and allows for population-specific assessments by focusing on multiple-choice questions (MCQs) relevant to pediatrics.
Dataset Details
Pediatric-Specific Dataset: PedMedQA… See the full description on the dataset page: https://huggingface.co/datasets/yma94/PedMedQA.NLP_Insights_2023_2024
Key Insignts from NLP papers (2023 -2024)
This dataset is processed and compiled by @hu_yifei as part of open-source effort from the Open Research Assistant Project.
It includes key insights extracted from top tier NLP conference papers:
Year
Venue
Paper Count
2024
eacl
225
2024
naacl
564
2023
acl
1077
2023
conll
41
2023
eacl
281
2023
emnlp
1048
2023
semeval
319
2023
wmt
101
Dataset Stats
Total number of papers: 3,640
Total rows (key… See the full description on the dataset page: https://huggingface.co/datasets/yifeihu/NLP_Insights_2023_2024.bert-dataset
Road Traffic Act QA Dataset
This dataset is automatically generated question-answer pairs based on the official Road Traffic Act (Republic of Korea). The dataset is designed to support RAG (Retrieval-Augmented Generation) and legal NLP tasks.
Dataset Summary
Source: Road Traffic Act (English version)
Task: Question Answering (QA)
Type: Automatically generated by GPT-4o with custom multi-QA prompt
Size: 2,000+ QA pairs
Language: English
Format: CSV (Question, Answer)… See the full description on the dataset page: https://huggingface.co/datasets/YeahOuts/bert-dataset.simpleqa-verified
SimpleQA Verified
A 1,000-prompt factuality benchmark from Google DeepMind and Google Research, designed to reliably evaluate LLM parametric knowledge.
▶ SimpleQA Verified Leaderboard on Kaggle▶ Technical Report▶ Evaluation Starter Code
Benchmark
SimpleQA Verified is a 1,000-prompt benchmark for reliably evaluating Large Language Models (LLMs) on short-form factuality
and parametric knowledge. The authors from Google DeepMind and Google Research build on… See the full description on the dataset page: https://huggingface.co/datasets/yxx94/simpleqa-verified.tombench_merged
TomBench Merged Dataset (Exact Matching)
This dataset contains the merged results of TomBench evaluation with the original TomBench dataset, using exact string matching.
Dataset Statistics
Total records: 2860
Exact matches: 2860
Manual matches: 0
Average model score: 0.5066
Matching Strategy
This version uses exact string matching after text normalization:
Remove extra whitespace and normalize formatting
Match stories exactly between datasets
Report any… See the full description on the dataset page: https://huggingface.co/datasets/ycfNTU/tombench_merged.turkce-atasozleri-coktan-secmeli-benchmark
Türkçe Atasözleri Çoktan Seçmeli Benchmark
Türkçe atasözlerinin anlamını ölçmek amacıyla hazırlanmış 100 soruluk çoktan seçmeli bir benchmark veri setidir.
Her soruda dört seçenek bulunur. Doğru cevaplar dengeli dağıtılmıştır:
A: 25
B: 25
C: 25
D: 25
Rastgele tahmin seviyesi: %25
Sorular, YusufSimsek/turkce-atasozleri-dataset veri setindeki 100 benzersiz atasözünden türetilmiştir.
Veri alanları
Alan
Açıklama
id
Sorunun benzersiz kimliği
proverb… See the full description on the dataset page: https://huggingface.co/datasets/YusufSimsek/turkce-atasozleri-coktan-secmeli-benchmark.pittsburgh_floods_street_levelThe dataset provides fine-grained spatiotemporal information on urban floods occurring inside the city of Pittsburgh, PA, USA, from 2015 to 2024 by integrating publicly available data sources. The data sources include NOAA storm events database and Pittsburgh 311 flooding requests. Each row corresponds to one segment flood event, characterized by the street segment defined by a distinct combination of "u_node", "v_node", "length_m", and time. Each flood event was mapped to the street segments… See the full description on the dataset page: https://huggingface.co/datasets/yueq92/pittsburgh_floods_street_level.customer-support-tickets
Featuring Labeled Customer Emails and Support Responses
🔧 Synthetic IT Ticket Generator — Custom Dataset
Create a dataset tailored to your own queues & priorities (no PII).
👉 Generate custom data
Define your queues, priorities, language
Need an on-prem AI to auto-classify tickets?→ Open Ticket AI
There are 2 Versions of the dataset, the new version has more tickets, but only languages english and german. So please look at both files, to find what best fits… See the full description on the dataset page: https://huggingface.co/datasets/yadavmana/customer-support-tickets.create_qa_news
질문 생성: kullm3 모델 이용
답변 생성: GPT3.5 turbo API 이용
지문 원본: AI HUB 뉴스 기계독해 데이터셋
east_java_dialect_instruct
Complaints From The East Javanese Dialect community
This dataset created manually by humans with reference to public complaints in the comments column of the local government's Instagram account and another platform like X and TikTok Comments.
Turkish-STEM-DPO-Dataset
Turkish STEM DPO Dataset
Dataset Summary
The Turkish STEM DPO (Direct Preference Optimization) dataset is a comprehensive synthetic resource containing 16,177 high-quality preference pairs designed to enhance the reasoning capabilities of Turkish language models in mathematics, physics, and programming.
The dataset leverages a preference-based learning approach: each instance pairs a carefully crafted, expert-level solution with a deliberately flawed or incomplete… See the full description on the dataset page: https://huggingface.co/datasets/yusufbaykaloglu/Turkish-STEM-DPO-Dataset.VivaBench
VivaBench: Simulating Viva Voce Examinations to Evaluate Clinical Reasoning in LLMs
This repository is the official implementation of VivaBench—“Simulating Viva Voce Examinations to Evaluate Clinical Reasoning in Large Language Models.”
VivaBench is a multi-turn benchmark of 1,152 physician-curated clinical vignettes that simulates a viva voce (oral) exam: agents must iteratively gather H&P findings and order investigations to arrive at a diagnosis.
📋 Requirements… See the full description on the dataset page: https://huggingface.co/datasets/yiihan/VivaBench.yuvam
