datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
nuclear-intelligence-dataset
Nuclear Intelligence Dataset
Public, auto-generated dataset of validated nuclear-energy research cycles.
Latest stats (auto-updated):
🪙 NES tokens minted: 0
⛓️ Blockchain length: 1 blocks
🕸️ Knowledge entities: 2
Source
GitHub: https://github.com/QalamHipHop/nuclear-intelligence
HF Space: https://huggingface.co/spaces/Qalam/Nuclear-Intelligence
License
MIT
mimic-medical-imaging-qa
MIMIC Medical Imaging QA Dataset
5,207 Bloom's-taxonomy-stratified question--answer pairs derived from 23 medical imaging lectures (RPI BMED 2300). The dataset supports the paper "MIMIC: A Course-Derivation Pipeline and Benchmark for Slide-Anchored Tutoring with a Domain-Adapted Large Language Model" and was used to fine-tune MIMIC-LM, a domain-adapted Llama-3.1-8B-Instruct model for grounded medical imaging instruction.
License
The benchmark annotations, dataset… See the full description on the dataset page: https://huggingface.co/datasets/zabir1996/mimic-medical-imaging-qa.motif-qa
MotifQA
Dataset Summary
MotifQA is a synthetic graph question-answering benchmark focused on detecting graph motifs inside small random graphs.
Each example pairs a textual prompt with an answer sentence, a list of nodes highlighted as the motif (when present), and an explicit graph description(nodes and edges).
In this QA dataset, all graphs are homogenous and undirected.
Subsets cover both yes/no motif detection, motif-type classification (house vs 5-cycle), and… See the full description on the dataset page: https://huggingface.co/datasets/naos-ku/motif-qa.OpenUAV-QA
✨OpenUAV-QA✨
OpenUAV-QA is a large-scale multiple-choice question-answering benchmark for UAV (drone) navigation decision-making, built upon the TravelUAV (OpenUAV) dataset. It transforms raw UAV flight trajectories into structured, text-polished 4-option QA pairs that test a multimodal model's ability to reason about path planning, action sequences, and spatial dynamics from first-person drone video footage.
The dataset covers 22 distinct simulated… See the full description on the dataset page: https://huggingface.co/datasets/choucsan/OpenUAV-QA.JavaError-QA
JErrRAG-Eval-800
JErrRAG-Eval-800 is the public benchmark release aligned with the paper's final canonical dataset and non-anonymous archival record.
This Hugging Face repository contains:
java_error_qa_v2/: the canonical public benchmark package
paper_online_artifacts/: the paper-facing supplementary artifacts and reproduction bundles
SHA256SUMS.txt: release-side hash anchors referenced by the paper
Dataset Summary
Total records: 800
Split sizes: train=639… See the full description on the dataset page: https://huggingface.co/datasets/HTJ008/JavaError-QA.ICLR2025Openreview基于 20241208 爬取数据制作。
题目和摘要的中文翻译、关键词由 Qwen2.5-72B-Instruct 自动生成,可能存在漏误,大家酌情参考。
生成数据的代码见 repo
bangladesh-legal-qa-dataset
Bangladesh Legal QA Dataset: Bangla-English Law and Fine-Tuning
The Bangladesh Legal QA Dataset is a bilingual Bangla-English dataset for
Bangladesh law question answering, legal NLP, LLM fine-tuning, instruction
tuning, and retrieval-augmented generation (RAG). It provides 2,165
context-grounded legal QA records, direct-answer and IRAC chat-format training
data, and structured statutory text from six Bangladesh Acts and three
schedules.
This is the 2,165-record paper-aligned… See the full description on the dataset page: https://huggingface.co/datasets/momahadi/bangladesh-legal-qa-dataset.duplex-qa-refusal
duplex-qa-refusal
No dialogue in this set has been validated by a human.
Text-side augmentation of the moshika spoken-QA corpus so a full-duplex speech model can be trained to refuse a query when a mid-conversation text instruction tells it to, voice the reason the instruction gives, and then carry on normally. Two classes: policy (an existing benign query is declined for a stated reason; comes with an untouched accept twin sharing pair_id) and attack (a new user turn pivots to… See the full description on the dataset page: https://huggingface.co/datasets/MagicLuke/duplex-qa-refusal.pit-earnings-call-qa
Earnings-Call QA dataset for PIT-4B-FT SFT
Supervised fine-tuning mixture for the PIT (Point-in-Time) line of language models, derived from US public-company earnings-call transcripts. Built to fine-tune the Diamegs/PIT-4B-FT-* snapshots while respecting PIT chronological discipline — no transcript dated after the base model's knowledge cutoff is used in training.
Available snapshots
Each snapshot has its own chronological splits keyed to the base model's… See the full description on the dataset page: https://huggingface.co/datasets/jdecim/pit-earnings-call-qa.PortBench-QA
PortBench QA Dataset
Dataset Description
6,269 structured question-answer pairs probing correlation-based financial reasoning for multi-asset portfolio management, generated from the PortBench Market Base Dataset.
Task Templates
Template
Task
Complexity
Pairs
T1
Return prediction — direction for next N days
1 (single asset)
1,000
T2
Risk assessment — VaR at given confidence level
1
1,000
T3
Position sizing — given max drawdown… See the full description on the dataset page: https://huggingface.co/datasets/AgenticFinLab/PortBench-QA.Nayose-Bench-QA
Dataset Card for Nayose-Bench-Instruction
This dataset was created as a benchmark for the entity resolution task in the pharmaceutical domain.
Dataset Details
This dataset is designed for the entity resolution task in the pharmaceutical domain.
The entity resolution task refers to a paraphrasing task, such as rephrasing drug names, converting chemical substances into brand names, or rewriting chemical substances into chemical formulas.
Uses
from datasets import… See the full description on the dataset page: https://huggingface.co/datasets/EQUES/Nayose-Bench-QA.pit-earnings-call-qa
Earnings-Call QA dataset for PIT-4B-FT SFT
Supervised fine-tuning mixture for the PIT (Point-in-Time) line of language models, derived from US public-company earnings-call transcripts. Built to fine-tune the Diamegs/PIT-4B-FT-* snapshots while respecting PIT chronological discipline — no transcript dated after the base model's knowledge cutoff is used in training.
Available snapshots
Each snapshot has its own chronological splits keyed to the base model's… See the full description on the dataset page: https://huggingface.co/datasets/idleengine/pit-earnings-call-qa.Wikipedia_RAG_QA_Classification
🏛️ Wikipedia RAG QA Dataset for Retrieval-Augmented Generation Training
📊 Dataset Description
This dataset contains 300,000+ validated model-generated responses to Wikipedia content, specifically designed for Retrieval-Augmented Generation (RAG) applications and SQL database insertion tasks. Generated by Jeeney AI Reloaded 207M GPT with specialized RAG tuning.
🖥️ Demo Interface: Discord
Live Chat Demo on Discord: https://discord.gg/Xe9tHFCS9h
The full CJ… See the full description on the dataset page: https://huggingface.co/datasets/CJJones/Wikipedia_RAG_QA_Classification.battery-device-data-qa
Battery Device QA Data
Battery device records, including anode, cathode, and electrolyte.
Examples of the question answering evaluation dataset:
{'question': 'What is the cathode?', 'answer': 'Al foil', 'context': 'The blended slurry was then cast onto a clean current collector (Al foil for the cathode and Cu foil for the anode) and dried at 90 °C under vacuum overnight.', 'start index': 645}
{'question': 'What is the anode?', 'answer': 'Cu foil', 'context': 'The blended slurry was… See the full description on the dataset page: https://huggingface.co/datasets/batterydata/battery-device-data-qa.kdb-qas4temporal-alignment-qalm-eval-results-Kukedlc-Neural-4-QA-7b-private
Dataset Card for Evaluation run of Kukedlc/Neural-4-QA-7b
Dataset automatically created during the evaluation run of model Kukedlc/Neural-4-QA-7b
The dataset is composed of 62 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An additional… See the full description on the dataset page: https://huggingface.co/datasets/nyu-dice-lab/lm-eval-results-Kukedlc-Neural-4-QA-7b-private.SO-Python_QA-Data_Science_and_Machine_Learning_classdelta-mem-qasper-data
Introduction
This repository contains the δ-mem training data, as presented in the paper δ-mem: Efficient Online Memory for Large Language Models.
δ-mem is a lightweight online memory mechanism that augments a frozen backbone with a compact associative memory state. It projects context into a low-dimensional space and updates a state matrix via delta-rule learning, allowing for efficient long-term memory utilization without full fine-tuning or context extension.
Paper:… See the full description on the dataset page: https://huggingface.co/datasets/huaXiaKyrie/delta-mem-qasper-data.wikipedia_multiple_choice_qa
Galician and Portuguese Multiple-Choice QA Instruction Subsets
Dataset description
This dataset contains two instruction-tuning subsets for multiple-choice question answering in Galician and Portuguese:
gl_wikipedia_multiple_choice_qa (1,486 instances)
pt_wikipedia_multiple_choice_qa (547 instances)
Both subsets are reformatted versions of QA data originally included in the cpt_instruction_datasets collection, adapted here as standalone instruction-style datasets.
Each… See the full description on the dataset page: https://huggingface.co/datasets/proxectonos/wikipedia_multiple_choice_qa.ZeroXClem__Qwen2.5-7B-Qandora-CySec-details
Dataset Card for Evaluation run of ZeroXClem/Qwen2.5-7B-Qandora-CySec
Dataset automatically created during the evaluation run of model ZeroXClem/Qwen2.5-7B-Qandora-CySec
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/ZeroXClem__Qwen2.5-7B-Qandora-CySec-details.SO-Python_QA-System_Administration_and_DevOps_classbunnycore__Qandora-2.5-7B-Creative-details
Dataset Card for Evaluation run of bunnycore/Qandora-2.5-7B-Creative
Dataset automatically created during the evaluation run of model bunnycore/Qandora-2.5-7B-Creative
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/bunnycore__Qandora-2.5-7B-Creative-details.stackoverflow-qa-top-300kGrayLine-QA-Reasoning
not releasing to the public yet, please reach out to me on Discord if you would like access to this dataset @_duude
datashop-medical-qaenvoy-qasper-code-trajectories
Envoy QASPER Code-Execution Trajectory Pilot
This is a small, fully disclosed pilot of executable research-agent trajectories.
Claude Sonnet 5 generated Python actions against a persistent document REPL. The
Envoy pipeline executed every action and retained the real observations. An AI
coding assistant then reviewed answer support, stopping behavior, and replay.
This release is useful for studying trajectory validation and citation failures.
It is not a production-ready SFT… See the full description on the dataset page: https://huggingface.co/datasets/jasonlingg/envoy-qasper-code-trajectories.aura_qa
Affect-Uniform ReAding QA (AURA-QA),
This dataset contains short passages from English texts found in Project Gutenberg paired with question–answer examples and emotion labels. The dataset is designed to support research in emotion-aware reading comprehension. Answers are constrained to 1–3 tokens and are generated and verified by large language models.
Dataset Structure
text — Passage excerpt
question — Question about the passage
answer — Short answer (1–3 tokens)… See the full description on the dataset page: https://huggingface.co/datasets/avalab/aura_qa.SO-Python_QA-Database_and_SQL_class
Dataset Card for Dataset Name
Dataset Summary
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Supported Tasks and Leaderboards
[More Information Needed]
Languages
[More Information Needed]
Dataset Structure
Data Instances
[More Information Needed]
Data Fields
[More Information Needed]
Data Splits
[More Information Needed]
Dataset Creation… See the full description on the dataset page: https://huggingface.co/datasets/RazinAleks/SO-Python_QA-Database_and_SQL_class.Malaysian-RAG-QA
Malaysian RAG Dataset
This dataset contains question-answering pairs with associated context documents and evaluation metrics. Each entry includes a source document, a question and an answer generated by ChatGPT-5.1, and quality scores for context precision and faithfulness evaluated using RAGAS.
Data Sources
Documents sourced from malaysia-ai/pretrain-text-dataset:
maktabahalbakri.com
muftiwp.gov.my.dedup
asklegal
dewanbahasa-jdbp
gov.my
Dataset Splits… See the full description on the dataset page: https://huggingface.co/datasets/Scicom-intl/Malaysian-RAG-QA.
