datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
CT-RATE
The CT-RATE Team organizes the VLM3D Challenge
VLM3D 2026 (2nd Edition) → Challenge Finals at MICCAI 2026
VLM3D 2025 (1st Edition) → Challenge Finals at MICCAI 2025 • Workshop at ICCV 2025
The CT-RATE Team is developing the MR-RATE Dataset
A large-scale brain MRI dataset with paired radiology reports for training 3D vision-language models.
GitHub |
Dataset |
Metadata Dashboard
Generalist Foundation Models from a Multimodal Dataset for 3D Computed Tomography… See the full description on the dataset page: https://huggingface.co/datasets/ibrahimhamamci/CT-RATE.gpqa
Dataset Card for GPQA
GPQA is a multiple-choice, Q&A dataset of very hard questions written and validated by experts in biology, physics, and chemistry. When attempting questions out of their own domain (e.g., a physicist answers a chemistry question), these experts get only 34% accuracy, despite spending >30m with full access to Google.
We request that you do not reveal examples from this dataset in plain text or images online, to reduce the risk of leakage into foundation… See the full description on the dataset page: https://huggingface.co/datasets/Idavidrein/gpqa.hacker-news
Hacker News - Complete Archive
Every Hacker News item since 2006, live-updated every 5 minutes
What is it?
This dataset contains the complete Hacker News archive: every story, comment, Ask HN, Show HN, job posting, and poll ever submitted to the site. Hacker News is one of the longest-running and most influential technology communities on the internet, operated by Y Combinator since 2007. It has become the de facto gathering place for founders, engineers… See the full description on the dataset page: https://huggingface.co/datasets/open-index/hacker-news.general-instruction-augmented-corpora
Instruction Pre-Training: Language Models are Supervised Multitask Learners (EMNLP 2024)
This repo contains the general instruction-augmented corpora (containing 200M instruction-response pairs covering 40+ task categories) used in our paper Instruction Pre-Training: Language Models are Supervised Multitask Learners.
We explore supervised multitask pre-training by proposing Instruction Pre-Training, a framework that scalably augments massive raw corpora with instruction-response… See the full description on the dataset page: https://huggingface.co/datasets/instruction-pretrain/general-instruction-augmented-corpora.python_code_instructions_18k_alpaca
Dataset Card for python_code_instructions_18k_alpaca
The dataset contains problem descriptions and code in python language.
This dataset is taken from sahil2801/code_instructions_120k, which adds a prompt column in alpaca style. Refer to the source here.
CodeFeedback-Filtered-Instruction OpenCodeInterpreter: Integrating Code Generation with Execution and Refinement
[🏠Homepage]
|
[🛠️Code]
OpenCodeInterpreter
OpenCodeInterpreter is a family of open-source code generation systems designed to bridge the gap between large language models and advanced proprietary systems like the GPT-4 Code Interpreter. It significantly advances code generation capabilities by integrating execution and iterative refinement functionalities.
For further information and… See the full description on the dataset page: https://huggingface.co/datasets/m-a-p/CodeFeedback-Filtered-Instruction.molecule_property_instruction
Dataset Card for "molecule_property_instruction"
More Information needed
mse-text-img-dataset
Dataset Card for MSE-text-img-dataset
We have created a custom dataset
that is extracted as a subset of the Math Stack Exchange (MSE)
dataset. This text-image dataset contains 64,860 questions with
their respective list of answers, scores, acceptance marking,
and image versions of each question and answer generated
from the stored text with embedded LaTeX math markup.
In this dataset there are 117,380 answers in total, with 1.81
answers per question on average. Each image… See the full description on the dataset page: https://huggingface.co/datasets/Zenos5/mse-text-img-dataset.open-india-law
Open India Law
Open, structured Indian primary law - plus the scrapers that build it.
Every judgment of the Supreme Court of India and all 25 High Courts, the decisions of 15
tribunals and regulators, and Central, State and Union Territory legislation down to the
individual section. Normalized to one schema, exclusively from official government sources.
Volume
Period
Court judgments
12,848,644
1950 to 2025
Tribunal and regulator matters
813,168
1985 to 2026… See the full description on the dataset page: https://huggingface.co/datasets/vaquill/open-india-law.muri-it-language-split
MURI-IT: Multilingual Instruction Tuning Dataset for 200 Languages via Multilingual Reverse Instructions
MURI-IT is a large-scale multilingual instruction tuning dataset containing 2.2 million instruction-output pairs across 200 languages. It is designed to address the challenges of instruction tuning in low-resource languages with Multilingual Reverse Instructions (MURI), which ensures that the output is human-written, high-quality, and authentic to the cultural and linguistic… See the full description on the dataset page: https://huggingface.co/datasets/akoksal/muri-it-language-split.tmmluplus
TMMLU+ : Large scale traditional chinese massive multitask language understanding
iKala presents TMMLU+, a large-scale benchmark for evaluating LLM capabilities in Traditional Chinese, with content primarily reflecting Taiwan's linguistic, educational, and professional contexts. It covers 66 subjects, from elementary to professional domains, and is approximately six times larger than TMMLU with broader, more balanced coverage.
TMMLU+ v1.1 improves benchmark quality through… See the full description on the dataset page: https://huggingface.co/datasets/ikala/tmmluplus.WildClawBenchWildClawBench
Hard, practical, end-to-end evaluation for AI agents — in the wild.
WildClawBench is an agent benchmark that tests what actually matters: can an AI agent do real work, end-to-end, without hand-holding?
We drop agents into a live OpenClaw environment — the same open-source personal AI assistant that real users rely on daily — and throw 60 original tasks at them: clipping goal highlights from a football match, negotiating meeting times over multi-round… See the full description on the dataset page: https://huggingface.co/datasets/internlm/WildClawBench.XLRS-Bench-lite
🐙GitHub
Information or evaluatation on this dataset can be found in this repo: https://github.com/AI9Stars/XLRS-Bench
📜Dataset License
Annotations of this dataset is released under a Creative Commons Attribution-NonCommercial 4.0 International License. For images from:
DOTARGB images from Google Earth and CycloMedia (for academic use only; commercial use is prohibited, and Google Earth terms of use apply).
ITCVDLicensed under CC-BY-NC-SA-4.0.
MiniFrance… See the full description on the dataset page: https://huggingface.co/datasets/initiacms/XLRS-Bench-lite.Trendyol-Cybersecurity-Instruction-Tuning-Dataset
Trendyol Cybersecurity Defense Instruction-Tuning Dataset (v2.0)
🚀 TL;DR
53,202 meticulously curated system/user/assistant instruction-tuning examples covering 200+ specialized cybersecurity domains. Built by the Trendyol Security Team for training state-of-the-art defensive security AI assistants. Expanded from 21K to 53K rows with comprehensive coverage of modern security challenges including cloud-native threats, AI/ML security, quantum computing risks… See the full description on the dataset page: https://huggingface.co/datasets/Trendyol/Trendyol-Cybersecurity-Instruction-Tuning-Dataset.medical_meadow_wikidoc_patient_information
Dataset Card for WikiDoc
For the dataset containing rephrased content from the living textbook refer to this dataset
Dataset Summary
This dataset containes medical question-answer pairs extracted from WikiDoc,
a collaborative platform for medical professionals to share and contribute to up-to-date medical knowledge.
The platform has to main subsites, the "Living Textbook" and "Patient Information". The "Living Textbook"
contains chapters for various medical specialties… See the full description on the dataset page: https://huggingface.co/datasets/medalpaca/medical_meadow_wikidoc_patient_information.isafpressreleases
ISAF Press Releases Dataset Description
Homepage: [N/A]
Repository: [N/A]
Paper: A Knock on the Door: 22 Months of ISAF Press Releases
Point of Contact: Alex Strick van Linschoten (@strickvl)
Dataset Summary
The ISAF Press Releases dataset contains data used as the basis for the research
paper "A Knock on the Door: 22 Months of ISAF Press Releases". The dataset
provides a comprehensive collection of press releases issued by the
International Security Assistance… See the full description on the dataset page: https://huggingface.co/datasets/strickvl/isafpressreleases.FinFIRST
FinFIRST: Financial Information Retrieval, Sourcing and Traceability
Released alongside Ling-3.0-flash-Fin, FinFIRST is an open benchmark for evaluating whether financial search agents can produce answers that are not only correct, but also supported by authoritative, timely, and verifiable evidence. It was developed by Ant Group, with professional support from the investment banking team at China International Capital Corporation Limited (CICC).
Financial research requires more… See the full description on the dataset page: https://huggingface.co/datasets/inclusionAI/FinFIRST.ID_Legal_QA_SynDeepThink
🧠 Indonesian Legal QA SynDeepThink Dataset
This repository hosts a specialized Indonesian Legal QA dataset that incorporates a Deep Thinking Phase. It is engineered for researchers and developers focusing on high-level judicial reasoning and complex regulatory analysis. 🏛️
💡 The Concept: Deep Thinking vs. Standard QA
While standard models often provide "System 1" (snap) judgments, the SynDeepThink approach simulates "System 2" (slow, deliberate) thinking. This dataset… See the full description on the dataset page: https://huggingface.co/datasets/Azzindani/ID_Legal_QA_SynDeepThink.xcsr
Dataset Card for X-CSR
Dataset Summary
To evaluate multi-lingual language models (ML-LMs) for commonsense reasoning in a cross-lingual zero-shot transfer setting (X-CSR), i.e., training in English and test in other languages, we create two benchmark datasets, namely X-CSQA and X-CODAH. Specifically, we automatically translate the original CSQA and CODAH datasets, which only have English versions, to 15 other languages, forming development and test sets for studying X-CSR.… See the full description on the dataset page: https://huggingface.co/datasets/INK-USC/xcsr.LLaVA-Instruct-150K
LLaVA Visual Instruct 150K Dataset Card
Dataset details
Dataset type:
LLaVA Visual Instruct 150K is a set of GPT-generated multimodal instruction-following data.
It is constructed for visual instruction tuning and for building large multimodal towards GPT-4 vision/language capability.
Dataset date:
LLaVA Visual Instruct 150K was collected in April 2023, by prompting GPT-4-0314 API.
Paper or resources for more information:
https://llava-vl.github.io/
License:… See the full description on the dataset page: https://huggingface.co/datasets/liuhaotian/LLaVA-Instruct-150K.image2struct-latex-v1
Image2Struct - Latex
Paper | Website | Datasets (Webpages, Latex, Music sheets) | Leaderboard | HELM repo | Image2Struct repo
License: Apache License Version 2.0, January 2004
Dataset description
Image2struct is a benchmark for evaluating vision-language models in practical tasks of extracting structured information from images.
This subdataset focuses on LaTeX code. The model is given an image of the expected output with the prompt:
Please provide the LaTex code used to… See the full description on the dataset page: https://huggingface.co/datasets/stanford-crfm/image2struct-latex-v1.MAmmoTH-VL-Instruct-12M
MAmmoTH-VL-Instruct-12M
🏠 Homepage | 🤖 MAmmoTH-VL-8B | 💻 Code | 📄 Arxiv | 📕 PDF | 🖥️ Demo
Introduction
Our simple yet scalable visual instruction data rewriting pipeline consists of three steps: manual data source collection, rewriting using MLLMs/LLMs, and filtering via the same MLLM as a judge. Examples below illustrate transformations in math and science categories, showcasing detailed, step-by-step responses.
The data distribution of… See the full description on the dataset page: https://huggingface.co/datasets/MAmmoTH-VL/MAmmoTH-VL-Instruct-12M.duorc
Dataset Card for duorc
Dataset Summary
The DuoRC dataset is an English language dataset of questions and answers gathered from crowdsourced AMT workers on Wikipedia and IMDb movie plots. The workers were given freedom to pick answer from the plots or synthesize their own answers. It contains two sub-datasets - SelfRC and ParaphraseRC. SelfRC dataset is built on Wikipedia movie plots solely. ParaphraseRC has questions written from Wikipedia movie plots and the answers are… See the full description on the dataset page: https://huggingface.co/datasets/ibm-research/duorc.terminal-bench
Terminal-Bench Dataset
This dataset contains tasks from Terminal-Bench, a benchmark for evaluating AI agents in real terminal environments. Each task is packaged as a complete, self-contained archive that preserves the exact directory structure, binary files, Docker configurations, and test scripts needed for faithful reproduction.
The archive column contains a gzipped tarball of the entire task directory.
Dataset Overview
Terminal-Bench evaluates AI agents on… See the full description on the dataset page: https://huggingface.co/datasets/ia03/terminal-bench.train_video_and_instruction
ShareGPTVideo Training Data
All dataset and models can be found at ShareGPTVideo.
Contents:
Train 300k video frames: contains video frames used for SFT and DPO model, which is a subset of total 900k.
ActivityNet 50k + vidal 150k + webvid 100k.
Train 600k video frames: contains the rest 600k frames, the total 900k frames are used for pre-training stage. If you just do finetuning using our video QA, you can just download the 300k above.
900k composition is 400k WebVid +… See the full description on the dataset page: https://huggingface.co/datasets/ShareGPTVideo/train_video_and_instruction.acp_bench
ACP Bench
🏠 Homepage •
📄 Paper •
📄 Paper
ACPBench is a benchmark dataset designed to evaluate the reasoning capabilities of large language models (LLMs) in the context of Action, Change, and Planning. It spans 13 diverse domains:
Blocksworld
Logistics
Grippers
Grid
Ferry
FloorTile
Rovers
VisitAll
Depot
Goldminer
Satellite
Swap
Alfworld
Task Types in ACPBench
ACPBench includes the following 8 reasoning tasks:
Action Applicability (app)… See the full description on the dataset page: https://huggingface.co/datasets/ibm-research/acp_bench.ifc-bench
IFC-Bench
A benchmark dataset for evaluating BIM (Building Information Modeling) comprehension and reasoning capabilities in AI systems. Provides curated IFC models with question-answer pairs across 4 complexity categories for testing BIM-related AI implementations.
Dataset snapshot:
question
ground_truth
ifc_model
project
category
0
What modelling program and IFC standard were used to create this model?
The model was created using...
arc
4351
1
1
What are the… See the full description on the dataset page: https://huggingface.co/datasets/sylvainHellin/ifc-bench.epstein-files-ocr-datasets-1-8-early-release
Epstein Files OCR — Datasets 1–8 (Early Release)
ARCHIVE NOTICE
This dataset is no longer maintained. Please refer to the Epstein Files — Complete OCR Dataset.
Dataset Summary
This dataset contains page-level OCR output (as Markdown) from a public release of documents related to Jeffrey Epstein / the Epstein case.
Each Markdown file represents one scanned page converted to text using an automated OCR pipeline. The dataset is designed for:
Question answering
Information… See the full description on the dataset page: https://huggingface.co/datasets/ishumilin/epstein-files-ocr-datasets-1-8-early-release.ID_Legal_QA_SynThink
🧠 Indonesian Legal QA Synthetic Think Dataset (ID_Legal_QA_SynThink)
This repository features an advanced Synthetic Question-and-Answer dataset for the Indonesian legal domain, distinguished by the inclusion of an explicit Thinking Phase (Chain-of-Thought). 🏛️
💡 The Concept: Transparent Legal Reasoning
Standard QA datasets often provide just the "final answer." This dataset goes deeper by capturing the internal reasoning process of the model before it arrives at a… See the full description on the dataset page: https://huggingface.co/datasets/Azzindani/ID_Legal_QA_SynThink.InfoSeek
InfoSeek: Open Data Synthesis For Deep Research
Paper | Code
Dataset Information
data/InfoSeek.jsonl
Contains the full research tree structures of InfoSeek. Each sample starts from a root node with a research question, its corresponding entity, and process information for sub-questions (stored in root). Also expands into intermediate tree structure during each step of construction (stored in all_tree_list). Totally 52K samples.
data/InfoSeekQA.jsonl
A collection… See the full description on the dataset page: https://huggingface.co/datasets/Lk123/InfoSeek.
