datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
healthspend-dataMedReason
MedReason: Eliciting Factual Medical Reasoning Steps in LLMs via Knowledge Graphs
📃 Paper |🤗 MedReason-8B | 📚 MedReason Data
✨ Latest News
[05/27/2025] 🎉 MedReason wins 3rd prize🏆 in the Huggingface Reasoning Datasets Competition!
⚡Introduction
MedReason is a large-scale high-quality medical reasoning dataset designed to enable faithful and explainable medical problem-solving in large language models (LLMs).
We utilize a structured medical knowledge… See the full description on the dataset page: https://huggingface.co/datasets/UCSC-VLAA/MedReason.VlaserW2-VLA-CoT
World-to-Wrist: Offline CoT Labels
This dataset contains frame-aligned offline chain-of-thought annotations used
to train W²-VLA policies on LIBERO, RoboTwin, and four real-world manipulation
tasks. Matching LeRobot action data is available in W2-VLA-Training-Data.
Dataset Structure
W2-VLA-CoT/
├── libero/
│ ├── libero_10_no_noops_1.0.0_lerobot/
│ ├── libero_goal_no_noops_1.0.0_lerobot/
│ ├── libero_object_no_noops_1.0.0_lerobot/
│ └──… See the full description on the dataset page: https://huggingface.co/datasets/yuuu94/W2-VLA-CoT.Robot-VLA-R1ClinSeek-Bench
ClinSeek-Bench
ClinSeek-Bench is the evaluation suite introduced in
ClinSeekAgent: Automating Multimodal Evidence Seeking for Agentic Clinical
Reasoning. It evaluates clinical reasoning
under two paired settings with the same task definitions and answer labels:
Curated Input: the model answers from the evidence package provided by
the source benchmark.
Automated Evidence-Seeking: the curated context is removed, and the model
must retrieve evidence from raw clinical data using… See the full description on the dataset page: https://huggingface.co/datasets/UCSC-VLAA/ClinSeek-Bench.vla-reasoningthinkflow-vla-features-b2tsiolkovsky-papers
Tsiolkovsky Papers: the complete personal archive as a machine-readable corpus
Machine transcriptions of all 51,008 sheets of fond 555 of the Archive of the
Russian Academy of Sciences — the personal archive of Konstantin Tsiolkovsky
(1857–1935), who derived the rocket equation and described the multistage rocket
decades before anyone could test either.
The archive had been scanned and put online, but without a catalogue you could
query, full-text search, or a dataset. This is… See the full description on the dataset page: https://huggingface.co/datasets/vladimirbesk/tsiolkovsky-papers.VLM-CapCurriculum-Perception-Data
VLM-CapCurriculum-Perception (D_perc)
Stage-1 visual perception data for the staged post-training recipe in
"From Seeing to Thinking: Decoupling Perception and Reasoning Improves Post-Training of Vision-Language Models"
(ICML 2026).
Each sample is a 4-way multiple-choice question over an image where the question can be answered from a fine-grained image caption but is missed by a strong VLM looking only at the image — by construction, these samples isolate perception failures from… See the full description on the dataset page: https://huggingface.co/datasets/UCSC-VLAA/VLM-CapCurriculum-Perception-Data.ViVQA-X
Dataset Card for ViVQA-X
Dataset Description
ViVQA-X is the Vietnamese version of the VQA-X dataset, designed for tasks in VQA-NLE in the Vietnamese language. This dataset was created to support research in Vietnamese VQA, offering translated examples with explanations, in alignment with the original English dataset. The ViVQA-X dataset is generated using a multi-stage pipeline for translation and evaluation, including both automated translation and post-processing, with… See the full description on the dataset page: https://huggingface.co/datasets/VLAI-AIVN/ViVQA-X.robotrace-vla-robustness-traces
RoboTrace Evidence Bundle
This dataset repository contains the public evidence bundle for RoboTrace, a low-cost deployment-stress evaluation scaffold for robot-learning and VLA-style inference pipelines.
The current release evaluates lerobot/pusht and includes reports, metrics, plots, summaries, and release manifests from a complete staged run.
What this bundle is for
Use this repository to inspect evidence from RoboTrace:
action-trace stability metrics
visual… See the full description on the dataset page: https://huggingface.co/datasets/i-am-shaurya05/robotrace-vla-robustness-traces.vlabench_primitive_ft_lerobot_video
VLABench Primitive Tasks — LeRobot v3.0 (TsFile)
Apache TsFile version of VLABench/vlabench_primitive_ft_lerobot_video.
Overview
This dataset is organized in the LeRobot v3.0 format and is used for integrating VLABench into the LeRobot framework officially. Compared with the v2.0 and the RLDS versions, this release stores the visual observations in a video-compressed format rather than as individual image files, giving better storage efficiency and data-loading… See the full description on the dataset page: https://huggingface.co/datasets/THULab/vlabench_primitive_ft_lerobot_video.Air-Chat
Format
Each entry contains:
question: the user’s question
answer: the assistant’s response
Example
{"question": "Hello!", "answer": "Hello to you too!"}
File architecture
dialogue_dataset/
├── README.md
├── dataset_infos.json
├── data/
│ ├──ru/
│ │ └── train-ru-00001-of-00001.jsonl
│ └──en/
│ └── train-en-00001-of-00001.jsonl
└── .gitattributes
Load Russian version
dataset =… See the full description on the dataset page: https://huggingface.co/datasets/Vladimir0-1/Air-Chat.urban-vla-expert-v1
Urban VLA Expert v1
Urban VLA Expert v1 is a simulator dataset for language-conditioned urban driving. Each frame pairs a 256 x 256 front-camera image with ego state, a natural-language instruction, and continuous driving controls.
This is a small research dataset, not evidence that a policy is ready for a real vehicle. The expert is a deterministic simulator controller, and the language prompts are curated paraphrases rather than speech collected from drivers.
What… See the full description on the dataset page: https://huggingface.co/datasets/Mayank022/urban-vla-expert-v1.VLM-CapCurriculum-TextReasoning-Data
VLM-CapCurriculum-TextReasoning (D_text)
Stage-2 textual-reasoning data for the staged post-training recipe in
"From Seeing to Thinking: Decoupling Perception and Reasoning Improves Post-Training of Vision-Language Models"
(ICML 2026).
A curated ORZ-Math-13k subset — challenging text-only math problems used to consolidate textual reasoning between the perception (Stage 1) and visual-reasoning (Stage 3) RLVR stages of our recipe. Every row also ships with a precomputed pass_rate so… See the full description on the dataset page: https://huggingface.co/datasets/UCSC-VLAA/VLM-CapCurriculum-TextReasoning-Data.opensubs-collocations
OpenSubtitles Collocations
NPMI-scored bigram collocations extracted from the OpenSubtitles parallel corpus. Three languages, three relation types, ~43K bigrams total.
Languages & Corpus Size
Language
Code
Corpus lines
Bigrams
English
en
~100M
15,000
Dutch
nl
~105M
15,000
Serbian
sr
~50M
13,586
Relation Types
ADJ+NOUN — adjective-noun pairs: "slim contract", "kreditan kartica"
VERB+ADP — phrasal verbs / verb-preposition: "come on", "houden… See the full description on the dataset page: https://huggingface.co/datasets/vladvlasov256/opensubs-collocations.ViLReward-73KProcess Reward Data for ViLBench: A Suite for Vision-Language Process Reward Modeling
Paper | Project Page
There are 73K vision-language process reward data sourcing from five training sets.
DAM-QA-annotations
DAM-QA Unified Annotations
22,675 question-answer pairs from 6 major VQA benchmarks, unified for the DAM-QA framework. This collection consolidates annotations from InfographicVQA, TextVQA, VQAv2, DocVQA, ChartQA, and ChartQA-Pro into standardized JSONL formats.
📖 Paper: Describe Anything Model for Visual Question Answering on Text-rich Images⚠️ Note: Images not included - obtain from original sources with proper licensing
Repository Structure
DAM-QA-annotations/… See the full description on the dataset page: https://huggingface.co/datasets/VLAI-AIVN/DAM-QA-annotations.action-evidence-vla-phase-state-cacheAlpha-Instruct
Alpha-Instruct
A synthetic instruction-tuning dataset for quantitative finance, covering formulaic alphas, technical indicators, and academic factor definitions. Designed to fine-tune language models on the vocabulary and reasoning patterns of quant researchers.
Dataset Summary
336 rows of instruction–response pairs in chat format, generated from three distinct quant finance source corpora and post-processed to remove noise and near-duplicates.
Each example is a messages… See the full description on the dataset page: https://huggingface.co/datasets/VladHong/Alpha-Instruct.edupres
Dataset Card for Edupres.ru Presentations
Dataset Summary
This dataset contains metadata about 44,210 presentations from the edupres.ru platform, with 21,941 presentations available in their original format. The dataset includes information such as presentation titles, descriptions, authors, publication dates, and file sizes. The presentations are primarily in Russian and cover various educational topics.
Languages
The dataset is multilingual, with Russian… See the full description on the dataset page: https://huggingface.co/datasets/vladtest88888888/edupres.vla-evaluation-v3vlangchatbotVladdy_dataset_slimfinalVladdy_Dataset_RefactoredVladdy_AwarenessMNLP_M2_mcqa_dataset
