datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
MedReason
MedReason: Eliciting Factual Medical Reasoning Steps in LLMs via Knowledge Graphs
📃 Paper |🤗 MedReason-8B | 📚 MedReason Data
✨ Latest News
[05/27/2025] 🎉 MedReason wins 3rd prize🏆 in the Huggingface Reasoning Datasets Competition!
⚡Introduction
MedReason is a large-scale high-quality medical reasoning dataset designed to enable faithful and explainable medical problem-solving in large language models (LLMs).
We utilize a structured medical knowledge… See the full description on the dataset page: https://huggingface.co/datasets/UCSC-VLAA/MedReason.MedTrinity-25M
Tutorial of using Medtrinity-25M
MedTrinity-25M, a comprehensive, large-scale multimodal dataset for medicine, covering over 25 million images across 10 modalities, with multigranular annotations for more than 65 diseases. These enriched annotations encompass both global textual information, such as disease/lesion type, modality, region-specific descriptions, and inter-regional relationships, as well as detailed local annotations for regions of interest (ROIs), including bounding… See the full description on the dataset page: https://huggingface.co/datasets/UCSC-VLAA/MedTrinity-25M.VisualClawArena
VisualClawArena
Webpage: https://ucsc-vlaa.github.io/VisualClaw/ |
Paper: https://arxiv.org/abs/2606.16295 |
Code: https://github.com/UCSC-VLAA/VisualClaw
VisualClawArena is a 200-scenario benchmark for multimodal computer-use agents.
Each scenario pairs a video clip with a persistent workspace, dynamic updates,
multi-round instructions, and executable checkers. The benchmark is designed to
test whether an agent can use video evidence, workspace files, and later
environment… See the full description on the dataset page: https://huggingface.co/datasets/UCSC-VLAA/VisualClawArena.efficient-vla-extracts
efficient-vla-extracts
Parser-derived artifacts for the efficient-vla-wiki project.
This dataset is the data layer companion to the main repository:
GitHub: guanweifan/efficient-vla-wiki
Dataset repo: efficient-vla-extracts
What is inside
The staged upload keeps the local directory structure:
extracts/
├── meta/
│ ├── extract_build_index.jsonl
│ └── extract_build_status.json
└── parses/
└── <paper_id>/
Each paper directory may contain:
pdftotext.txt… See the full description on the dataset page: https://huggingface.co/datasets/guanweifan/efficient-vla-extracts.m23k-tokenized
m1: Unleash the Potential of Test-Time Scaling for Medical Reasoning in Large Language Models
A simple test-time scaling strategy, with minimal fine-tuning, can unlock strong medical reasoning within large language models.
⚡ Introduction
Hi! Welcome to the huggingface repository for m1 (https://github.com/UCSC-VLAA/m1)!
m1 is a medical LLM designed to enhance reasoning through efficient test-time scaling. It enables lightweight models to match or exceed the performance of… See the full description on the dataset page: https://huggingface.co/datasets/UCSC-VLAA/m23k-tokenized.m1k-tokenized
m1: Unleash the Potential of Test-Time Scaling for Medical Reasoning in Large Language Models
A simple test-time scaling strategy, with minimal fine-tuning, can unlock strong medical reasoning within large language models.
⚡ Introduction
Hi! Welcome to the huggingface repository for m1 (Github, Paper)!
m1 is a medical LLM designed to enhance reasoning through efficient test-time scaling. It enables lightweight models to match or exceed the performance of much larger… See the full description on the dataset page: https://huggingface.co/datasets/UCSC-VLAA/m1k-tokenized.VLM-CapCurriculum-TextReasoning-Data
VLM-CapCurriculum-TextReasoning (D_text)
Stage-2 textual-reasoning data for the staged post-training recipe in
"From Seeing to Thinking: Decoupling Perception and Reasoning Improves Post-Training of Vision-Language Models"
(ICML 2026).
A curated ORZ-Math-13k subset — challenging text-only math problems used to consolidate textual reasoning between the perception (Stage 1) and visual-reasoning (Stage 3) RLVR stages of our recipe. Every row also ships with a precomputed pass_rate so… See the full description on the dataset page: https://huggingface.co/datasets/UCSC-VLAA/VLM-CapCurriculum-TextReasoning-Data.DAM-QA-annotations
DAM-QA Unified Annotations
22,675 question-answer pairs from 6 major VQA benchmarks, unified for the DAM-QA framework. This collection consolidates annotations from InfographicVQA, TextVQA, VQAv2, DocVQA, ChartQA, and ChartQA-Pro into standardized JSONL formats.
📖 Paper: Describe Anything Model for Visual Question Answering on Text-rich Images⚠️ Note: Images not included - obtain from original sources with proper licensing
Repository Structure
DAM-QA-annotations/… See the full description on the dataset page: https://huggingface.co/datasets/VLAI-AIVN/DAM-QA-annotations.polaris-support-tickets-v2
Polaris Support Tickets — Synthetic Dataset (v2)
~24,000 synthetic customer-support tickets for Polaris, a fictional
multichannel analytics SaaS. Each ticket carries coherent ground-truth labels for
triage (topic · type · priority · routing · sentiment) plus a noisy intake
category, and the collection spans Jan 2024 → Jun 2026 with a realistic
temporal event layer (service outages and product launches).
Synthetic data, generated on a real pipeline. No real users, no PII. Built… See the full description on the dataset page: https://huggingface.co/datasets/VladislavMarinovich/polaris-support-tickets-v2.my-distiset-bf9522fa
Dataset Card for my-distiset-bf9522fa
This dataset has been created with distilabel.
Dataset Summary
This dataset contains a pipeline.yaml which can be used to reproduce the pipeline that generated it in distilabel using the distilabel CLI:
distilabel pipeline run --config "https://huggingface.co/datasets/vladimirtan/my-distiset-bf9522fa/raw/main/pipeline.yaml"
or explore the configuration:
distilabel pipeline info --config… See the full description on the dataset page: https://huggingface.co/datasets/vladimirtan/my-distiset-bf9522fa.isforsmol
Dataset Card for isforsmol
This dataset has been created with distilabel.
Dataset Summary
This dataset contains a pipeline.yaml which can be used to reproduce the pipeline that generated it in distilabel using the distilabel CLI:
distilabel pipeline run --config "https://huggingface.co/datasets/vladimirlavrentev/isforsmol/raw/main/pipeline.yaml"
or explore the configuration:
distilabel pipeline info --config… See the full description on the dataset page: https://huggingface.co/datasets/vladimirlavrentev/isforsmol.ViInfographicVQA
ViInfographicVQA
Overview
ViInfographicVQA is a Vietnamese Visual Question Answering (VQA) benchmark for infographic understanding.It evaluates models’ ability to read, reason, and synthesize information from data-rich, layout-heavy visuals that mix text, charts, maps, and design elements.
Two settings are provided:
Single-image VQA – questions answered from one infographic.
Multi-image VQA – questions requiring reasoning across multiple, semantically related… See the full description on the dataset page: https://huggingface.co/datasets/VLAI-AIVN/ViInfographicVQA.
