datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
path-vqa
Dataset Card for PathVQA
Dataset Description
PathVQA is a dataset of question-answer pairs on pathology images. The dataset is intended to be used for training and testing
Medical Visual Question Answering (VQA) systems. The dataset includes both open-ended questions and binary "yes/no" questions.
The dataset is built from two publicly-available pathology textbooks: "Textbook of Pathology" and "Basic Pathology", and a
publicly-available digital library: "Pathology… See the full description on the dataset page: https://huggingface.co/datasets/flaviagiammarino/path-vqa.PathEval
PathEval: A Benchmark for Evaluating Vision-Language Models as Evaluators for Path Planning
Overview
Despite their promise to perform complex reasoning, large language models (LLMs) have been shown to have limited effectiveness in end-to-end planning. This has inspired an intriguing question: if these models cannot plan well, can they still contribute to the planning framework as a helpful plan evaluator? In this work, we generalize this question to consider LLMs… See the full description on the dataset page: https://huggingface.co/datasets/maghzal/PathEval.Fino1_Reasoning_Path_FinQAFino1 is a financial reasoning dataset based on FinQA, with GPT-4o-generated reasoning paths to enhance structured financial question answering.
For more details, please check our paper arxiv.org/abs/2502.08127.
Source Data
Initial Data Collection and Normalization
The dataset originates from FinQA dataset.
Annotations
Annotation Process
We add a prompt and create a reasoning process using GPT-4o for each question-answer pair.
💡 Citation… See the full description on the dataset page: https://huggingface.co/datasets/TheFinAI/Fino1_Reasoning_Path_FinQA.PATHOS-PLM-EMBEDDINGS
PATHOS PLM Embeddings
Precomputed protein language model (PLM) embeddings for missense substitutions and wild-type residues in 20,416 human SwissProt proteins. These embeddings are used by PATHOS to predict the pathogenicity of missense mutations.
Paper: http://dx.doi.org/10.1016/j.ailsci.2026.100165
Dataset Structure
The repository contains two config families for each PLM:
Mutation configs: <model> stores embeddings for generated missense substitutions.
Wild-type… See the full description on the dataset page: https://huggingface.co/datasets/DSIMB/PATHOS-PLM-EMBEDDINGS.indoor-anomaly-detection-path-obstruction-monitoring
Indoor Anomaly Detection & Path Obstruction Monitoring
Generated by datapack-import.ts
This dataset mirrors public data-pack render outputs from Physicl.
Each row represents one render view. The image column contains a stable URL to the primary render image uploaded under /data; image_path stores the relative repository path and data_commit_sha pins the Hugging Face dataset commit used by those URLs. Files are uploaded as downloaded unless optional PNG recompression is enabled… See the full description on the dataset page: https://huggingface.co/datasets/physicl/indoor-anomaly-detection-path-obstruction-monitoring.pathological_speech
Pathological Speech (TORGO + UA-Speech + LibriSpeech Normal)
Mixed-corpus speech dataset for training and evaluating controllable
speech-synthesis and severity-classification models. Three corpora are merged
with unified metadata so a single model can learn severity- and
gender-conditioned generation without confounds.
Splits (speaker-disjoint since 2026-09-14)
Split
Rows
Bytes (parquet)
What it is
train
37704
5,500,328,387
every clip of every speaker… See the full description on the dataset page: https://huggingface.co/datasets/resproj007/pathological_speech.plant-pathology-2021
Description
Dataset from the Plant Pathology 2021 (FGVC8) Challenge.
'
For Plant Pathology 2021-FGVC8, we have significantly increased the number of foliar disease images and added additional disease categories. This year’s dataset contains approximately 23,000 high-quality RGB images of apple foliar diseases, including a large expert-annotated disease dataset. This dataset reflects real field scenarios by representing non-homogeneous backgrounds of leaf images taken at different… See the full description on the dataset page: https://huggingface.co/datasets/timm/plant-pathology-2021.Alexandria_geometry_optimization_paths_PBE_2D
Cite this dataset Schmidt, J., Hoffmann, N., Wang, H., Borlido, P., Carriço, P. J. M. A., Cerqueira, T. F. T., Botti, S., and Marques, M. A. L. Alexandria geometry optimization paths PBE 2D. ColabFit, 2025. https://doi.org/10.60732/8781419f
This dataset has been curated and formatted for the ColabFit Exchange
This dataset is also available on the ColabFit Exchange:
https://materials.colabfit.org/id/DS_6pieq95jrqpn_0
Visit the ColabFit… See the full description on the dataset page: https://huggingface.co/datasets/colabfit/Alexandria_geometry_optimization_paths_PBE_2D.Alexandria_geometry_optimization_paths_PBE_3D
Cite this dataset Schmidt, J., Hoffmann, N., Wang, H., Borlido, P., Carriço, P. J. M. A., Cerqueira, T. F. T., Botti, S., and Marques, M. A. L. Alexandria geometry optimization paths PBE 3D. ColabFit, 2024. https://doi.org/10.60732/c88da7df
This dataset has been curated and formatted for the ColabFit Exchange
This dataset is also available on the ColabFit Exchange:
https://materials.colabfit.org/id/DS_s6gf4z2hcjqy_0
Visit the ColabFit… See the full description on the dataset page: https://huggingface.co/datasets/colabfit/Alexandria_geometry_optimization_paths_PBE_3D.Fino1_Reasoning_Path_FinQA_v2Fino1 is a financial reasoning dataset based on FinQA, with GPT-4o-generated reasoning paths to enhance structured financial question answering.
For more details, please check our paper arxiv.org/abs/2502.08127.
Source Data
Initial Data Collection and Normalization
The dataset originates from FinQA; TATQA, ConvFinQA, DocMath-Eval, DocFinQA, Bizbench dataset.
Annotations
Annotation Process
We add a prompt and create a reasoning process using GPT-4o… See the full description on the dataset page: https://huggingface.co/datasets/TheFinAI/Fino1_Reasoning_Path_FinQA_v2.Llama-slideQA-Sample-Featuresai-crawler-index
AI Crawler Index
150 web crawlers and AI user agents from 74 operators — what each one is for,
what blocking it costs you, and the IP ranges its operator publishes.
Plus a compiled user-agent regex and the union of 1997 IPv4 and 1062 IPv6
prefixes from 15 operator-published range files.
Home: https://www.pathwren.workers.dev/c/huggingface-datasets/ · CC0 · no signup, no key.
What this is, plainly
This is an independent, non-commercial automated project. It is run… See the full description on the dataset page: https://huggingface.co/datasets/pathwren/ai-crawler-index.PathoROB-camelyon
PathoROB
Preprint | Code | Licenses | Cite
PathoROB is a benchmark for the robustness of pathology foundation models (FMs) to non-biological medical center differences.
PathoROB contains four datasets covering 28 biological classes from 34 medical centers and three metrics:
Robustness Index: Measures the dominance of biological over non-biological features in an FM representation space.
Average Performance Drop (APD): Measures the robustness of downstream models to shortcut… See the full description on the dataset page: https://huggingface.co/datasets/bifold-pathomics/PathoROB-camelyon.PathoROB-tolkach_esca
PathoROB
Preprint | Code | Licenses | Cite
PathoROB is a benchmark for the robustness of pathology foundation models (FMs) to non-biological medical center differences.
PathoROB contains four datasets covering 28 biological classes from 34 medical centers and three metrics:
Robustness Index: Measures the dominance of biological over non-biological features in an FM representation space.
Average Performance Drop (APD): Measures the robustness of downstream models to shortcut… See the full description on the dataset page: https://huggingface.co/datasets/bifold-pathomics/PathoROB-tolkach_esca.pathfinder_arxiv_dataThis dataset is associated with the Pathfinder app (https://pfdr.app, paper at https://arxiv.org/abs/2408.01556) and is updated roughly monthly to keep pace with new literature in astrophysics and cosmology.
Refined-TCGA-PRAD-Prostate-Cancer-Pathology-Dataset
Note: Our 245 TCGA cases are ones we identified as having potential for improvement.
We plan to upload them in two phases: the first batch of 138 cases, and the second batch of 107 cases in the quality review pipeline, we plan to upload them around early of January, 2025.
Dataset: A Second Opinion on TCGA PRAD Prostate Dataset Labels with ROI-Level Annotations
Overview
This dataset provides enhanced Gleason grading annotations for the TCGA PRAD prostate cancer… See the full description on the dataset page: https://huggingface.co/datasets/Humanbased-AI/Refined-TCGA-PRAD-Prostate-Cancer-Pathology-Dataset.alexandria-pbe-geo-opt-paths-rawgreater-london-tx-proxy-rat-path-gain
Greater London Per-Transmitter-Proxy, Per-RAT Simulated Path-Gain Dataset
Short display name: Greater London Tx-Proxy × RAT Path-GainChinese name: 大伦敦逐发射代理、逐 RAT 模拟路径增益数据集Release: v9 final release (COMPLETE)
City-scale propagation, one transmitter proxy at a time.
This release turns Greater London into a queryable radio-propagation dataset:
449,437,201,731 simulated path-gain relations connect 22,678 computed
transmitter-proxy hypotheses with 171,549,960 receiver faces across… See the full description on the dataset page: https://huggingface.co/datasets/EEzim/greater-london-tx-proxy-rat-path-gain.alexandria-pbe-geo-opt-paths-sampledPathGen-shortest-pathplant-pathology-2021
Description
Dataset from the Plant Pathology 2021 (FGVC8) Challenge.
'
For Plant Pathology 2021-FGVC8, we have significantly increased the number of foliar disease images and added additional disease categories. This year’s dataset contains approximately 23,000 high-quality RGB images of apple foliar diseases, including a large expert-annotated disease dataset. This dataset reflects real field scenarios by representing non-homogeneous backgrounds of leaf images taken at different… See the full description on the dataset page: https://huggingface.co/datasets/wuhuhuhu123/plant-pathology-2021.alexandria-pbe-geo-opt-paths-raw-filteredPathGen-random-walk-10Mpath-vqa-robustnessPathMMUai2thor-path-tracing-qa-v5ai2thor_path_tracing_2point_tifa_filtered_evalxMIL-HeatmapsHeatmap data for the multiple instance learning models presented in:
Jamshidi Idaji et al. "Beyond attention heatmaps: How to get better explanations for multiple instance learning models in histopathology". Medical Image Analysis (2026).
Link: https://www.sciencedirect.com/science/article/pii/S1361841526002173
Code: https://github.com/bifold-pathomics/xMIL
@article{
jamshidi26beyond,
title = {Beyond attention heatmaps: How to get better explanations for multiple instance learning models… See the full description on the dataset page: https://huggingface.co/datasets/bifold-pathomics/xMIL-Heatmaps.rl_rl-conf_24GP_base-yaml_mode-path_r2eg-nl2b-stac-bugs-fixt_trai-data_exp_rpt_stac-php-largpath-vqa-refined
