datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
MedQA-USMLE-4-optionsOriginal dataset introduced by Jin et al. in What Disease does this Patient Have? A Large-scale Open Domain Question Answering Dataset from Medical Exams
Citation information:
@article{jin2020disease,
title={What Disease does this Patient Have? A Large-scale Open Domain Question Answering Dataset from Medical Exams},
author={Jin, Di and Pan, Eileen and Oufattole, Nassim and Weng, Wei-Hung and Fang, Hanyi and Szolovits, Peter},
journal={arXiv preprint arXiv:2009.13081},
year={2020}
}
capstone_sakuga_preproc_optical_flowMedQA-USMLE-4-options-hfOriginal dataset introduced by Jin et al. in What Disease does this Patient Have? A Large-scale Open Domain Question Answering Dataset from Medical Exams
Citation information:
@article{jin2020disease,
title={What Disease does this Patient Have? A Large-scale Open Domain Question Answering Dataset from Medical Exams},
author={Jin, Di and Pan, Eileen and Oufattole, Nassim and Weng, Wei-Hung and Fang, Hanyi and Szolovits, Peter},
journal={arXiv preprint arXiv:2009.13081},
year={2020}
}
3d_optical_flow_droid
3D Optical Flow DROID Dataset
Processed DROID robotics dataset with optical flow and scene flow annotations.
Dataset Structure
Organized by lab, each trajectory in separate tar.gz archive:
IPRL/IPRL+2023-06-19+Mon_Jun_19_23:27:48_2023.tar.gz
CLVR/CLVR+2023-...tar.gz
... (15 labs, ~33K trajectories)
Each trajectory contains:
metadata.json - Trajectory metadata
trajectory.h5 - Robot state and actions
camera_left/, camera_right/ - Camera data
rgb/ - RGB images
depth/ -… See the full description on the dataset page: https://huggingface.co/datasets/Salesforce/3d_optical_flow_droid.cpuClimbLabClimbLab is a high-quality pre-training corpus released by NVIDIA. Here is the description:
ClimbLab is a filtered 1.2-trillion-token corpus with 20 clusters.
Based on Nemotron-CC and SmolLM-Corpus, we employed our proposed CLIMB-clustering to semantically reorganize and filter this combined dataset into 20 distinct clusters, leading to a 1.2-trillion-token high-quality corpus. Specifically, we first grouped the data into 1,000 groups based on topic information. Then we applied two… See the full description on the dataset page: https://huggingface.co/datasets/OptimalScale/ClimbLab.india-index-options-1m
India Index & Options - 1-minute OHLC
1-minute OHLCV(+OI) bars for NSE/BSE index spot and option chains: NIFTY, BANKNIFTY, SENSEX (~2021-2026).
Powers the open-source TradeMarkk backtester (https://thetrademarkk.com).
Educational use only. Provided as-is, no warranty. Verify against official exchange data before relying on it.
Structure
index/{SYMBOL}.parquet - 1-min spot OHLC per index.
options/{SYMBOL}/{EXPIRY}.parquet - 1-min OHLC per option contract (with… See the full description on the dataset page: https://huggingface.co/datasets/thetrademarkk/india-index-options-1m.documentation-imagesThis dataset contains images used in the documentation of HuggingFace's Optimum library.
pinnacle-optuna-dbllm-perf-leaderboardalpaca-options-datacudaClimbMixClimbMix is a high-quality pre-training corpus released by NVIDIA. Here is the description:
ClimbMix is a compact yet powerful 400-billion-token dataset designed for efficient pre-training that delivers superior performance under an equal token budget. It was introduced in this paper.
We proposed a new algorithm to filter and mix the dataset. First, we grouped the data into 1,000 groups based on topic information. Then we applied two classifiers: one to detect advertisements and another to… See the full description on the dataset page: https://huggingface.co/datasets/OptimalScale/ClimbMix.UnlearnCanvas
Dataset Card for UnlearnCanvas
This dataset card introduces "UnlearnCanvas", a high-resolution stylized image dataset for benchmarking generative modeling tasks, in particular for machine unlearning in diffusion models. Developed to address the societal concerns arising from diffusion models, such as harmful content generation, copyright disputes, and the perpetuation of stereotypes and biases, UnlearnCanvas aims at facilitating the evaluation and improvement of machine unlearning… See the full description on the dataset page: https://huggingface.co/datasets/OPTML-Group/UnlearnCanvas.transformers_daily_citransformers_pr_cirocmagentic-ai-options-resultscrypto-options-surface
Crypto options surface and order books
Snapshots of listed options on Aevo, including marks, implied volatility, Greeks and a selected set of order books. The tables support analysis of volatility surfaces and the relationship between published marks and quoted prices.
Contents
Table
Record
options_surface
An instrument's strike, expiry, mark, forward, implied volatility and Greeks
options_order_book
Best bid and ask, available sizes and quoted… See the full description on the dataset page: https://huggingface.co/datasets/dataforge-labs/crypto-options-surface.optiq-lab-traces
OptiQ Lab Traces
Research and tool-calling sessions produced by OptiQ Lab, the local web UI that ships with mlx-optiq. Each session is a complete run: a deep-research report built from live web sources, or a multi-turn agent loop driving the Lab's own sandboxed tools.
The dataset is 866 sessions in HuggingFace Session-Traces format (the agent-traces viewer). Each .jsonl file is one session: a header line carrying the run's metadata, then one message per turn.
The two… See the full description on the dataset page: https://huggingface.co/datasets/mlx-community/optiq-lab-traces.Optical-SAR-Infrared
MMDiff: Multi-modal Remote Sensing Image Generation via Cross-Modality Spatial Feature Transfer
ISPRS 2026 🔥
Haojun Tang1 · Wenda Zhao1,* · Hengshuai Cui1 · Haipeng Wang2
1 Dalian University of Technology2 Unit 92728 of PLA
* Corresponding author:
Abstract
Collecting spatially consistent multi-modal remote sensing (MMRS) images remains challenging due to different sensors vary in the imaging principles and acquisition times. This hinders… See the full description on the dataset page: https://huggingface.co/datasets/XinRan-Tang/Optical-SAR-Infrared.LLM_opt_backup_406OpticalRS-13M
Harnessing Massive Satellite Imagery with Efficient Masked Image Modeling
Fengxiang Wang1
Hongzhen Wang2,‡
Di Wang3
Zonghao Guo2
Zhenyu Zhong4
Long Lan1,‡
Wenjing Yang1,‡
Jing Zhang3,‡
1 National University of Defense Technology
2Tsinghua University
3Wuhan University
4Nankai University
ICCV 2025
📃 Paper |
🤗 OpticalRS-4M|
🤗 OpticalRS-13M |
🤗 Models
🎯Intruduction… See the full description on the dataset page: https://huggingface.co/datasets/initiacms/OpticalRS-13M.hle-multimodalhosts100-labelled-optcphi-winogrande_inverted_option-results
Dataset Card for "phi-winogrande_inverted_option-results"
More Information needed
Alexandria_geometry_optimization_paths_PBE_3D
Cite this dataset Schmidt, J., Hoffmann, N., Wang, H., Borlido, P., Carriço, P. J. M. A., Cerqueira, T. F. T., Botti, S., and Marques, M. A. L. Alexandria geometry optimization paths PBE 3D. ColabFit, 2024. https://doi.org/10.60732/c88da7df
This dataset has been curated and formatted for the ColabFit Exchange
This dataset is also available on the ColabFit Exchange:
https://materials.colabfit.org/id/DS_s6gf4z2hcjqy_0
Visit the ColabFit… See the full description on the dataset page: https://huggingface.co/datasets/colabfit/Alexandria_geometry_optimization_paths_PBE_3D.Alexandria_geometry_optimization_paths_PBE_2D
Cite this dataset Schmidt, J., Hoffmann, N., Wang, H., Borlido, P., Carriço, P. J. M. A., Cerqueira, T. F. T., Botti, S., and Marques, M. A. L. Alexandria geometry optimization paths PBE 2D. ColabFit, 2025. https://doi.org/10.60732/8781419f
This dataset has been curated and formatted for the ColabFit Exchange
This dataset is also available on the ColabFit Exchange:
https://materials.colabfit.org/id/DS_6pieq95jrqpn_0
Visit the ColabFit… See the full description on the dataset page: https://huggingface.co/datasets/colabfit/Alexandria_geometry_optimization_paths_PBE_2D.oercommons-v1-optimized
OERCommons v1 Optimized
Authors: Junjie Wang and Yuhan SunHosted by: PIN TeamDataset: pin-team/oercommons-v1-optimized
OERCommons v1 Optimized is a provenance-preserving multimodal pretraining corpus built on the OERCommons subset of The Common Pile v0.1, which serves as its upstream data and licensing baseline. We extend it with full-page recovery, canonical Markdown, ordered image/PDF/link metadata, conservative corrections, and integrity evidence.
At a glance… See the full description on the dataset page: https://huggingface.co/datasets/pin-team/oercommons-v1-optimized.optimus-dataset
OPTIMUS Dataset
This dataset contains approximately 600K image time series of 40-50 Sentinel-2 satellite images captured between January 2016 and December 2023.
It also includes 300 time series that are labeled with binary "change" or "no change" labels.
It is used to train and evaluate OPTIMUS (https://arxiv.org/abs/2506.13902).
The time series are distributed globally, with half of the time series selected at random locations covered by Sentinel-2, and the other half sampled… See the full description on the dataset page: https://huggingface.co/datasets/optimus-change/optimus-dataset.
