opt
Datasets
All datasets matching “opt”MedQA-USMLE-4-optionsOriginal dataset introduced by Jin et al. in What Disease does this Patient Have? A Large-scale Open Domain Question Answering Dataset from Medical Exams
Citation information:
@article{jin2020disease,
title={What Disease does this Patient Have? A Large-scale Open Domain Question Answering Dataset from Medical Exams},
author={Jin, Di and Pan, Eileen and Oufattole, Nassim and Weng, Wei-Hung and Fang, Hanyi and Szolovits, Peter},
journal={arXiv preprint arXiv:2009.13081},
year={2020}
}
capstone_sakuga_preproc_optical_flowMedQA-USMLE-4-options-hfOriginal dataset introduced by Jin et al. in What Disease does this Patient Have? A Large-scale Open Domain Question Answering Dataset from Medical Exams
Citation information:
@article{jin2020disease,
title={What Disease does this Patient Have? A Large-scale Open Domain Question Answering Dataset from Medical Exams},
author={Jin, Di and Pan, Eileen and Oufattole, Nassim and Weng, Wei-Hung and Fang, Hanyi and Szolovits, Peter},
journal={arXiv preprint arXiv:2009.13081},
year={2020}
}
3d_optical_flow_droid
3D Optical Flow DROID Dataset
Processed DROID robotics dataset with optical flow and scene flow annotations.
Dataset Structure
Organized by lab, each trajectory in separate tar.gz archive:
IPRL/IPRL+2023-06-19+Mon_Jun_19_23:27:48_2023.tar.gz
CLVR/CLVR+2023-...tar.gz
... (15 labs, ~33K trajectories)
Each trajectory contains:
metadata.json - Trajectory metadata
trajectory.h5 - Robot state and actions
camera_left/, camera_right/ - Camera data
rgb/ - RGB images
depth/ -… See the full description on the dataset page: https://huggingface.co/datasets/Salesforce/3d_optical_flow_droid.cpuClimbLabClimbLab is a high-quality pre-training corpus released by NVIDIA. Here is the description:
ClimbLab is a filtered 1.2-trillion-token corpus with 20 clusters.
Based on Nemotron-CC and SmolLM-Corpus, we employed our proposed CLIMB-clustering to semantically reorganize and filter this combined dataset into 20 distinct clusters, leading to a 1.2-trillion-token high-quality corpus. Specifically, we first grouped the data into 1,000 groups based on topic information. Then we applied two… See the full description on the dataset page: https://huggingface.co/datasets/OptimalScale/ClimbLab.
