datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
jetson1-pca-roi-doris081026
jetson1-pca-roi-doris081026
Recorded dataset — captured on jetson1 — 2 episodes · 293 frames @ 20 fps (~0 min of demonstration).
Tasks
Instruction
Episodes
place the cucumber on the middle of the cutting board, aligned vertically in the image, with the back end of the cucumber under the blue gripper
1
Recording
Rig
jetson1 (calibration sidecar)
Recorded
2026-08-10
Operator
dorischen
Episode sources
2 teleop… See the full description on the dataset page: https://huggingface.co/datasets/VibeCuisine/jetson1-pca-roi-doris081026.RoITrainingROI-1555_Rebar_Detection_and_Instance_Segmentation_DatasetROI-1555: Rebar Detection and Instance Segmentation Dataset
ROI-1555 for rebar object detection and instance segmentation contains 1555 rebar images and their fine-labeled bounding boxes and pixel-wise masks.
Diverse rebar specifications, layouts, application scenarios, and environmental conditions.
Usage
Here is an example to convert the annotations to MSCOCO 2017 format
python
cp -r 1260/img_label tools/data_annotated/train2017
cd tools
python labelme2coco_instance.py… See the full description on the dataset page: https://huggingface.co/datasets/tsrobcvai/ROI-1555_Rebar_Detection_and_Instance_Segmentation_Dataset.testdataVietPET-RoI
VietPET-RoI
VietPET-RoI is a Vietnamese whole-body PET/CT dataset containing paired
cropped 3D volumes, regional reports, and modality-specific 3D ROI bounding
boxes. It is intended for medical multimodal research, report generation,
visual question answering, and ROI grounding.
Research use only. This dataset is not intended for diagnosis, treatment
decisions, or direct patient care.
Summary
Split
Patients
CT/PET region pairs
ROIs
Train
160
480
1,544… See the full description on the dataset page: https://huggingface.co/datasets/scarlettlin/VietPET-RoI.store_play_v2_en_phase6_roiROI-1555_Rebar_Detection_and_Instance_Segmentation_DatasetROI-1555: Rebar Detection and Instance Segmentation Dataset
ROI-1555 for rebar object detection and instance segmentation contains 1555 rebar images and their fine-labeled bounding boxes and pixel-wise masks.
Diverse rebar specifications, layouts, application scenarios, and environmental conditions.
Usage
Here is an example to convert the annotations to MSCOCO 2017 format
python
cp -r 1260/img_label tools/data_annotated/train2017
cd tools
python… See the full description on the dataset page: https://huggingface.co/datasets/archispace/ROI-1555_Rebar_Detection_and_Instance_Segmentation_Dataset.store_play_v2_en_phase4_roiVietPET-RoI
VietPET-RoI
VietPET-RoI is a Vietnamese whole-body PET/CT dataset containing paired
cropped 3D volumes, regional reports, and modality-specific 3D ROI bounding
boxes. It is intended for medical multimodal research, report generation,
visual question answering, and ROI grounding.
Research use only. This dataset is not intended for diagnosis, treatment
decisions, or direct patient care.
Summary
Split
Patients
CT/PET region pairs
ROIs
Train
160
480
1,544… See the full description on the dataset page: https://huggingface.co/datasets/b00l26/VietPET-RoI.torchani-tests-pickled-filescdl-devai-results-ds003604-roiauditory
ds003604 (Wang et al. 2022) — brain × interpretability × localisation, per model per checkpoint
Wang et al. 2022 — auditory language comprehension in children. Cohort: children scanned at 5, 7 and 9; presentation: auditory.
Tasks Sem, Phon, Gram, Plaus × sessions ses-5, ses-7, ses-9 = 12 task × session cells,
all of them scored here.
Incomplete families --- do not read these as scale-ladder points. pythia-2.8b-full has 4 checkpoints, pythia-6.9b-full has 1 checkpoint.… See the full description on the dataset page: https://huggingface.co/datasets/BrainAlign/cdl-devai-results-ds003604-roiauditory.cdl-devai-results-ds003604-roiphonology
ds003604 (Wang et al. 2022) — brain × interpretability × localisation, per model per checkpoint
Wang et al. 2022 — auditory language comprehension in children. Cohort: children scanned at 5, 7 and 9; presentation: auditory.
Tasks Sem, Phon, Gram, Plaus × sessions ses-5, ses-7, ses-9 = 12 task × session cells,
all of them scored here.
Incomplete families --- do not read these as scale-ladder points. pythia-2.8b-full has 4 checkpoints. pythia-6.9b-full's single checkpoint is step… See the full description on the dataset page: https://huggingface.co/datasets/BrainAlign/cdl-devai-results-ds003604-roiphonology.cdl-devai-results-ds003604-roimotor
ds003604 (Wang et al. 2022) — brain × interpretability × localisation, per model per checkpoint
Wang et al. 2022 — auditory language comprehension in children. Cohort: children scanned at 5, 7 and 9; presentation: auditory.
Tasks Sem, Phon, Gram, Plaus × sessions ses-5, ses-7, ses-9 = 12 task × session cells,
all of them scored here.
Incomplete families --- do not read these as scale-ladder points. pythia-2.8b-full has 4 checkpoints. pythia-6.9b-full's single checkpoint is step… See the full description on the dataset page: https://huggingface.co/datasets/BrainAlign/cdl-devai-results-ds003604-roimotor.black_clothes260708_r1lite_roi_depth_gr00t_N17store_play_v2_en_phase5_roiwhite_blue_clothes_400llm-pretrain-fineweb-edu-2BTwhite_blue_clotheszetagpt-cot-countdown-game-20kwhite_black_clothesMedical-ROIs-K2.6
Medical-ROIs-K2.6
Medical visual grounding SFT data: for each clinical VQA sample, a teacher model proposes answer-supporting ROIs (regions of interest) as 2D bounding boxes.
Teacher: moonshotai/Kimi-K2.6.Upstream images & QA: MBZUAI/medix-rl-data.
How the data is generated
MBZUAI/medix-rl-data (train)
│
│ problem / solution / image / source / id
▼
Teacher: moonshotai/Kimi-K2.6
(vision + text; given question + gold answer)
│… See the full description on the dataset page: https://huggingface.co/datasets/erow/Medical-ROIs-K2.6.the-stack-smol-xs\blue_clothescollege-roi-data
College ROI Data — what U.S. colleges and majors actually pay back
Clean, citable tables on the lifetime financial return of U.S. colleges and majors —
30-year net present value by school and state, ROI by major category, the out-of-state
premium, and how exposed each major's career paths are to today's AI. Maintained by
LE TEEN, a college-ROI data project. Every number traces to a
public source; nothing is scraped, modeled behind closed doors, or vibes.
The headline the… See the full description on the dataset page: https://huggingface.co/datasets/le-teen/college-roi-data.260708_r1lite_roi_dist260708_r1lite_roi_depth_gr00tllm-ultrachat
Dataset Card for Dataset Name
Dataset Description
An open-source, large-scale, and multi-round dialogue data powered by Turbo APIs. In consideration of factors such as safeguarding privacy, we do not directly use any data available on the Internet as prompts.
To ensure generation quality, two separate ChatGPT Turbo APIs are adopted in generation, where one plays the role of the user to generate queries and the other generates the response.
We instruct the user… See the full description on the dataset page: https://huggingface.co/datasets/roisincrtai/llm-ultrachat.RoITD
Dataset Summary
We introduce a Romanian IT Dataset (RoITD) resembling SQuAD 1.1. RoITD consists of 9575 Romanian QA pairs formulated by crowd workers. QA pairs are based on 5043 articles from Romanian Wikipedia articles describing IT and household products. Of the total number of questions, 5103 are possible (i.e. the correct answer can be found within the paragraph) and 4472 are not possible (i.e. the given answer is a "plausible answer" and not correct)
Dataset… See the full description on the dataset page: https://huggingface.co/datasets/dragosnicolae555/RoITD.ellipsis-lrs3-B-roicore
