R.O.C
Datasets
All datasets matching “R.O.C”OmniReasoner-SFT
OmniReasoner-SFT
OmniReasoner-SFT is a mixed-source, research-only supervised fine-tuning dataset
for audio-visual and long-video reasoning. It contains two-stage cold-start SFT
trajectories with interval selection, zoom-in evidence, and final answers.
Contents
data/train.jsonl: HF-ready training JSONL with repo-relative media paths.
media/: raw and derived media referenced by train.jsonl.
manifests/media_manifest.jsonl: media inventory with repo paths, source
family… See the full description on the dataset page: https://huggingface.co/datasets/Rocky131/OmniReasoner-SFT.rockwaladynotashinamideshite
Bangumi Image Base of Rock Wa Lady No Tashinami Deshite
This is the image base of bangumi Rock wa Lady no Tashinami deshite, we detected 104 characters, 7595 images in total. The full dataset is here.
Please note that these image bases are not guaranteed to be 100% cleaned, they may be noisy actual. If you intend to manually train models using this dataset, we recommend performing necessary preprocessing on the downloaded dataset to eliminate potential noisy samples… See the full description on the dataset page: https://huggingface.co/datasets/BangumiBase/rockwaladynotashinamideshite.rocochallenge2025
Dataset Description
This dataset serves for the roco challenge @AAAI26, focusing on collaborative gearbox assembly scenarios in human-centric manufacturing. It contains demonstrations for the gearbox assembly task, conducted with Galaxea R1 robot in IsaacSim environment. The code for data collection is given here.
Homepage: https://rocochallenge.github.io/RoCo2026/
Dataset Structure
The dataset contains RGBD observation from 3 views (head, left hand, right hand)… See the full description on the dataset page: https://huggingface.co/datasets/rocochallenge2025/rocochallenge2025.ROCOv2-radiology
ROCOv2: Radiology Object in COntext version 2
Introduction
ROCOv2 is a multimodal dataset consisting of radiological images and associated medical concepts and captions extracted from the PMC Open Access Subset. It is an updated version of the ROCO dataset, adding 35,705 new images and improving concept extraction and filtering.
Dataset Overview
The ROCOv2 dataset contains 79,789 radiological images, each with a corresponding caption and medical concepts. The… See the full description on the dataset page: https://huggingface.co/datasets/eltorio/ROCOv2-radiology.pile-of-rocq
theostos/pile-of-rocq
Pile-of-Rocq exported as normalized parquet tables with docstring + env_toc.
Each config is <env>-<table> and loads one parquet table for one env.
Load examples
from datasets import load_dataset
toc = load_dataset('theostos/pile-of-rocq', 'coq-actuary-toc_nodes', split='train')
steps = load_dataset('theostos/pile-of-rocq', 'coq-actuary-proof_steps', split='train')
print(len(toc), len(steps))
Environments in this export
count: 48… See the full description on the dataset page: https://huggingface.co/datasets/theostos/pile-of-rocq.rocm
