datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
pack_toothbrush_Nov19-advantages
Advantage Values for villekuosmanen/pack_toothbrush_Nov19
Pre-computed advantage values for offline RL training.
Source
Dataset: villekuosmanen/pack_toothbrush_Nov19
Value Model: villekuosmanen/rewact_toothbrush_pistar_1.5.0
N-step lookahead: 50
Files
This dataset contains per-episode parquet files with advantage values for each frame.
Usage
from pathlib import Path
import pandas as pd
# Load advantages for a specific episode
advantage_df =… See the full description on the dataset page: https://huggingface.co/datasets/villekuosmanen/pack_toothbrush_Nov19-advantages.dAgger_pack_toothbrush_Nov26-advantages
Advantage Values for villekuosmanen/dAgger_pack_toothbrush_Nov26
Pre-computed advantage values for offline RL training.
Source
Dataset: villekuosmanen/dAgger_pack_toothbrush_Nov26
Value Model: villekuosmanen/rewact_toothbrush_pistar_1.5.0
N-step lookahead: 50
Files
This dataset contains per-episode parquet files with advantage values for each frame.
Usage
from pathlib import Path
import pandas as pd
# Load advantages for a specific episode… See the full description on the dataset page: https://huggingface.co/datasets/villekuosmanen/dAgger_pack_toothbrush_Nov26-advantages.dAgger_pack_toothbrush_Nov22-advantages
Advantage Values for villekuosmanen/dAgger_pack_toothbrush_Nov22
Pre-computed advantage values for offline RL training.
Source
Dataset: villekuosmanen/dAgger_pack_toothbrush_Nov22
Value Model: villekuosmanen/rewact_toothbrush_pistar_1.5.0
N-step lookahead: 50
Files
This dataset contains per-episode parquet files with advantage values for each frame.
Usage
from pathlib import Path
import pandas as pd
# Load advantages for a specific episode… See the full description on the dataset page: https://huggingface.co/datasets/villekuosmanen/dAgger_pack_toothbrush_Nov22-advantages.STS-3D-Tooth
STS-3D-Tooth
The 3D Cone-Beam CT (CBCT) subset of the STS (Semi-supervised Teeth
Segmentation) multi-modal dental dataset, as released in
Wang et al., Scientific Data 12, 117 (2025)
and used in the MICCAI 2023/2024 STS Challenges.
The companion 2D panoramic X-ray subset is hosted at
Angelou0516/STS-2D-Tooth.
Dataset Summary
Field
Details
Modality
Cone-Beam CT (CBCT), NIfTI (.nii.gz)
Body Part
Teeth (32 permanent teeth, FDI numbering)
Volumes
371… See the full description on the dataset page: https://huggingface.co/datasets/MedOtter/STS-3D-Tooth.pack_toothbrush_Nov26-advantages
Advantage Values for villekuosmanen/pack_toothbrush_Nov26
Pre-computed advantage values for offline RL training.
Source
Dataset: villekuosmanen/pack_toothbrush_Nov26
Value Model: villekuosmanen/rewact_toothbrush_pistar_1.5.0
N-step lookahead: 50
Files
This dataset contains per-episode parquet files with advantage values for each frame.
Usage
from pathlib import Path
import pandas as pd
# Load advantages for a specific episode
advantage_df =… See the full description on the dataset page: https://huggingface.co/datasets/villekuosmanen/pack_toothbrush_Nov26-advantages.toothbrush-v2-dataset
Toothbrushing Detection Dataset (v2)
Video and image data for detecting toothbrushing behavior, collected for a Raspberry Pi Zero 2W toothbrush-detection project (toothbrush_v2). A single-class object detector is trained on this data to output [x, y, w, h, confidence] for the toothbrush in frame.
Dataset structure
Files are packed into tar shards (rather than uploaded individually) to stay within the Hub's per-repo file-count guidelines. To reconstruct the… See the full description on the dataset page: https://huggingface.co/datasets/Schrodin-purrrrr/toothbrush-v2-dataset.dAgger_pack_toothbrush_Nov30-advantages
Advantage Values for villekuosmanen/dAgger_pack_toothbrush_Nov30
Pre-computed advantage values for offline RL training.
Source
Dataset: villekuosmanen/dAgger_pack_toothbrush_Nov30
Value Model: villekuosmanen/rewact_toothbrush_pistar_1.5.0
N-step lookahead: 50
Files
This dataset contains per-episode parquet files with advantage values for each frame.
Usage
from pathlib import Path
import pandas as pd
# Load advantages for a specific episode… See the full description on the dataset page: https://huggingface.co/datasets/villekuosmanen/dAgger_pack_toothbrush_Nov30-advantages.lfv-raw-toothbrush-in-cup
LFV raw recording: Toothbrush_In_Cup
Public raw robot-teleoperation recording archived by the LFV TUI.
Episodes: 32
Rosbag files: 32
Integrated-health BAD episodes: 0
Camera/action contract: config/lfv_camera_contract.json
Download on a local or cloud training machine:
hf auth login
hf download Onol/lfv-raw-toothbrush-in-cup --repo-type dataset \
--local-dir ./rosbag_data/Toothbrush_In_Cup
The detailed archive receipt is in lfv_hf_upload_manifest.json. Episode-level… See the full description on the dataset page: https://huggingface.co/datasets/Onol/lfv-raw-toothbrush-in-cup.grab_tooth_brushdAgger_pack_toothbrush_Nov28-advantages
Advantage Values for villekuosmanen/dAgger_pack_toothbrush_Nov28
Pre-computed advantage values for offline RL training.
Source
Dataset: villekuosmanen/dAgger_pack_toothbrush_Nov28
Value Model: villekuosmanen/rewact_toothbrush_pistar_1.5.0
N-step lookahead: 50
Files
This dataset contains per-episode parquet files with advantage values for each frame.
Usage
from pathlib import Path
import pandas as pd
# Load advantages for a specific episode… See the full description on the dataset page: https://huggingface.co/datasets/villekuosmanen/dAgger_pack_toothbrush_Nov28-advantages.pack_toothbrush_Nov19This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "arx5",
"total_episodes": 30,
"total_frames": 23721,
"total_tasks": 1,
"total_videos": 90,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 20,
"splits": {
"train": "0:30"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/villekuosmanen/pack_toothbrush_Nov19.tooth_extraction_4This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.0",
"robot_type": "so100",
"total_episodes": 200,
"total_frames": 76053,
"total_tasks": 1,
"total_videos": 400,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:200"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/samsam0510/tooth_extraction_4.ToothXpert.MM-OPG-Annotationstooth_extraction_3This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.0",
"robot_type": "so100",
"total_episodes": 100,
"total_frames": 32879,
"total_tasks": 1,
"total_videos": 200,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:100"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/samsam0510/tooth_extraction_3.toothbrushThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": null,
"total_episodes": 2,
"total_frames": 785,
"total_tasks": 1,
"total_videos": 10,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 33,
"splits": {
"train": "0:2"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/Jinyu220/toothbrush.pickup_toothpicks_2_plus_recovery_bboxesSTS-2D-Tooth
STS-2D-Tooth
The 2D panoramic dental X-ray subset of the STS (Semi-supervised Teeth
Segmentation) multi-modal dataset, as released in
Wang et al., Scientific Data 12, 117 (2025)
and used in the MICCAI 2023 STS Challenge.
Composition
4,000 panoramic X-ray images (PNG, 640x320, 3-channel grayscale-as-RGB) split
across two demographic subsets:
Subset
Total
Labeled
Unlabeled
A-PXI (adult)
3,500
850
2,650
C-PXI (child)
500
50
450
Total
4,000
900
3,100… See the full description on the dataset page: https://huggingface.co/datasets/MedOtter/STS-2D-Tooth.toothbruh_bboxes
toothbruh
This dataset was generated using a phospho starter pack.
This dataset contains a series of episodes recorded with a robot and multiple cameras. It can be directly used to train a policy using imitation learning. It's compatible with LeRobot and RLDS.
dAgger_pack_toothbrush_Dec2This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "arx5",
"total_episodes": 30,
"total_frames": 32477,
"total_tasks": 1,
"total_videos": 90,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 20,
"splits": {
"train": "0:30"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/villekuosmanen/dAgger_pack_toothbrush_Dec2.toothbrush_cup_organizationffw_bg2_rev4_pick_tooth_brushThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "ffw_bg2_rev4",
"total_episodes": 30,
"total_frames": 4406,
"total_tasks": 1,
"total_videos": 90,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 10,
"splits": {
"train": "0:30"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/Dongkkka/ffw_bg2_rev4_pick_tooth_brush.dAgger_pack_toothbrush_Dec2-advantages
Advantage Values for villekuosmanen/dAgger_pack_toothbrush_Dec2
Pre-computed advantage values for offline RL training.
Source
Dataset: villekuosmanen/dAgger_pack_toothbrush_Dec2
Value Model: villekuosmanen/rewact_toothbrush_pistar_1.5.0
N-step lookahead: 50
Files
This dataset contains per-episode parquet files with advantage values for each frame.
Usage
from pathlib import Path
import pandas as pd
# Load advantages for a specific episode… See the full description on the dataset page: https://huggingface.co/datasets/villekuosmanen/dAgger_pack_toothbrush_Dec2-advantages.dAgger_pack_toothbrush_Nov28This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "arx5",
"total_episodes": 30,
"total_frames": 35163,
"total_tasks": 1,
"total_videos": 90,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 20,
"splits": {
"train": "0:30"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/villekuosmanen/dAgger_pack_toothbrush_Nov28.dAgger_pack_toothbrush_Nov30This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "arx5",
"total_episodes": 30,
"total_frames": 32896,
"total_tasks": 1,
"total_videos": 90,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 20,
"splits": {
"train": "0:30"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/villekuosmanen/dAgger_pack_toothbrush_Nov30.dAgger_pack_toothbrush_Nov22This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "arx5",
"total_episodes": 3,
"total_frames": 3261,
"total_tasks": 1,
"total_videos": 9,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 20,
"splits": {
"train": "0:3"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/villekuosmanen/dAgger_pack_toothbrush_Nov22.pack_toothbrushThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.0",
"robot_type": "arx5",
"total_episodes": 15,
"total_frames": 13845,
"total_tasks": 1,
"total_videos": 45,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 20,
"splits": {
"train": "0:15"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/villekuosmanen/pack_toothbrush.toothbrush_lfThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": null,
"total_episodes": 10,
"total_frames": 2778,
"total_tasks": 1,
"total_videos": 20,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 33,
"splits": {
"train": "0:10"},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/iantc104/toothbrush_lf.ToothFairytask_table30v2_pack_the_toothbrush_holdertooth_classification
