datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Java-GitHub-CodesMPM-Verse-MaterialSim-Small
Dataset Card for MPMVerse Physics Simulation Dataset
Dataset Summary
This dataset contains Material-Point-Method (MPM) simulations for various materials, including water, sand, plasticine, elasticity, jelly, rigid collisions, and melting. Each material is represented as point-clouds that evolve over time. The dataset is designed for learning and predicting MPM-based physical simulations.
Supported Tasks and Leaderboards
The dataset supports tasks such as:… See the full description on the dataset page: https://huggingface.co/datasets/hrishivish23/MPM-Verse-MaterialSim-Small.PerceptionComp
PerceptionComp: A Benchmark for Complex Perception-Centric Video Reasoning
PerceptionComp is a benchmark for complex perception-centric video reasoning. It focuses on questions that cannot be solved from a single frame, a short clip, or a shallow caption. Models must revisit visually complex videos, gather evidence across temporally separated segments, and combine multiple perceptual cues before answering.
Dataset Details
Dataset… See the full description on the dataset page: https://huggingface.co/datasets/hrinnnn/PerceptionComp.MPM-Verse-MaterialSim-Large
MPM-Verse-MaterialSim-Large
Dataset Summary
This dataset contains Material-Point-Method (MPM) simulations for various materials, including water, sand, plasticine, and jelly.
Each material is represented as point-clouds that evolve over time. The dataset is designed for learning and predicting MPM-based
physical simulations. The dataset is rendered using five geometric models - Stanford-bunny, Spot, Dragon, Armadillo, and Blub.
Each setting has 10 trajectories per… See the full description on the dataset page: https://huggingface.co/datasets/hrishivish23/MPM-Verse-MaterialSim-Large.lerobot_hriwikitext-tags-deberta-basewikitext-tags-modernbertwikitext-tags-robertahr-interview-datasetwikitext-tags-deberta-v3YT-100K
A larger version of YT-100K dataset -> YT-30M dataset with 30 million YouTube multilingual multicategory comments is available here: YTCommentVerse
Bibtex
@inproceedings{dutta2025ytcommentverse,
title={YTCommentVerse: A Multi-Category Multi-Lingual YouTube Comment Corpus},
author={Dutta, Hridoy Sankar and Khan, Biswadeep},
booktitle={Proceedings of the 34th ACM International Conference on Information and Knowledge Management},
pages={6351--6355},
year={2025}
}… See the full description on the dataset page: https://huggingface.co/datasets/hridaydutta123/YT-100K.sciercAI_dataso101_sra_hri_tasks_2camThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "so101_follower",
"total_episodes": 48,
"total_frames": 16577,
"total_tasks": 1,
"total_videos": 96,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:48"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/cchenalds17/so101_sra_hri_tasks_2cam.HR-Instruct-Math-v0.1
Dataset Summary
HAERAE-HUB/HR-Instruct-Math-v0.1 is a Math instruction dataset written in the Korean language. This dataset contains evolved instructions aimed at enhancing the learning experience in mathematical concepts. The responses in this dataset are generated from open-source Language Models (LLMs). This is a Proof of Concept (PoC) version, meaning there may be errors or unexpected problems in the dataset. Future iterations will be made to improve the dataset quality.… See the full description on the dataset page: https://huggingface.co/datasets/HAERAE-HUB/HR-Instruct-Math-v0.1.metal_block_box_correct_cameraThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "custom_manipulator",
"total_episodes": 10,
"total_frames": 1395,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 10,
"splits": {
"train": "0:10"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/HSP-HRII/metal_block_box_correct_camera.hrithik_roshan_imagesTUC-HRI-CS
University of Technology Chemnitz, Germany
Department Robotics and Human Machine Interaction
Author: Robert Schulz
TUC-HRI Dataset Card
TUC-AR is an action recognition dataset, containing 10(+1) action categories for human machine interaction. This version contains video sequences, stored as images, frame by frame.
We introduce two validation types: random validation and cross-subject validation. This is the cross-subject validation dataset. For random validation, please use… See the full description on the dataset page: https://huggingface.co/datasets/SchulzR97/TUC-HRI-CS.acl-arcAll_Puzzles_5k_New_Context_Hritikeval_groot_metal_block_boxThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "custom_manipulator",
"total_episodes": 10,
"total_frames": 1436,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 10,
"splits": {
"train": "0:10"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/HSP-HRII/eval_groot_metal_block_box.TUC-HRI
University of Technology Chemnitz, Germany
Department Robotics and Human Machine Interaction
Author: Robert Schulz
TUC-HRI Dataset Card
TUC-AR is an action recognition dataset, containing 10(+1) action categories for human machine interaction. This version contains video sequences, stored as images, frame by frame.
We introduce two validation types: random validation and cross-subject validation. This is the random validation dataset. For cross-subject validation, please use… See the full description on the dataset page: https://huggingface.co/datasets/SchulzR97/TUC-HRI.so101_sra_hri_tasks_3camThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "so101_follower",
"total_episodes": 48,
"total_frames": 16577,
"total_tasks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:48"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path": "videos/{video_key}/chunk-{chunk_index:03d}/file-{file_index:03d}.mp4"… See the full description on the dataset page: https://huggingface.co/datasets/cchenalds17/so101_sra_hri_tasks_3cam.HRII_test_pick_obj_with_mi_centroidsThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "custom_manipulator",
"total_episodes": 50,
"total_frames": 4288,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 500,
"fps": 10,
"splits": {
"train": "0:50"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/lucazanett/HRII_test_pick_obj_with_mi_centroids.metal_block_boxThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "custom_manipulator",
"total_episodes": 51,
"total_frames": 6335,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 10,
"splits": {
"train": "0:51"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/HSP-HRII/metal_block_box.h_rings_peg3This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "so101_follower",
"total_episodes": 40,
"total_frames": 26935,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 500,
"fps": 30,
"splits": {
"train": "0:40"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/bogdan2435/h_rings_peg3.autotrain-data-meme-classification
AutoTrain Dataset for project: meme-classification
Dataset Description
This dataset has been automatically processed by AutoTrain for project meme-classification.
Languages
The BCP-47 code for the dataset's language is unk.
Dataset Structure
Data Instances
A sample from this dataset looks as follows:
[
{
"image": "<657x657 RGB PIL image>",
"target": 1
},
{
"image": "<1124x700 RGB PIL image>",
"target": 0
}]… See the full description on the dataset page: https://huggingface.co/datasets/Hrishikesh332/autotrain-data-meme-classification.syntaxgym-hexatagged
Dataset Card for "syntaxgym-hexatagged"
More Information needed
blimp-hexatagged-incremental
configs:
eval_act_metal_block_boxThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "custom_manipulator",
"total_episodes": 10,
"total_frames": 2578,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 10,
"splits": {
"train": "0:10"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/HSP-HRII/eval_act_metal_block_box.
