datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
construction-traversability-dataset
Construction Traversability Dataset
A construction-site RGB semantic segmentation dataset developed for research on
terrain understanding, traversability estimation, and multimodal RGB–LiDAR
perception for mobile robots.
Overview
This dataset contains RGB images and pixel-wise semantic segmentation masks
collected in construction-site environments. The dataset is intended to support
research on construction-site scene understanding and traversability-aware
robot… See the full description on the dataset page: https://huggingface.co/datasets/manojkarnekar/construction-traversability-dataset.speech-enhancement-dfn-16ksweep_manoThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "human",
"total_episodes": 498,
"total_frames": 96415,
"total_tasks": 1,
"total_videos": 0,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 50,
"splits": {
"train": "0:498"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path": null,
"features": {… See the full description on the dataset page: https://huggingface.co/datasets/Jgold90/sweep_mano.sweep_mano_recompileThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "human",
"total_episodes": 498,
"total_frames": 96415,
"total_tasks": 1,
"total_videos": 0,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 50,
"splits": {
"train": "0:498"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path": null,
"features": {… See the full description on the dataset page: https://huggingface.co/datasets/Jgold90/sweep_mano_recompile.ASVspoof2021_DF
ASVspoof 2021 DF
Benchmark-ready packaging of the DeepFake (DF) evaluation partition from ASVspoof 2021 for speech anti-spoofing and synthetic / deepfake voice detection.
Overview
This dataset contains the DF evaluation subset of the ASVspoof 2021 challenge. The task is binary classification: bonafide (genuine human speech) vs. spoof (synthetic, converted, or otherwise manipulated speech). The original dataset is available at… See the full description on the dataset page: https://huggingface.co/datasets/manojkumarainala152/ASVspoof2021_DF.mano_feedbacknoisev2construction-rosbags
Construction-Site Multimodal Robotics ROS 2 Recordings
Construction-Site Multimodal Robotics ROS 2 Recordings
This repository contains the raw ROS 2 MCAP recordings collected from construction-site environments for the accompanying study on failure-mode-aware traversability mapping for autonomous mobile robots (AMRs).
The recordings preserve the multimodal sensor streams and their temporal relationships during real robot operation. All four recording sessions are… See the full description on the dataset page: https://huggingface.co/datasets/manojkarnekar/construction-rosbags.facial_emotion_detection_dataset
Face Emotion Classification Dataset
This dataset contain about 35000 images which are belongs to 7 classes. This dataset can be used to train deep learning models for human emotion classification problems.
big_patent
Dataset Card for Big Patent
Dataset Summary
BIGPATENT, consisting of 1.3 million records of U.S. patent documents along with human written abstractive summaries.
Each US patent application is filed under a Cooperative Patent Classification (CPC) code.
There are nine such classification categories:
a: Human Necessities
b: Performing Operations; Transporting
c: Chemistry; Metallurgy
d: Textiles; Paper
e: Fixed Constructions
f: Mechanical Engineering; Lightning; Heating;… See the full description on the dataset page: https://huggingface.co/datasets/manoj8890/big_patent.ca-sco-properties
CA SCO Unclaimed Property
California State Controller's Office unclaimed property records.
Bucketed into alphabetical splits by owner last name first letter
so every split stays under HF's 5 GB filter-index limit.
Updated weekly via GitHub Actions.
code-audio-video
Code Audio Video Data Notes
Dataset summary
This repository contains a preparation pipeline and a small metadata sample for Code work with Audio Video inputs. It does not claim to be a complete benchmark release; the loader documents how source data is normalized and validated.
Included material
load_data.py — loading, cleaning, and split preparation code.
dataset_infos.json — schema and split metadata.
metadata_sample.jsonl — small, human-readable… See the full description on the dataset page: https://huggingface.co/datasets/mano-jpmp/code-audio-video.lingbot_robotool_knife_spread_tomatosauce_mano_all_lerobotsweep_mano_newThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "human",
"total_episodes": 10,
"total_frames": 1822,
"total_tasks": 1,
"total_videos": 0,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 50,
"splits": {
"train": "0:10"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path": null,
"features": {
"lang":… See the full description on the dataset page: https://huggingface.co/datasets/Jgold90/sweep_mano_new.indic-diarbench
Indic DiarBench
A multilingual joint diarization and ASR benchmark for Indian languages, spanning all 22 scheduled languages of India with approximately 108 hours of natural multi-speaker audio.
Dataset Summary
Indic DiarBench is a conversational speech benchmark designed to evaluate speaker-attributed ASR in realistic multi-speaker settings for Indian languages. All annotations are human-corrected with time-aligned, speaker-attributed transcriptions. The dataset… See the full description on the dataset page: https://huggingface.co/datasets/manojkumarcs/indic-diarbench.INCLUDE_Dataset
🤟 INCLUDE: A Large-Scale Dataset for Indian Sign Language Recognition
A comprehensive video dataset for Indian Sign Language Recognition
🌐 Computer Vision •
🧠 Deep Learning •
🎥 Video Recognition •
🤟 Sign Language Recognition
👋 Welcome
Hello and welcome!
Thank you for your interest in INCLUDE: A Large-Scale Dataset for Indian Sign Language Recognition.
We are pleased to make this dataset available to the… See the full description on the dataset page: https://huggingface.co/datasets/manojkumarcs/INCLUDE_Dataset.sweep_mano_mini_updatedThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "human",
"total_episodes": 10,
"total_frames": 1822,
"total_tasks": 1,
"total_videos": 0,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 50,
"splits": {
"train": "0:10"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path": null,
"features": {
"lang":… See the full description on the dataset page: https://huggingface.co/datasets/Jgold90/sweep_mano_mini_updated.Manosaba_Benchmark_Reorderedtachin-vitra-precut-mano-v1
Tachin VITRA Pre-cut MANO v1
This is the unsegmented Stage-01 release derived from
Tachintech/TachinTactileGlove-Demo01,
pinned at revision 2eb0e45e6d2103f33a7329c657a161c067d516c9.
The release contains all 102 public processed segment videos for which aligned
pose data and fitted MANO are available. It covers 31 source recordings and
54,237 RGB frames. It does not contain atomic episode cuts or generated text
instructions. Use this release as the input to post-MANO… See the full description on the dataset page: https://huggingface.co/datasets/MIT-Media-Lab/tachin-vitra-precut-mano-v1.gigahands-vitra-mano
GigaHands → VITRA Stage-1, official-MANO annotations
VITRA Stage-1 hand annotations for GigaHands, with all joint positions taken from GigaHands'
official MANO fit instead of mixing in triangulated keypoints. Annotations only — no videos
(get those from GigaHands; the mapping is described in §5).
episodes
13,247 (train 11,904 / test 1,343)
frames
3,395,733
camera
brics-odroid-001_cam0 (static rig; one constant extrinsic per scene)
source
GigaHands params/ +… See the full description on the dataset page: https://huggingface.co/datasets/yuzhench/gigahands-vitra-mano.football-players
Dataset Labels
['football', 'player']
Number of Images
{'valid': 87, 'train': 119}
How to Use
Install datasets:
pip install datasets
Load the dataset:
from datasets import load_dataset
ds = load_dataset("manot/football-players", name="full")
example = ds['train'][0]
Roboflow Dataset Page
https://universe.roboflow.com/konstantin-sargsyan-wucpb/football-players-2l81z/dataset/1
Citation
@misc{… See the full description on the dataset page: https://huggingface.co/datasets/manot/football-players.pothole-segmentation
Dataset Labels
['potholes', 'object', 'pothole', 'potholes']
Number of Images
{'valid': 157, 'test': 80, 'train': 582}
How to Use
Install datasets:
pip install datasets
Load the dataset:
from datasets import load_dataset
ds = load_dataset("manot/pothole-segmentation", name="full")
example = ds['train'][0]
Roboflow Dataset Page
https://universe.roboflow.com/abdulmohsen-fahad-f7pdw/road-damage-xvt2d/dataset/3
Citation… See the full description on the dataset page: https://huggingface.co/datasets/manot/pothole-segmentation.so100_marker1This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "so100",
"total_episodes": 2,
"total_frames": 598,
"total_tasks":1,
"total_videos": 2,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:2"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/manoj92/so100_marker1.MedHallu
Dataset Card for MedHallu
MedHallu is a comprehensive benchmark dataset designed to evaluate the ability of large language models to detect hallucinations in medical question-answering tasks.
Dataset Details
Dataset Description
MedHallu is intended to assess the reliability of large language models in a critical domain—medical question-answering—by measuring their capacity to detect hallucinated outputs. The dataset includes two distinct splits:… See the full description on the dataset page: https://huggingface.co/datasets/Manoharareddy/MedHallu.stack_manoThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "human",
"total_episodes": 100,
"total_frames": 13254,
"total_tasks": 1,
"total_videos": 0,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 50,
"splits": {
"train": "0:100"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path": null,
"features": {… See the full description on the dataset page: https://huggingface.co/datasets/mhyatt000/stack_mano.claude-opus-4.6-4.7-reasoning-8.7k
Background
Ended up with some tokens to burn on a Claude Max plan. Assembly began during 4.6 and moved to 4.7. Model is tagged. The development evolved as it went along. The dataset has not been manually reviewed. It's entirely Claude developed.
Clarification on Reasoning
The reasoning is not Claude's actual chain-of-thought (cot) and is not summarized cot. It's a fully synthetic cot created as part of the Assistant response to mimic the type of "thinking"… See the full description on the dataset page: https://huggingface.co/datasets/manojdahal191gom/claude-opus-4.6-4.7-reasoning-8.7k.stack_mano_miniThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "human",
"total_episodes": 10,
"total_frames": 1344,
"total_tasks": 1,
"total_videos": 0,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 50,
"splits": {
"train": "0:10"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path": null,
"features": {
"lang":… See the full description on the dataset page: https://huggingface.co/datasets/mjuarez4/stack_mano_mini.Manosaba_Benchmarkmanosaba-audios
Mahou Shoujo no Majo Saiban - Audio Extract
datas.csv:chapter index, filename, audio path, character name, Japanese text, Chinese translation, etc. This file contains all the text and translation in the game, 34355 rows in total. Some rows do not have corresponding audio (e.g. narration), fill with None. A total of 26251 audios.
Datas/BGM:Background music
Datas/Ambient:Ambient sound
Datas/Sfx:Sound effects
Datas/Song:Songs
Other folders are named after the characters, containing… See the full description on the dataset page: https://huggingface.co/datasets/Sucial/manosaba-audios.pothole-segmentation2
Dataset Labels
['pothole']
Number of Images
{'valid': 133, 'test': 66, 'train': 466}
How to Use
Install datasets:
pip install datasets
Load the dataset:
from datasets import load_dataset
ds = load_dataset("manot/pothole-segmentation2", name="full")
example = ds['train'][0]
Roboflow Dataset Page
https://universe.roboflow.com/gurgen-hovsepyan-mbrnv/pothole-detection-gilij/dataset/2
Citation
@misc{… See the full description on the dataset page: https://huggingface.co/datasets/manot/pothole-segmentation2.
