datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
peacock-data-public-datasets-idcpeacock-data-public-datasets-ywang30egodistillHypoTranslateThis repo releases the HypoTranslate dataset in paper "GenTranslate: Large Language Models are Generative Multilingual Speech and Machine Translators".
Code: https://github.com/YUCHEN005/GenTranslate
Model: https://huggingface.co/PeacefulData/GenTranslate
Data: This repo
Filename format: [split]_[data_source]_[src_language_code]_[tgt_language_code]_[task]_[seamlessm4t_size].pt
e.g. train_fleurs_en_cy_st_large.pt
Note:
Language code look-up: Table 15 & 17 in… See the full description on the dataset page: https://huggingface.co/datasets/PeacefulData/HypoTranslate.Agilex_Cobot_Magic_basket_storage_peach
Agilex_Cobot_Magic_basket_storage_peach
📋 Overview
This dataset uses an extended format based on LeRobot and is fully compatible with LeRobot.
Robot Type: Agilex_Cobot_Magic
| Codebase Version: v2.1
End-Effector Type: two_finger_gripper
🏠 Scene Types
This dataset covers the following scene types:
home
kitchen
🤖 Atomic Actions
This dataset includes the following atomic actions:
grasp
place
pick
📊 Dataset Statistics… See the full description on the dataset page: https://huggingface.co/datasets/RoboCOIN/Agilex_Cobot_Magic_basket_storage_peach.opendatalab-experimental-nmr-peaks
OpenDataLab Experimental NMR Peaks Dataset
Dataset Description
This dataset contains experimental NMR (Nuclear Magnetic Resonance) peak sequences extracted from the OpenDataLab experimental spectra database. The dataset includes both H-NMR and C-NMR peak sequences for chemical compounds, along with their SMILES representations and molecular formulas.
Dataset Summary
Total Samples: 533,595 compounds
Batches: 333 batch files
Data Source: Experimental spectra… See the full description on the dataset page: https://huggingface.co/datasets/snehasis19/opendatalab-experimental-nmr-peaks.weapon-detection-runs-backup-2026-07-24peacock-data-public-datasets-idc-14.backup.outputpeacock-data-public-datasets-idc-cronscriptRobust-HyPoradise
HypothesesParadise
This repo releases the Robust HyPoradise dataset in paper "Large Language Models are Efficient Learners of Noise-Robust Speech Recognition."
GitHub: https://github.com/YUCHEN005/RobustGER
Model: https://huggingface.co/PeacefulData/RobustGER
Data: This repo
UPDATE (Apr-18-2024): We have released the training data, which follows the same format as test data.
Considering the file size, the uploaded training data does not contain the speech features (vast size).… See the full description on the dataset page: https://huggingface.co/datasets/PeacefulData/Robust-HyPoradise.peacock-data-public-datasets-idc-datasetscommon-voice-scripted-speech-26
Common Voice Scripted Speech
A row-normalized multilingual ASR dataset built from Mozilla Data Collective
Common Voice Scripted Speech. Each upstream archive is converted to appendable
parquet shards under data/<upstream_split>/, one shard per source archive and
split, with audio bytes embedded in an audio struct column.
Status
Manifest languages: 60
Languages uploaded: 18
Columns
audio (bytes, path)
sentence, locale, language, upstream_split… See the full description on the dataset page: https://huggingface.co/datasets/Peacockery/common-voice-scripted-speech-26.weapon-detection-workerssd-backup-2026-07-24CoVoGER
CoVoGER: A Multilingual Multitask Benchmark for Speech-to-text Generative Error Correction with Large Language Models
Dataset Description
Large language models (LLMs) can rewrite the N-best hypotheses from a speech-to-text model, often fixing recognition or translation errors that traditional rescoring cannot. Yet research on generative error correction (GER) has been focusing on monolingual automatic speech recognition (ASR), leaving its multilingual and multitask… See the full description on the dataset page: https://huggingface.co/datasets/PeacefulData/CoVoGER.robocasa-composite-raw-videosRMC-AIDA-L_basket_storage_peach
RMC-AIDA-L_basket_storage_peach
📋 Overview
This dataset uses an extended format based on LeRobot and is fully compatible with LeRobot.
Robot Type: realman_rmc_aidal
| Codebase Version: v2.1
End-Effector Type: two_finger_gripper
🏠 Scene Types
This dataset covers the following scene types:
home
kitchen
🤖 Atomic Actions
This dataset includes the following atomic actions:
grasp
pick
place
📊 Dataset Statistics… See the full description on the dataset page: https://huggingface.co/datasets/RoboCOIN/RMC-AIDA-L_basket_storage_peach.HD-EPICTask_800005_pick_place_peanut_stage1_MCAP
Task_800005_pick_place_peanut_stage1_merged_MCAP
Created with Cyclo Intelligence by ROBOTIS.
peacock-data-public-datasets-idc-llm_evalpeacock-data-public-datasets-sangrahaRealman_RMC_AIDA_L_storage_peach_box
Realman_RMC-AIDA-L_storage_peach_box
Dataset Description
This dataset uses an extended format based on LeRobot and is fully compatible with LeRobot.
Task Preview
View Video Directly
Overview
Total Episodes: 118
Total Frames: 78472
FPS: 30
Dataset Size: 751.68 MB
Robot Name: Realman_RMC-AIDA-L
End-Effector Type: two_finger_gripper
Teleoperation Type: Due to some reasons, this dataset temporarily cannot provide the teleoperation… See the full description on the dataset page: https://huggingface.co/datasets/RoboCOIN/Realman_RMC_AIDA_L_storage_peach_box.farsi-asr-iran-international-raw
Iran International raw Farsi audio archive
This repository preserves 11,993 individually addressable FLAC source files for incremental ASR relabeling and reproducible restoration.
Repository file layout
The first 9,990 FLAC files are stored at the repository root. The remaining 2,003 FLAC files are stored individually under overflow/ to respect Hugging Face's 10,000-entry-per-directory limit.
REMOTE_PATHS.jsonl records every source filename, remote path, byte size… See the full description on the dataset page: https://huggingface.co/datasets/Peacockery/farsi-asr-iran-international-raw.R1_Lite_peach_storage
R1_Lite_peach_storage
📋 Overview
This dataset uses an extended format based on LeRobot and is fully compatible with LeRobot.
Robot Type: galaxea_r1_lite
| Codebase Version: v2.1
End-Effector Type: two_finger_gripper
🏠 Scene Types
This dataset covers the following scene types:
kitchen
living_room
🤖 Atomic Actions
This dataset includes the following atomic actions:
grasp
place
pick
📊 Dataset Statistics… See the full description on the dataset page: https://huggingface.co/datasets/RoboCOIN/R1_Lite_peach_storage.CIDER
CIDER: A Dataset of Contextual Disclosure Boundaries for Privacy Preference Alignment
Paper | Code
Dataset for the COLM 2026 paper CIDER: A Dataset of Contextual Disclosure Boundaries for Privacy Preference Alignment
CIDER is a dataset of privacy disclosure decisions collected from real users. It consists of 14,850 annotations from 169 users, forming 1,650 contextual disclosure boundary sets across 60 interpersonal communication scenarios.
What can you do with… See the full description on the dataset page: https://huggingface.co/datasets/peach-lab/CIDER.OpusLM_v1_MLS_Englishconceptnet_en_simpleresumes-raw-pdflibrispeech-phoneme-featuresTask_800007_pick_place_peanut_stage1_MCAP
Task_800007_pick_place_peanut_stage1_MCAP
Created with Cyclo Intelligence by ROBOTIS.
Agilex_Cobot_Magic_storage_peach_brown_bag
Agilex_Cobot_Magic_storage_peach_brown_bag
Dataset Description
This dataset uses an extended format based on LeRobot and is fully compatible with LeRobot.
Task Preview
View Video Directly
Overview
Total Episodes: 100
Total Frames: 49252
FPS: 30
Dataset Size: 568.64 MB
Robot Name: Agilex_Cobot_Magic
End-Effector Type: two_finger_gripper
Teleoperation Type: Due to some reasons, this dataset temporarily cannot provide the… See the full description on the dataset page: https://huggingface.co/datasets/RoboCOIN/Agilex_Cobot_Magic_storage_peach_brown_bag.
