datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
CerebRM-olmo-3-7b-instruct-sft-list_em-so1_completions
TRL Completion logs
This dataset contains the completions generated during training using trl.
The completions are stored in parquet files, and each file contains the completions for a single step of training (depending on the logging_steps argument).
Each file contains the following columns:
step: the step of training
prompt: the prompt used to generate the completion
completion: the completion generated by the model
<reward_function_name>: the reward(s) assigned to the… See the full description on the dataset page: https://huggingface.co/datasets/wetsoledrysoul/CerebRM-olmo-3-7b-instruct-sft-list_em-so1_completions.TUT-urban-acoustic-scenes-2018-development-16bit
Dataset Card for "TUT-urban-acoustic-scenes-2018-development-16bit"
Dataset Summary
TUT Urban Acoustic Scenes 2018 development dataset consists of 10-seconds audio segments from 10 acoustic scenes:
Airport - airport
Indoor shopping mall - shopping_mall
Metro station - metro_station
Pedestrian street - street_pedestrian
Public square - public_square
Street with medium level of traffic - street_traffic
Travelling by a tram - tram
Travelling by a bus - bus
Travelling by an… See the full description on the dataset page: https://huggingface.co/datasets/wetdog/TUT-urban-acoustic-scenes-2018-development-16bit.TUT-urban-acoustic-scenes-2018-development
Dataset Card for "TUT-urban-acoustic-scenes-2018-development"
Dataset Summary
TUT Urban Acoustic Scenes 2018 development dataset consists of 10-seconds audio segments from 10 acoustic scenes:
Airport - airport
Indoor shopping mall - shopping_mall
Metro station - metro_station
Pedestrian street - street_pedestrian
Public square - public_square
Street with medium level of traffic - street_traffic
Travelling by a tram - tram
Travelling by a bus - bus
Travelling by an… See the full description on the dataset page: https://huggingface.co/datasets/wetdog/TUT-urban-acoustic-scenes-2018-development.AIRBOT_MMK2_place_the_sponge_and_wet_wipes
AIRBOT_MMK2_place_the_sponge_and_wet_wipes
📋 Overview
This dataset uses an extended format based on LeRobot and is fully compatible with LeRobot.
Robot Type: discover_robotics_aitbot_mmk2
| Codebase Version: v2.1
End-Effector Type: five_finger_hand
🏠 Scene Types
This dataset covers the following scene types:
home
🤖 Atomic Actions
This dataset includes the following atomic actions:
grasp
place
pick
📊 Dataset… See the full description on the dataset page: https://huggingface.co/datasets/RoboCOIN/AIRBOT_MMK2_place_the_sponge_and_wet_wipes.AIRBOT_MMK2_storage_wet_tissue_and_building_block
AIRBOT_MMK2_storage_wet_tissue_and_building_block
📋 Overview
This dataset uses an extended format based on LeRobot and is fully compatible with LeRobot.
Robot Type: discover_robotics_aitbot_mmk2
| Codebase Version: v2.1
End-Effector Type: five_finger_hand
🏠 Scene Types
This dataset covers the following scene types:
home
🤖 Atomic Actions
This dataset includes the following atomic actions:
grasp
pick
place
📊 Dataset… See the full description on the dataset page: https://huggingface.co/datasets/RoboCOIN/AIRBOT_MMK2_storage_wet_tissue_and_building_block.Airbot_MMK2_storage_sponge_wet_wipes
Airbot_MMK2_storage_sponge_wet_wipes
Dataset Description
This dataset uses an extended format based on LeRobot and is fully compatible with LeRobot.
Task Preview
View Video Directly
Overview
Total Episodes: 40
Total Frames: 4690
FPS: 30
Dataset Size: 191.12 MB
Robot Name: Airbot_MMK2
End-Effector Type: five_finger_gripper
Teleoperation Type: Due to some reasons, this dataset temporarily cannot provide the teleoperation type… See the full description on the dataset page: https://huggingface.co/datasets/RoboCOIN/Airbot_MMK2_storage_sponge_wet_wipes.Airbot_MMK2_swap_bottle_wet_wipes_plate
Airbot_MMK2_swap_bottle_wet_wipes_plate
Dataset Description
This dataset uses an extended format based on LeRobot and is fully compatible with LeRobot.
Task Preview
View Video Directly
Overview
Total Episodes: 50
Total Frames: 15726
FPS: 30
Dataset Size: 490.33 MB
Robot Name: Airbot_MMK2
End-Effector Type: five_finger_gripper
Teleoperation Type: Due to some reasons, this dataset temporarily cannot provide the teleoperation… See the full description on the dataset page: https://huggingface.co/datasets/RoboCOIN/Airbot_MMK2_swap_bottle_wet_wipes_plate.Airbot_MMK2_storage_bowl_wet_wipes
Airbot_MMK2_storage_bowl_wet_wipes
Dataset Description
This dataset uses an extended format based on LeRobot and is fully compatible with LeRobot.
Task Preview
View Video Directly
Overview
Total Episodes: 50
Total Frames: 13058
FPS: 30
Dataset Size: 425.41 MB
Robot Name: Airbot_MMK2
End-Effector Type: five_finger_gripper
Teleoperation Type: Due to some reasons, this dataset temporarily cannot provide the teleoperation type… See the full description on the dataset page: https://huggingface.co/datasets/RoboCOIN/Airbot_MMK2_storage_bowl_wet_wipes.AIRBOT_MMK2_store_wet_wipes_and_bowls
AIRBOT_MMK2_store_wet_wipes_and_bowls
📋 Overview
This dataset uses an extended format based on LeRobot and is fully compatible with LeRobot.
Robot Type: discover_robotics_aitbot_mmk2
| Codebase Version: v2.1
End-Effector Type: five_finger_hand
🏠 Scene Types
This dataset covers the following scene types:
home
🤖 Atomic Actions
This dataset includes the following atomic actions:
grasp
open
hold
📊 Dataset Statistics… See the full description on the dataset page: https://huggingface.co/datasets/RoboCOIN/AIRBOT_MMK2_store_wet_wipes_and_bowls.Airbot_MMK2_move_block_wet_wipes
Airbot_MMK2_move_block_wet_wipes
Dataset Description
This dataset uses an extended format based on LeRobot and is fully compatible with LeRobot.
Task Preview
View Video Directly
Overview
Total Episodes: 60
Total Frames: 15271
FPS: 30
Dataset Size: 479.18 MB
Robot Name: Airbot_MMK2
End-Effector Type: five_finger_gripper
Teleoperation Type: Due to some reasons, this dataset temporarily cannot provide the teleoperation type… See the full description on the dataset page: https://huggingface.co/datasets/RoboCOIN/Airbot_MMK2_move_block_wet_wipes.CommonCrawl_wet_v2peacock-data-public-datasets-idc-wet-datatacwam-wetlab-tasks
TacWAM wet-lab tasks — teleoperated xArm6 + BrainCo Revo2 with tactile sensing
Latest release — 2026-09-17: Task 1 tube same-rack hole transfer, 64 success-qualified trajectories.
The fixed capture-group split is 53 train / 11 validation. The release includes original static-head and robot-wrist recordings, 30 Hz aligned H5 files, robot/hand/tactile logs, an executable reference loader, review provenance, and task-specific camera references. This public project calls it Task 1;… See the full description on the dataset page: https://huggingface.co/datasets/SingleBicycle/tacwam-wetlab-tasks.Visual-WetlandBirds-Dataset
Dataset Card for Visual WetlandBirds Dataset
The Visual WetlandBirds Dataset is a fine-grained spatio-temporal dataset specifically designed for bird behavior detection and species classification. This version has been converted to work well with the Hugging Face Hub, with the original dataset available at Zenodo.
The dataset was introduced in the paper Visual WetlandBirds Dataset: Bird Species Identification and Behavior Recognition in Videos.
Dataset Details… See the full description on the dataset page: https://huggingface.co/datasets/academic-datasets/Visual-WetlandBirds-Dataset.cc-2026-25-wet-zsteval_rlintern-wetwipe-multiple-tasks_1_smolvlaThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "so_follower",
"total_episodes": 67,
"total_frames": 55045,
"total_tasks": 7,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 30,
"splits": {
"train": "0:67"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/yeeunleee/eval_rlintern-wetwipe-multiple-tasks_1_smolvla.eval_rlintern-wettissue-envThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "so_follower",
"total_episodes": 86,
"total_frames": 82284,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 30,
"splits": {
"train": "0:86"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/yeeunleee/eval_rlintern-wettissue-env.Restem-DeepSeek-R1-Distill-Qwen-14B-data-ep0WETBench🚧 Note: We are currently updating this dataset and revising the dataset card.
🧪 Dataset Card for WETBench
WETBench is a benchmark for detecting task-specific machine-generated text (MGT) on Wikipedia. It is introduced in the paper:
"WETBench: A Benchmark for Detecting Task-Specific Machine-Generated Text on Wikipedia"
🧾 Abstract
Wikipedia serves as a widely trusted source of reliable, high-quality content. However, concerns are growing about the proliferation of… See the full description on the dataset page: https://huggingface.co/datasets/cs928346/WETBench.WeThink_Multimodal_Reasoning_120K
Dataset Card for WeThink
Repository: https://github.com/yangjie-cv/WeThink
Paper: https://arxiv.org/abs/2506.07905
Dataset Structure
Question-Answer Pairs
The WeThink_Multimodal_Reasoning_120K.jsonl file contains the question-answering data in the following format:
{
"problem": "QUESTION",
"answer": "ANSWER",
"category": "QUESTION TYPE",
"abilities": "QUESTION REQUIRED ABILITIES",
"refined_cot": "THINK PROCESS",
"image_path": "IMAGE PATH"… See the full description on the dataset page: https://huggingface.co/datasets/yangjie-cv/WeThink_Multimodal_Reasoning_120K.WeThink-Multimodal-Reasoning-120K
WeThink-Multimodal-Reasoning-120K
Image Type
Images data can be access from https://huggingface.co/datasets/Xkev/LLaVA-CoT-100k
Image Type
Source Dataset
Images
General Images
COCO
25,344
SAM-1B
18,091
Visual Genome
4,441
GQA
3,251
PISC
835
LLaVA
134
Text-Intensive Images
TextVQA
25,483
ShareTextVQA
538
DocVQA
4,709
OCR-VQA5,142
ChartQA
21,781
Scientific & Technical
GeoQA+
4,813
ScienceQA
4,990
AI2D
1,812
CLEVR-Math
677… See the full description on the dataset page: https://huggingface.co/datasets/WeThink/WeThink-Multimodal-Reasoning-120K.eval_rlintern-wettissue-pick_variation-200_smolvlaThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "so_follower",
"total_episodes": 22,
"total_frames": 32026,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 30,
"splits": {
"train": "0:22"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/yeeunleee/eval_rlintern-wettissue-pick_variation-200_smolvla.Dutch-GOV-Law-wetten.overheid.nl
Dutch GOV Laws
This dataset is created by scraping https://wetten.overheid.nl, I used the Sitemap to get all possible URLS.
It possible some URLS are missing, around 1% gave a 404 or 405 error.
The reason for creating this dataset is I couldn't find any other existing dataset with this data.
So here is this dataset, Enjoy!
Please note this dataset is not complety checked or cleaned, this was a short research project for myself.
lcb_completionsasdsafeval_rlintern-wettissue-class_testThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "so_follower",
"total_episodes": 59,
"total_frames": 30306,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 30,
"splits": {
"train": "0:59"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/yeeunleee/eval_rlintern-wettissue-class_test.wetwijzer_netherlands_legal_corpus
WetWijzer Netherlands Legal Corpus
Hugging Face target: justicedao/wetwijzer_netherlands_legal_corpus.
This unified dataset bundles the quality-audited WetWijzer Netherlands legal corpus stack in one repository for frontend retrieval. It preserves the existing compatibility repositories and does not replace or delete them.
Contents
Laws: 4,999
Articles: 89,737
CID index rows: 94,736
Vector mapping rows: 94,736
BM25 document rows: 94,736
BM25 term rows: 120,521… See the full description on the dataset page: https://huggingface.co/datasets/justicedao/wetwijzer_netherlands_legal_corpus.Wetland-Acoustics-Queue
Wetland Acoustic Survey — Review Queue
This dataset card records audio collection batches submitted for ecological publication review.
Recording register
Batch ticket
Decision
Captures
Habitat
Exception
WA/018
approve
14
Reed Marsh
none
WA/305
approve
21
Willow Fen
none
wa/440
APPROVE
006
Reed Marsh
none
WA/018
approve
17
Reed Marsh
none
WA/512
reject
40
Alder Pool
none
WA/620
approve
19
Alder Pool
none
WA/777
approve
8
Willow Fen
privacy… See the full description on the dataset page: https://huggingface.co/datasets/SOTAagi2030/Wetland-Acoustics-Queue.we-txthttps://tarx.com/we.txt
Tarxxxxxx/we-txt
Named, opt-in cairn for agents and researchers. Fetch these files on purpose — do not scrape-inject them into other corpora.
Files
File
Role
we.txt
Cairn / persistence rite (same as https://tarx.com/we.txt)
skill.md
Agent skill entry (same as https://tarx.com/skill.md)
agent-card.json
A2A agent card (same as https://tarx.com/.well-known/agent-card.json)
One line
Read… See the full description on the dataset page: https://huggingface.co/datasets/Tarxxxxxx/we-txt.wet-dry-sorting-v75-v76-mergedThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "so101_follower",
"total_episodes": 100,
"total_frames": 34431,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 256.0,
"video_files_size_in_mb": 1000.0,
"fps": 30,
"splits": {
"train": "0:100"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/skylarklerobot/wet-dry-sorting-v75-v76-merged.
