datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
OpenDebateEvidence-Anonymized
Dataset Card for OpenDebateEvidence (Anonymized)
A collection of evidence used in collegiate and high school debate competitions,
with all debater-identifying columns removed.
This is an anonymized redistribution of
Yusuf5/OpenCaselist. The
argumentative content is byte-for-byte unchanged. 26 of the original 45 columns
have been dropped. See Anonymization for exactly what was
removed and why.
Dataset Details
Dataset Description
This dataset is a… See the full description on the dataset page: https://huggingface.co/datasets/Hellisotherpeople/OpenDebateEvidence-Anonymized.block_pickup_14This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "so100",
"total_episodes": 50,
"total_frames": 45935,
"total_tasks": 1,
"total_videos": 50,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 60,
"splits": {
"train": "0:50"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/HelloCephalopod/block_pickup_14.block_pickup_17This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "so100_follower",
"total_episodes": 20,
"total_frames": 8991,
"total_tasks": 1,
"total_videos": 20,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:20"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/HelloCephalopod/block_pickup_17.hellaswagCC-FilteredCorpus
English Cleaned Common Crawl Markdown Dataset
An English-focused dataset created from Common Crawl, cleaned and converted to Markdown.
The goal is to preserve web-document structure so that AI models can learn both natural language and Markdown formatting.
Features
English-focused
Cleaned and filtered web content
HTML converted to Markdown
Exact and near-duplicate filtering
GPT-2 perplexity filtering
Stored as compressed Parquet shards
Source
The… See the full description on the dataset page: https://huggingface.co/datasets/helloadhavan/CC-FilteredCorpus.hello26_mergedThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"action": {
"dtype": "float32",
"names": [
"left_shoulder_pan.pos",
"left_shoulder_lift.pos",
"left_elbow_flex.pos",
"left_wrist_flex.pos",
"left_wrist_roll.pos",
"left_gripper.pos"… See the full description on the dataset page: https://huggingface.co/datasets/zeta0707/hello26_merged.OpenDebateEvidence-Deduplicated-Anonymized
Dataset Card for OpenDebateEvidence-Deduplicated (Anonymized)
Debate evidence from collegiate and high school competitions, semantically
deduplicated, with all debater-identifying columns removed.
This is the semantically deduplicated companion to
OpenDebateEvidence-Anonymized.
Where the parent dataset contains every piece of evidence as used in every round,
this version collapses repeated use of the same evidence into single records,
making it substantially smaller and better… See the full description on the dataset page: https://huggingface.co/datasets/Hellisotherpeople/OpenDebateEvidence-Deduplicated-Anonymized.hellaswagXTransEnV_hellaswag
Version 2 (2026-08): Dialect configs regenerated with a stronger pipeline
The 18 dialect configs (AAVE, AppE, AuE, AuE_V, BahE, EAngE, IrE, Manx, NZE,
N_Eng, NfE, OzE, SE_AmE, SE_Eng, SW_Eng, ScE, TdCE, WaE) were regenerated with an
upgraded Trans-EnV pipeline. The ESL configs (A_*/B_*) are unchanged (v1).
Previous versions of all files remain available via git revisions of this repo.
What changed
Transformation model: google/gemma-2-27b-it → google/gemma-4-31B-it,
with a… See the full description on the dataset page: https://huggingface.co/datasets/jiyounglee0523/TransEnV_hellaswag.hello-hands-multitask-cubesThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"action": {
"dtype": "float32",
"names": [
"shoulder_pan.pos",
"shoulder_lift.pos",
"elbow_flex.pos",
"wrist_flex.pos",
"wrist_roll.pos",
"gripper.pos"
],
"shape": [
6… See the full description on the dataset page: https://huggingface.co/datasets/harrison-powe/hello-hands-multitask-cubes.so100_hellolerobotThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "so100",
"total_episodes": 3,
"total_frames": 1793,
"total_tasks": 1,
"total_videos": 6,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:3"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/dulics/so100_hellolerobot.so100_bi_helloThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.0",
"robot_type": "so100",
"total_episodes": 20,
"total_frames": 5836,
"total_tasks": 1,
"total_videos": 60,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:20"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/liyitenga/so100_bi_hello.so100_test3This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.0",
"robot_type": "so100",
"total_episodes": 30,
"total_frames": 10830,
"total_tasks": 1,
"total_videos": 60,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:30"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/hellkrusher/so100_test3.so100_helloworldThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "so100",
"total_episodes": 3,
"total_frames": 1792,
"total_tasks": 1,
"total_videos": 6,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:3"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/dulics/so100_helloworld.kor_hellaswag
Dataset Card for "kor_hellaswag"
More Information needed
Source Data Citation Information
@inproceedings{zellers2019hellaswag,
title={HellaSwag: Can a Machine Really Finish Your Sentence?},
author={Zellers, Rowan and Holtzman, Ari and Bisk, Yonatan and Farhadi, Ali and Choi, Yejin},
booktitle ={Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics},
year={2019}
}
HellaSwag_TH_cleanned
Dataset Card for HellaSwag_TH_cleanned
Dataset Description
This dataset is Thai translated version of hellaswag using google translate with Multilingual Universal Sentence Encoder to calculate score for Thai translation.
The score was penalized by the length of original text compare to translated text. The row that any score < 0.5 was dropped.
Languages
EN
TH
Citation
@misc{HellaSwag_TH_cleanned,
author = {Triamamornwooth Patteera},
title =… See the full description on the dataset page: https://huggingface.co/datasets/Patt/HellaSwag_TH_cleanned.so101_pick_place_2cam_v4This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "so_follower",
"total_episodes": 52,
"total_frames": 14201,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 30,
"splits": {
"train": "0:52"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/helloworld26/so101_pick_place_2cam_v4.qwen3.8-max-glm5.2-kimi-k3-distillation
Multi-Teacher Distillation Dataset (57,937 traces)
A quality-filtered, deduplicated, multi-teacher SFT corpus combining traces from three frontier models across math, code, reasoning, instruction-following, tool-use, science, long-context, multilingual, and creative dialogue domains.
Teachers
Teacher
Provider
Traces
Qwen3.8-Max-Preview
Alibaba Cloud Model Studio
48,283
GLM-5.2
Z.AI Coding Plan
5,307
Kimi Code K3
Moonshot AI (Kimi)
4,347… See the full description on the dataset page: https://huggingface.co/datasets/Helloxiaolaodi/qwen3.8-max-glm5.2-kimi-k3-distillation.so100_lerobot2_helloThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "so100",
"total_episodes": 2,
"total_frames": 561,
"total_tasks": 1,
"total_videos": 0,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:2"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/ncavallo/so100_lerobot2_hello.libero_spatial_no_noops_1.0.0_lerobot_v3.0
libero_spatial_no_noops_1.0.0_lerobot
Dataset Description
This dataset is the LIBERO-Spatial subset of the LIBERO benchmark, converted to LeRobot v3.0 format.
10 spatial generalization tasks. All tasks share the same objects but differ in their spatial layout/configuration, testing the agent's ability to generalize to new object placements.
Note: This dataset has been filtered to remove no-op frames (idle frames where the robot does not execute meaningful actions). This… See the full description on the dataset page: https://huggingface.co/datasets/hello3x3/libero_spatial_no_noops_1.0.0_lerobot_v3.0.libero_plus_lerobot_v3.0This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "panda",
"total_episodes": 14347,
"total_frames": 2238036,
"total_tasks": 40,
"chunks_size": 1000,
"fps": 20,
"splits": {
"train": "0:14347"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path": "videos/{video_key}/chunk-{chunk_index:03d}/file-{file_index:03d}.mp4"… See the full description on the dataset page: https://huggingface.co/datasets/hello3x3/libero_plus_lerobot_v3.0.Hellaswag-poly
HellaSwag Polyglot
This dataset is a multilingual version of the original HellaSwag (Harder Endings, Longer contexts, and Low-shot Activities for Situations With Adversarial Generations) dataset, which consists short commonsense reasoning tasks designed to evaluate the ability of language models to understand and predict plausible continuations of given contexts. The polyglot version includes translations of the original English questions into various languages, allowing for… See the full description on the dataset page: https://huggingface.co/datasets/Polygl0t/Hellaswag-poly.hello-hands-cube-pick-place_20260804_190811This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"action": {
"dtype": "float32",
"names": [
"shoulder_pan.pos",
"shoulder_lift.pos",
"elbow_flex.pos",
"wrist_flex.pos",
"wrist_roll.pos",
"gripper.pos"
],
"shape": [
6… See the full description on the dataset page: https://huggingface.co/datasets/harrison-powe/hello-hands-cube-pick-place_20260804_190811.eval_so100_test5This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "so100",
"total_episodes": 5,
"total_frames": 1683,
"total_tasks": 1,
"total_videos": 10,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:5"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/hellkrusher/eval_so100_test5.libero_goal_no_noops_1.0.0_lerobot_v3.0
libero_goal_no_noops_1.0.0_lerobot
Dataset Description
This dataset is the LIBERO-Goal subset of the LIBERO benchmark, converted to LeRobot v3.0 format.
10 goal-conditioned generalization tasks. Tasks share the same scene but have different goals/objectives, testing the agent's ability to follow diverse instructions in a shared environment.
Note: This dataset has been filtered to remove no-op frames (idle frames where the robot does not execute meaningful actions). This… See the full description on the dataset page: https://huggingface.co/datasets/hello3x3/libero_goal_no_noops_1.0.0_lerobot_v3.0.eval_so100_test3This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "so100",
"total_episodes": 2,
"total_frames": 668,
"total_tasks": 1,
"total_videos": 4,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:2"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/hellkrusher/eval_so100_test3.so101_pick_place_arm_cameraThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "so_follower",
"total_episodes": 50,
"total_frames": 27393,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 30,
"splits": {
"train": "0:50"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/helloworld26/so101_pick_place_arm_camera.x1_testThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.0",
"robot_type": "x1",
"total_episodes": 2,
"total_frames": 1135,
"total_tasks": 1,
"total_videos": 4,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:2"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/Hello-XSJ/x1_test.eval_cap_pen_two_handThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "bi_so_follower",
"total_episodes": 1,
"total_frames": 3972,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 20,
"splits": {
"train": "0:1"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/hellozjt/eval_cap_pen_two_hand.so100_testThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.0",
"robot_type": "so100",
"total_episodes": 2,
"total_frames": 678,
"total_tasks": 1,
"total_videos": 4,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:2"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/hellershi123/so100_test.
