datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
OpenDebateEvidence-Anonymized
Dataset Card for OpenDebateEvidence (Anonymized)
A collection of evidence used in collegiate and high school debate competitions,
with all debater-identifying columns removed.
This is an anonymized redistribution of
Yusuf5/OpenCaselist. The
argumentative content is byte-for-byte unchanged. 26 of the original 45 columns
have been dropped. See Anonymization for exactly what was
removed and why.
Dataset Details
Dataset Description
This dataset is a… See the full description on the dataset page: https://huggingface.co/datasets/Hellisotherpeople/OpenDebateEvidence-Anonymized.block_pickup_14This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "so100",
"total_episodes": 50,
"total_frames": 45935,
"total_tasks": 1,
"total_videos": 50,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 60,
"splits": {
"train": "0:50"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/HelloCephalopod/block_pickup_14.DebateSum
DebateSum
Corresponding code repo for the upcoming paper at ARGMIN 2020: "DebateSum: A large-scale argument mining and summarization dataset"
Arxiv pre-print available here: https://arxiv.org/abs/2011.07251
Check out the presentation date and time here: https://argmining2020.i3s.unice.fr/node/9
Full paper as presented by the ACL is here: https://www.aclweb.org/anthology/2020.argmining-1.1/
Video of presentation at COLING 2020:… See the full description on the dataset page: https://huggingface.co/datasets/Hellisotherpeople/DebateSum.opengpt-x_hellaswagxThis is a copy of the translations from openGPT-X/hellaswagx, but the repo is
modified so it doesn't require trusting remote code.
Citation Information
If you find benchmarks useful in your research, please consider citing the test and also the HellaSwag dataset it draws from:
@misc{thellmann2024crosslingual,
title={Towards Cross-Lingual LLM Evaluation for European Languages},
author={Klaudia Thellmann and Bernhard Stadler and Michael Fromm and Jasper Schulze Buschhoff… See the full description on the dataset page: https://huggingface.co/datasets/LumiOpen/opengpt-x_hellaswagx.hellaswagblock_pickup_17This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "so100_follower",
"total_episodes": 20,
"total_frames": 8991,
"total_tasks": 1,
"total_videos": 20,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:20"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/HelloCephalopod/block_pickup_17.CC-FilteredCorpus
English Cleaned Common Crawl Markdown Dataset
An English-focused dataset created from Common Crawl, cleaned and converted to Markdown.
The goal is to preserve web-document structure so that AI models can learn both natural language and Markdown formatting.
Features
English-focused
Cleaned and filtered web content
HTML converted to Markdown
Exact and near-duplicate filtering
GPT-2 perplexity filtering
Stored as compressed Parquet shards
Source
The… See the full description on the dataset page: https://huggingface.co/datasets/helloadhavan/CC-FilteredCorpus.hello26_mergedThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"action": {
"dtype": "float32",
"names": [
"left_shoulder_pan.pos",
"left_shoulder_lift.pos",
"left_elbow_flex.pos",
"left_wrist_flex.pos",
"left_wrist_roll.pos",
"left_gripper.pos"… See the full description on the dataset page: https://huggingface.co/datasets/zeta0707/hello26_merged.OpenDebateEvidence-Deduplicated-Anonymized
Dataset Card for OpenDebateEvidence-Deduplicated (Anonymized)
Debate evidence from collegiate and high school competitions, semantically
deduplicated, with all debater-identifying columns removed.
This is the semantically deduplicated companion to
OpenDebateEvidence-Anonymized.
Where the parent dataset contains every piece of evidence as used in every round,
this version collapses repeated use of the same evidence into single records,
making it substantially smaller and better… See the full description on the dataset page: https://huggingface.co/datasets/Hellisotherpeople/OpenDebateEvidence-Deduplicated-Anonymized.hellaswagXTransEnV_hellaswag
Version 2 (2026-08): Dialect configs regenerated with a stronger pipeline
The 18 dialect configs (AAVE, AppE, AuE, AuE_V, BahE, EAngE, IrE, Manx, NZE,
N_Eng, NfE, OzE, SE_AmE, SE_Eng, SW_Eng, ScE, TdCE, WaE) were regenerated with an
upgraded Trans-EnV pipeline. The ESL configs (A_*/B_*) are unchanged (v1).
Previous versions of all files remain available via git revisions of this repo.
What changed
Transformation model: google/gemma-2-27b-it → google/gemma-4-31B-it,
with a… See the full description on the dataset page: https://huggingface.co/datasets/jiyounglee0523/TransEnV_hellaswag.hello-hands-multitask-cubesThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"action": {
"dtype": "float32",
"names": [
"shoulder_pan.pos",
"shoulder_lift.pos",
"elbow_flex.pos",
"wrist_flex.pos",
"wrist_roll.pos",
"gripper.pos"
],
"shape": [
6… See the full description on the dataset page: https://huggingface.co/datasets/harrison-powe/hello-hands-multitask-cubes.so100_hellolerobotThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "so100",
"total_episodes": 3,
"total_frames": 1793,
"total_tasks": 1,
"total_videos": 6,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:3"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/dulics/so100_hellolerobot.so100_bi_helloThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.0",
"robot_type": "so100",
"total_episodes": 20,
"total_frames": 5836,
"total_tasks": 1,
"total_videos": 60,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:20"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/liyitenga/so100_bi_hello.so100_test3This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.0",
"robot_type": "so100",
"total_episodes": 30,
"total_frames": 10830,
"total_tasks": 1,
"total_videos": 60,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:30"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/hellkrusher/so100_test3.so100_helloworldThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "so100",
"total_episodes": 3,
"total_frames": 1792,
"total_tasks": 1,
"total_videos": 6,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:3"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/dulics/so100_helloworld.HellaSwag_TH_cleanned
Dataset Card for HellaSwag_TH_cleanned
Dataset Description
This dataset is Thai translated version of hellaswag using google translate with Multilingual Universal Sentence Encoder to calculate score for Thai translation.
The score was penalized by the length of original text compare to translated text. The row that any score < 0.5 was dropped.
Languages
EN
TH
Citation
@misc{HellaSwag_TH_cleanned,
author = {Triamamornwooth Patteera},
title =… See the full description on the dataset page: https://huggingface.co/datasets/Patt/HellaSwag_TH_cleanned.libero_spatial_no_noops_1.0.0_lerobot_v3.0
libero_spatial_no_noops_1.0.0_lerobot
Dataset Description
This dataset is the LIBERO-Spatial subset of the LIBERO benchmark, converted to LeRobot v3.0 format.
10 spatial generalization tasks. All tasks share the same objects but differ in their spatial layout/configuration, testing the agent's ability to generalize to new object placements.
Note: This dataset has been filtered to remove no-op frames (idle frames where the robot does not execute meaningful actions). This… See the full description on the dataset page: https://huggingface.co/datasets/hello3x3/libero_spatial_no_noops_1.0.0_lerobot_v3.0.kor_hellaswag
Dataset Card for "kor_hellaswag"
More Information needed
Source Data Citation Information
@inproceedings{zellers2019hellaswag,
title={HellaSwag: Can a Machine Really Finish Your Sentence?},
author={Zellers, Rowan and Holtzman, Ari and Bisk, Yonatan and Farhadi, Ali and Choi, Yejin},
booktitle ={Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics},
year={2019}
}
Hellaswag-poly
HellaSwag Polyglot
This dataset is a multilingual version of the original HellaSwag (Harder Endings, Longer contexts, and Low-shot Activities for Situations With Adversarial Generations) dataset, which consists short commonsense reasoning tasks designed to evaluate the ability of language models to understand and predict plausible continuations of given contexts. The polyglot version includes translations of the original English questions into various languages, allowing for… See the full description on the dataset page: https://huggingface.co/datasets/Polygl0t/Hellaswag-poly.hello-ot-imagenet-pca
HELLO representative ImageNet VA-VAE PCA features
This dataset contains four independently generated float32 NumPy matrices,
with shapes (262144, d) for d = 4, 32, 256, 2048, totaling about 2.29 GiB.
All use seed=42 and serve the public main-scaling and parameter-sensitivity
experiments. Larger sample counts are outside the published data scope.
The artifact contains numeric PCA-projected features only. It contains no
images, labels, captions, filenames, or ImageNet identifiers.… See the full description on the dataset page: https://huggingface.co/datasets/WenzhouXia/hello-ot-imagenet-pca.libero_goal_no_noops_1.0.0_lerobot_v3.0
libero_goal_no_noops_1.0.0_lerobot
Dataset Description
This dataset is the LIBERO-Goal subset of the LIBERO benchmark, converted to LeRobot v3.0 format.
10 goal-conditioned generalization tasks. Tasks share the same scene but have different goals/objectives, testing the agent's ability to follow diverse instructions in a shared environment.
Note: This dataset has been filtered to remove no-op frames (idle frames where the robot does not execute meaningful actions). This… See the full description on the dataset page: https://huggingface.co/datasets/hello3x3/libero_goal_no_noops_1.0.0_lerobot_v3.0.qwen3.8-max-glm5.2-kimi-k3-distillation
Multi-Teacher Distillation Dataset (57,937 traces)
A quality-filtered, deduplicated, multi-teacher SFT corpus combining traces from three frontier models across math, code, reasoning, instruction-following, tool-use, science, long-context, multilingual, and creative dialogue domains.
Teachers
Teacher
Provider
Traces
Qwen3.8-Max-Preview
Alibaba Cloud Model Studio
48,283
GLM-5.2
Z.AI Coding Plan
5,307
Kimi Code K3
Moonshot AI (Kimi)
4,347… See the full description on the dataset page: https://huggingface.co/datasets/Helloxiaolaodi/qwen3.8-max-glm5.2-kimi-k3-distillation.ai-translate-video
Video Localization Timing Dataset
Dataset Summary
This dataset is a self-authored synthetic benchmark for multilingual video localization workflows. It focuses on the timing pressure that appears when subtitle segments are translated across languages and then reviewed for dubbing fit, subtitle-window preservation, and lip-sync risk. The package is designed for repository-safe experimentation and documentation. It does not contain third-party video, third-party audio… See the full description on the dataset page: https://huggingface.co/datasets/hellohihiloy789/ai-translate-video.libero_plus_lerobot_v3.0This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "panda",
"total_episodes": 14347,
"total_frames": 2238036,
"total_tasks": 40,
"chunks_size": 1000,
"fps": 20,
"splits": {
"train": "0:14347"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path": "videos/{video_key}/chunk-{chunk_index:03d}/file-{file_index:03d}.mp4"… See the full description on the dataset page: https://huggingface.co/datasets/hello3x3/libero_plus_lerobot_v3.0.so100_lerobot2_helloThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "so100",
"total_episodes": 2,
"total_frames": 561,
"total_tasks": 1,
"total_videos": 0,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:2"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/ncavallo/so100_lerobot2_hello.cbis-ddsm-r
CBIS-DDSM-R: A Curated Radiomic Feature Dataset for Breast Cancer Classification
Dataset Summary
CBIS-DDSM-R is an open-source, radiomics-ready extension of the Curated Breast Imaging Subset of the Digital Database for Screening Mammography (CBIS-DDSM). It is designed to facilitate reproducible radiomics and quantitative imaging research in breast cancer analysis.
The dataset provides a standardized preprocessing pipeline for mammograms and includes IBSI-compliant… See the full description on the dataset page: https://huggingface.co/datasets/helloerikaaa/cbis-ddsm-r.so101_pick_place_2cam_v4This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "so_follower",
"total_episodes": 52,
"total_frames": 14201,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 30,
"splits": {
"train": "0:52"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/helloworld26/so101_pick_place_2cam_v4.eval_cap_pen_two_handThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "bi_so_follower",
"total_episodes": 1,
"total_frames": 3972,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 20,
"splits": {
"train": "0:1"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/hellozjt/eval_cap_pen_two_hand.HellaSwag_TH
Dataset Card for HellaSwag_TH
Dataset Description
The cleaned version is available at Patt/HellaSwag_TH_cleanned.
This dataset is Thai translated version of hellaswag using google translate with Multilingual Universal Sentence Encoder to calculate score for Thai translation.
Languages
EN
TH
Citation
@misc{HellaSwag_TH,
author = {Triamamornwooth Patteera},
title = {HellaSwag_TH},
year = {2023},
publisher = {Hugging Face}… See the full description on the dataset page: https://huggingface.co/datasets/Patt/HellaSwag_TH.
