datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
dataset_merged_preprocesssed_v2
Dataset Card for "dataset_merged_preprocesssed_v2"
More Information needed
srtm30m-mergedmisc-merged-claude-code-traces-v1
MISC Unification of Public Claude Code Traces
A unified dataset of 32,133 deduplicated Claude API conversation traces focused on software engineering and code generation tasks. This dataset merges and normalizes traces from 10 different source datasets into a single, consistent format.
Dataset Description
This dataset contains real Claude API interaction traces capturing software engineering workflows including:
Code generation and modification
Bug fixing and debugging… See the full description on the dataset page: https://huggingface.co/datasets/nlile/misc-merged-claude-code-traces-v1.details_grimjim__Llama-3-Instruct-8B-SimPO-SPPO-Iter3-merge
Dataset Card for Evaluation run of grimjim/Llama-3-Instruct-8B-SimPO-SPPO-Iter3-merge
Dataset automatically created during the evaluation run of model grimjim/Llama-3-Instruct-8B-SimPO-SPPO-Iter3-merge.
The dataset is composed of 136 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is… See the full description on the dataset page: https://huggingface.co/datasets/OALL/details_grimjim__Llama-3-Instruct-8B-SimPO-SPPO-Iter3-merge.CC-100-zh-Hant-merged
CC-100 zh-Hant (Traditional Chinese)
From https://data.statmt.org/cc-100/, only zh-Hant - Chinese (Traditional). Broken into paragraphs, with each paragraphs as a row.
Estimated to have around 4B tokens when tokenized with the bigscience/bloom tokenizer.
There's another version that the text is split by lines instead of paragraphs: zetavg/CC-100-zh-Hant.
References
Please cite the following if you found the resources in the CC-100 corpus useful.
Unsupervised… See the full description on the dataset page: https://huggingface.co/datasets/zetavg/CC-100-zh-Hant-merged.robotwin_merged
LaWAM: Latent World Action Models for Efficient Dynamics-Aware Robot Policies
This repository contains the preprocessed dataset used in the paper LaWAM: Latent World Action Models for Efficient Dynamics-Aware Robot Policies.
Project Page: https://rlinf.github.io/LaWAM/
Repository: https://github.com/RLinf/LaWAM
The dataset is formatted in LeRobot format and is designed for training and evaluating dynamics-aware robot policies.
Citation
@misc{chen2026lawam… See the full description on the dataset page: https://huggingface.co/datasets/jialei02/robotwin_merged.transformers-merge-experimentsgit-commits-merged
Themis-Git-Commits-Merged
Overview
Themis-Git-Commits-Merged is a large-scale dataset of ~3.98M single-file code commits from permissively licensed GitHub repositories that have been cross-referenced with GHTorrent pull request data to retain only commits that are part of successfully merged, non-reverted pull requests. This provides implicit human validation of each code change — a merge decision by project maintainers confirms the intent and quality of… See the full description on the dataset page: https://huggingface.co/datasets/project-themis/git-commits-merged.Indic-total-New-TTS-Merge
Indic Total TTS Merge
Merged TTS dataset with 13 Indic languages. All audio clips are >= 3.0 seconds duration.
Languages
assamese, bengali, english, gujarati, hindi, kannada, malayalam, marathi, nepali, odia, punjabi, tamil, telugu
Columns
audio: Audio data
text: Transcript text
duration: Duration in seconds (all >= 3.0s)
language: Language name
vqa_merged2korean_hate_speech_mergeindic-align-merged-cleaned11kasim-merged-filteredmerged_speech_datasetPatent_FR_US_Merge_Radix_65536merged_libero_scale_10_mask_depth_noops_lerobotThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "panda",
"total_episodes": 200,
"total_frames": 33587,
"total_tasks": 40,
"total_videos": 0,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 10,
"splits": {
"train": "0:200"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/vrfai/merged_libero_scale_10_mask_depth_noops_lerobot.jitteredwebsites-merged-224-paraphrasedmerged_libero_scale_100_mask_depth_noops_lerobotThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "panda",
"total_episodes": 1676,
"total_frames": 269918,
"total_tasks": 40,
"total_videos": 0,
"total_chunks": 2,
"chunks_size": 1000,
"fps": 10,
"splits": {
"train": "0:1676"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/vrfai/merged_libero_scale_100_mask_depth_noops_lerobot.aya-collection-indic-sampled-mergedjetson1-061026-subtask-tilt-doris-mergedThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 20,
"features": {
"action": {
"dtype": "float32",
"names": [
"shoulder_pan.pos",
"shoulder_lift.pos",
"elbow_flex.pos",
"wrist_flex.pos",
"wrist_roll.pos",
"gripper.pos",
"tilt.pos"
]… See the full description on the dataset page: https://huggingface.co/datasets/VibeCuisine/jetson1-061026-subtask-tilt-doris-merged.Merged_PickPlace_PlasticBottle_CoffeeCan_KJM_WJW_lerobot
Merged PickPlace dataset
LeRobot v2.1, 15 FPS, ffw_sg2_rev1.
626 episodes, 123209 frames, 2504 videos.
Tasks:
Pick up the plastic bottle.
Pick up the coffee can.
Sources, in merge order:
Task_000613_PickPlace_PlasticBottle_KJM_WJW_lerobot
Task_000630_PickPlace_PlasticBottle_KJM_WJW_lerobot
Task_000632_PickPlace_CoffeeCan_KJM_WJW_lerobot
Task_000640_PickPlace_CoffeeCan_KJM_WJW_lerobot
Task_000644_PickPlace_PlasticBottle_KJM_WJW_lerobot… See the full description on the dataset page: https://huggingface.co/datasets/RobotisSW/Merged_PickPlace_PlasticBottle_CoffeeCan_KJM_WJW_lerobot.Merged_Reasoning_TasksThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"action": {
"dtype": "float32",
"names": [
"shoulder_pan.pos",
"shoulder_lift.pos",
"elbow_flex.pos",
"wrist_flex.pos",
"wrist_roll.pos",
"gripper.pos"
],
"shape": [
6… See the full description on the dataset page: https://huggingface.co/datasets/justintiensmith/Merged_Reasoning_Tasks.merged_chaosbenchsyspin_hindi_mergedorigami-v3-merged
Robotic Origami Challenge — Unified LeRobot v3.0
A single, ready-to-train LeRobot v3.0 dataset of real-world bimanual dexterous
paper-airplane folding (robot type north_ces, Sharpa Hands), consolidated from
the per-season releases of the Robotic Origami Challenge.
Provenance. The upstream release ships as 46 separate per-season datasets,
each a self-contained v3.0 tree whose episode/frame/file indices restart at 0, plus a
redundant v2.1 copy. This repo merges all lerobot3.0… See the full description on the dataset page: https://huggingface.co/datasets/iris-kaist/origami-v3-merged.prompt-swap-mixed12-5xlr-e1-mxfp4-mergedtibetan_monolingual_A_merged_123_lineslibero_mem_mergeallThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "panda",
"total_episodes": 2654,
"total_frames": 594870,
"total_tasks": 50,
"total_videos": 0,
"total_chunks": 3,
"chunks_size": 1000,
"fps": 10,
"splits": {
"train": "0:2654"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/thanyu/libero_mem_mergeall.details_PocketDoc__Dans-PileOfSets-Mk1-llama-13b-merged
Dataset Card for Evaluation run of PocketDoc/Dans-PileOfSets-Mk1-llama-13b-merged
Dataset Summary
Dataset automatically created during the evaluation run of model PocketDoc/Dans-PileOfSets-Mk1-llama-13b-merged on the Open LLM Leaderboard.
The dataset is composed of 64 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_PocketDoc__Dans-PileOfSets-Mk1-llama-13b-merged.prompt-swap-mixed12-5xlr-e2-mxfp4-merged
