datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
dataset_merged_preprocesssed_v2
Dataset Card for "dataset_merged_preprocesssed_v2"
More Information needed
misc-merged-claude-code-traces-v1
MISC Unification of Public Claude Code Traces
A unified dataset of 32,133 deduplicated Claude API conversation traces focused on software engineering and code generation tasks. This dataset merges and normalizes traces from 10 different source datasets into a single, consistent format.
Dataset Description
This dataset contains real Claude API interaction traces capturing software engineering workflows including:
Code generation and modification
Bug fixing and debugging… See the full description on the dataset page: https://huggingface.co/datasets/nlile/misc-merged-claude-code-traces-v1.details_grimjim__Llama-3-Instruct-8B-SimPO-SPPO-Iter3-merge
Dataset Card for Evaluation run of grimjim/Llama-3-Instruct-8B-SimPO-SPPO-Iter3-merge
Dataset automatically created during the evaluation run of model grimjim/Llama-3-Instruct-8B-SimPO-SPPO-Iter3-merge.
The dataset is composed of 136 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is… See the full description on the dataset page: https://huggingface.co/datasets/OALL/details_grimjim__Llama-3-Instruct-8B-SimPO-SPPO-Iter3-merge.CC-100-zh-Hant-merged
CC-100 zh-Hant (Traditional Chinese)
From https://data.statmt.org/cc-100/, only zh-Hant - Chinese (Traditional). Broken into paragraphs, with each paragraphs as a row.
Estimated to have around 4B tokens when tokenized with the bigscience/bloom tokenizer.
There's another version that the text is split by lines instead of paragraphs: zetavg/CC-100-zh-Hant.
References
Please cite the following if you found the resources in the CC-100 corpus useful.
Unsupervised… See the full description on the dataset page: https://huggingface.co/datasets/zetavg/CC-100-zh-Hant-merged.git-commits-merged
Themis-Git-Commits-Merged
Overview
Themis-Git-Commits-Merged is a large-scale dataset of ~3.98M single-file code commits from permissively licensed GitHub repositories that have been cross-referenced with GHTorrent pull request data to retain only commits that are part of successfully merged, non-reverted pull requests. This provides implicit human validation of each code change — a merge decision by project maintainers confirms the intent and quality of… See the full description on the dataset page: https://huggingface.co/datasets/project-themis/git-commits-merged.Indic-total-New-TTS-Merge
Indic Total TTS Merge
Merged TTS dataset with 13 Indic languages. All audio clips are >= 3.0 seconds duration.
Languages
assamese, bengali, english, gujarati, hindi, kannada, malayalam, marathi, nepali, odia, punjabi, tamil, telugu
Columns
audio: Audio data
text: Transcript text
duration: Duration in seconds (all >= 3.0s)
language: Language name
vqa_merged211kasim-merged-filteredmerged_speech_datasetPatent_FR_US_Merge_Radix_65536merged_libero_scale_10_mask_depth_noops_lerobotThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "panda",
"total_episodes": 200,
"total_frames": 33587,
"total_tasks": 40,
"total_videos": 0,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 10,
"splits": {
"train": "0:200"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/vrfai/merged_libero_scale_10_mask_depth_noops_lerobot.jitteredwebsites-merged-224-paraphrasedmerged_libero_scale_100_mask_depth_noops_lerobotThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "panda",
"total_episodes": 1676,
"total_frames": 269918,
"total_tasks": 40,
"total_videos": 0,
"total_chunks": 2,
"chunks_size": 1000,
"fps": 10,
"splits": {
"train": "0:1676"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/vrfai/merged_libero_scale_100_mask_depth_noops_lerobot.aya-collection-indic-sampled-mergedjetson1-061026-subtask-tilt-doris-mergedThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 20,
"features": {
"action": {
"dtype": "float32",
"names": [
"shoulder_pan.pos",
"shoulder_lift.pos",
"elbow_flex.pos",
"wrist_flex.pos",
"wrist_roll.pos",
"gripper.pos",
"tilt.pos"
]… See the full description on the dataset page: https://huggingface.co/datasets/VibeCuisine/jetson1-061026-subtask-tilt-doris-merged.Merged_Reasoning_TasksThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"action": {
"dtype": "float32",
"names": [
"shoulder_pan.pos",
"shoulder_lift.pos",
"elbow_flex.pos",
"wrist_flex.pos",
"wrist_roll.pos",
"gripper.pos"
],
"shape": [
6… See the full description on the dataset page: https://huggingface.co/datasets/justintiensmith/Merged_Reasoning_Tasks.syspin_hindi_mergedtibetan_monolingual_A_merged_123_lineslibero_mem_mergeallThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "panda",
"total_episodes": 2654,
"total_frames": 594870,
"total_tasks": 50,
"total_videos": 0,
"total_chunks": 3,
"chunks_size": 1000,
"fps": 10,
"splits": {
"train": "0:2654"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/thanyu/libero_mem_mergeall.pt_mergevqa_plant-disease-classification-merged-datasetso101_merged_20260701_fixedThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"action": {
"dtype": "float32",
"shape": [
6
],
"names": [
"shoulder_pan.pos",
"shoulder_lift.pos",
"elbow_flex.pos",
"wrist_flex.pos",
"wrist_roll.pos",
"gripper.pos"… See the full description on the dataset page: https://huggingface.co/datasets/makermods/so101_merged_20260701_fixed.gr00t-g1-grab-bottle-right-hand-radius-20-merged
Grab-Bottle (right hand) - radius-20 merged
LeRobot v2.1 dataset for the Unitree G1 right-hand bottle-grab task. This is a
merge of two source teleoperation datasets, curated with the zero-wandering
pipeline at smoothing half-width 20 and safe_frames 50, then renumbered into
one contiguous set of 502 episodes. It is the largest and best-curated set of the
grab-bottle lineage and the training data for the
v6 GR00T N1.7 fine-tune.
Each surviving clean sub-segment is emitted as its… See the full description on the dataset page: https://huggingface.co/datasets/cloudwalk-research/gr00t-g1-grab-bottle-right-hand-radius-20-merged.merged_200ep_24corr3x_10trans_blue_cube_orange_trayThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"action": {
"dtype": "float32",
"names": [
"shoulder_pan.pos",
"shoulder_lift.pos",
"elbow_flex.pos",
"wrist_flex.pos",
"wrist_roll.pos",
"gripper.pos"
],
"shape": [
6… See the full description on the dataset page: https://huggingface.co/datasets/makermods/merged_200ep_24corr3x_10trans_blue_cube_orange_tray.full-fold-the-rag-parquet-merged0222instruction_merge_set
Dataset Card for "instruction_merge_set"
本数据集由以下数据集构成:
数据(id in the merged set)
Hugging face 地址
notes
OIG (unified-任务名称) 15k
https://huggingface.co/datasets/laion/OIG
Open Instruction Generalist Dataset
Dolly databricks-dolly-15k
https://huggingface.co/datasets/databricks/databricks-dolly-15k
an open-source dataset of instruction-following records generated by thousands of Databricks employees in several of the behavioral categories
UltraChat… See the full description on the dataset page: https://huggingface.co/datasets/LinkSoul/instruction_merge_set.merged_200ep_100corr3x_blue_cube_orange_boxThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"action": {
"dtype": "float32",
"shape": [
6
],
"names": [
"shoulder_pan.pos",
"shoulder_lift.pos",
"elbow_flex.pos",
"wrist_flex.pos",
"wrist_roll.pos",
"gripper.pos"… See the full description on the dataset page: https://huggingface.co/datasets/makermods/merged_200ep_100corr3x_blue_cube_orange_box.highlevel_thinking_with_grounding_annotation_split1000_v3_merged_promptstibetan_monolingual_A_merged_135_linesnew_camera_orange_color_mergedThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"action": {
"dtype": "float32",
"shape": [
6
],
"names": [
"shoulder_pan.pos",
"shoulder_lift.pos",
"elbow_flex.pos",
"wrist_flex.pos",
"wrist_roll.pos",
"gripper.pos"… See the full description on the dataset page: https://huggingface.co/datasets/makermods/new_camera_orange_color_merged.
