merged
Datasets
All datasets matching “merged”dataset_merged_preprocesssed_v2
Dataset Card for "dataset_merged_preprocesssed_v2"
More Information needed
srtm30m-mergedmisc-merged-claude-code-traces-v1
MISC Unification of Public Claude Code Traces
A unified dataset of 32,133 deduplicated Claude API conversation traces focused on software engineering and code generation tasks. This dataset merges and normalizes traces from 10 different source datasets into a single, consistent format.
Dataset Description
This dataset contains real Claude API interaction traces capturing software engineering workflows including:
Code generation and modification
Bug fixing and debugging… See the full description on the dataset page: https://huggingface.co/datasets/nlile/misc-merged-claude-code-traces-v1.robotwin_merged
LaWAM: Latent World Action Models for Efficient Dynamics-Aware Robot Policies
This repository contains the preprocessed dataset used in the paper LaWAM: Latent World Action Models for Efficient Dynamics-Aware Robot Policies.
Project Page: https://rlinf.github.io/LaWAM/
Repository: https://github.com/RLinf/LaWAM
The dataset is formatted in LeRobot format and is designed for training and evaluating dynamics-aware robot policies.
Citation
@misc{chen2026lawam… See the full description on the dataset page: https://huggingface.co/datasets/jialei02/robotwin_merged.CC-100-zh-Hant-merged
CC-100 zh-Hant (Traditional Chinese)
From https://data.statmt.org/cc-100/, only zh-Hant - Chinese (Traditional). Broken into paragraphs, with each paragraphs as a row.
Estimated to have around 4B tokens when tokenized with the bigscience/bloom tokenizer.
There's another version that the text is split by lines instead of paragraphs: zetavg/CC-100-zh-Hant.
References
Please cite the following if you found the resources in the CC-100 corpus useful.
Unsupervised… See the full description on the dataset page: https://huggingface.co/datasets/zetavg/CC-100-zh-Hant-merged.git-commits-merged
Themis-Git-Commits-Merged
Overview
Themis-Git-Commits-Merged is a large-scale dataset of ~3.98M single-file code commits from permissively licensed GitHub repositories that have been cross-referenced with GHTorrent pull request data to retain only commits that are part of successfully merged, non-reverted pull requests. This provides implicit human validation of each code change — a merge decision by project maintainers confirms the intent and quality of… See the full description on the dataset page: https://huggingface.co/datasets/project-themis/git-commits-merged.
