Merge
Datasets
All datasets matching “Merge”dataset_merged_preprocesssed_v2
Dataset Card for "dataset_merged_preprocesssed_v2"
More Information needed
srtm30m-mergedmisc-merged-claude-code-traces-v1
MISC Unification of Public Claude Code Traces
A unified dataset of 32,133 deduplicated Claude API conversation traces focused on software engineering and code generation tasks. This dataset merges and normalizes traces from 10 different source datasets into a single, consistent format.
Dataset Description
This dataset contains real Claude API interaction traces capturing software engineering workflows including:
Code generation and modification
Bug fixing and debugging… See the full description on the dataset page: https://huggingface.co/datasets/nlile/misc-merged-claude-code-traces-v1.details_grimjim__Llama-3-Instruct-8B-SimPO-SPPO-Iter3-merge
Dataset Card for Evaluation run of grimjim/Llama-3-Instruct-8B-SimPO-SPPO-Iter3-merge
Dataset automatically created during the evaluation run of model grimjim/Llama-3-Instruct-8B-SimPO-SPPO-Iter3-merge.
The dataset is composed of 136 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is… See the full description on the dataset page: https://huggingface.co/datasets/OALL/details_grimjim__Llama-3-Instruct-8B-SimPO-SPPO-Iter3-merge.CC-100-zh-Hant-merged
CC-100 zh-Hant (Traditional Chinese)
From https://data.statmt.org/cc-100/, only zh-Hant - Chinese (Traditional). Broken into paragraphs, with each paragraphs as a row.
Estimated to have around 4B tokens when tokenized with the bigscience/bloom tokenizer.
There's another version that the text is split by lines instead of paragraphs: zetavg/CC-100-zh-Hant.
References
Please cite the following if you found the resources in the CC-100 corpus useful.
Unsupervised… See the full description on the dataset page: https://huggingface.co/datasets/zetavg/CC-100-zh-Hant-merged.robotwin_merged
LaWAM: Latent World Action Models for Efficient Dynamics-Aware Robot Policies
This repository contains the preprocessed dataset used in the paper LaWAM: Latent World Action Models for Efficient Dynamics-Aware Robot Policies.
Project Page: https://rlinf.github.io/LaWAM/
Repository: https://github.com/RLinf/LaWAM
The dataset is formatted in LeRobot format and is designed for training and evaluating dynamics-aware robot policies.
Citation
@misc{chen2026lawam… See the full description on the dataset page: https://huggingface.co/datasets/jialei02/robotwin_merged.
