datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
aloha_unscrew_cwtThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 50,
"features": {
"observation.state": {
"dtype": "float32",
"shape": [
14
],
"names": [
"q0",
"q1",
"q2",
"q3",
"q4",
"q5",
"q6",
"q7"… See the full description on the dataset page: https://huggingface.co/datasets/mjkim00/aloha_unscrew_cwt.cwt-multilingual-pretrain-mix
CWT Multilingual Pretrain Mix (reproducible recipe)
A deterministic multilingual-including-English pretraining corpus for controlled vocabulary-scaling studies.
The raw text is not stored here — it is reproduced byte-identically from manifest.json + the pinned
dataset revisions, with no sampling and no RNG (the first-N documents of each stream).
Recipe
English: HuggingFaceFW/fineweb config sample-10BT, revision 9bb295ddab0e05d785b879661af7260fed5140fc… See the full description on the dataset page: https://huggingface.co/datasets/DaveGabe/cwt-multilingual-pretrain-mix.
