language
Datasets
All datasets matching “language”language_table_lerobotThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.0",
"robot_type": "xarm",
"total_episodes": 442226,
"total_frames": 7045476,
"total_tasks": 127605,
"total_videos": 442226,
"total_chunks": 443,
"chunks_size": 1000,
"fps": 10,
"splits": {
"train": "0:442226"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/IPEC-COMMUNITY/language_table_lerobot.Open-Sora-Plan-v1.1.0
Annotation
We resized the dataset to 1080p for easier uploading. Therefore, the original annotation file might not match the video names. Please refer to this https://github.com/PKU-YuanGroup/Open-Sora-Plan/issues/312#issuecomment-2197312973
Pexels
Pexels consists of multiple folders, but each folder exceeds the size limit for Huggingface uploads. Therefore, we divided each folder into 5 parts. You need to merge the 5 parts of each folder first, and then extract each… See the full description on the dataset page: https://huggingface.co/datasets/LanguageBind/Open-Sora-Plan-v1.1.0.github-code-2025-language-split
📜 Source Data & Attribution
This dataset is a processed derivative of nick007x/github-code-2025.
Origination
The original data was aggregated by nick007x from public GitHub repositories. We have retained the original content, file paths, and metadata while restructuring the format for easier consumption by language-specific models.
Processing Steps
To create this dataset, we performed the following processing on the source data:
Language… See the full description on the dataset page: https://huggingface.co/datasets/hasankursun/github-code-2025-language-split.aya_collection_language_split
This is a re-upload of the aya_collection, and only differs in the structure of upload. While the original aya_collection is structured by folders split according to dataset name, this dataset is split by language. We recommend you use this version of the dataset if you are only interested in downloading all of the Aya collection for a single or smaller set of languages.
Dataset Summary
The Aya Collection is a massive multilingual collection consisting of 513 million instances of… See the full description on the dataset page: https://huggingface.co/datasets/CohereLabs/aya_collection_language_split.rbo_oxe_base_language_table_lerobot
Language Table (LeRobot) — Task-Pruned, Reindexed Subset
This release is a task-pruned subset of the original
IPEC-COMMUNITY/language_table_lerobot.
We subsampled by task text and rebuilt the package so it remains internally consistent
(indices, splits, stats, paths).
Robot: xArm
Modality: RGB video + states + actions
FPS / Resolution: 10 FPS, 360×640, AV1
License: apache-2.0 (inherits from source)
What’s different in this subset
Kept ~0.85% of unique tasks… See the full description on the dataset page: https://huggingface.co/datasets/saaduddinM/rbo_oxe_base_language_table_lerobot.muri-it-language-split
MURI-IT: Multilingual Instruction Tuning Dataset for 200 Languages via Multilingual Reverse Instructions
MURI-IT is a large-scale multilingual instruction tuning dataset containing 2.2 million instruction-output pairs across 200 languages. It is designed to address the challenges of instruction tuning in low-resource languages with Multilingual Reverse Instructions (MURI), which ensures that the output is human-written, high-quality, and authentic to the cultural and linguistic… See the full description on the dataset page: https://huggingface.co/datasets/akoksal/muri-it-language-split.
