datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
language_table_lerobotThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.0",
"robot_type": "xarm",
"total_episodes": 442226,
"total_frames": 7045476,
"total_tasks": 127605,
"total_videos": 442226,
"total_chunks": 443,
"chunks_size": 1000,
"fps": 10,
"splits": {
"train": "0:442226"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/IPEC-COMMUNITY/language_table_lerobot.Open-Sora-Plan-v1.1.0
Annotation
We resized the dataset to 1080p for easier uploading. Therefore, the original annotation file might not match the video names. Please refer to this https://github.com/PKU-YuanGroup/Open-Sora-Plan/issues/312#issuecomment-2197312973
Pexels
Pexels consists of multiple folders, but each folder exceeds the size limit for Huggingface uploads. Therefore, we divided each folder into 5 parts. You need to merge the 5 parts of each folder first, and then extract each… See the full description on the dataset page: https://huggingface.co/datasets/LanguageBind/Open-Sora-Plan-v1.1.0.github-code-2025-language-split
📜 Source Data & Attribution
This dataset is a processed derivative of nick007x/github-code-2025.
Origination
The original data was aggregated by nick007x from public GitHub repositories. We have retained the original content, file paths, and metadata while restructuring the format for easier consumption by language-specific models.
Processing Steps
To create this dataset, we performed the following processing on the source data:
Language… See the full description on the dataset page: https://huggingface.co/datasets/hasankursun/github-code-2025-language-split.aya_collection_language_split
This is a re-upload of the aya_collection, and only differs in the structure of upload. While the original aya_collection is structured by folders split according to dataset name, this dataset is split by language. We recommend you use this version of the dataset if you are only interested in downloading all of the Aya collection for a single or smaller set of languages.
Dataset Summary
The Aya Collection is a massive multilingual collection consisting of 513 million instances of… See the full description on the dataset page: https://huggingface.co/datasets/CohereLabs/aya_collection_language_split.rbo_oxe_base_language_table_lerobot
Language Table (LeRobot) — Task-Pruned, Reindexed Subset
This release is a task-pruned subset of the original
IPEC-COMMUNITY/language_table_lerobot.
We subsampled by task text and rebuilt the package so it remains internally consistent
(indices, splits, stats, paths).
Robot: xArm
Modality: RGB video + states + actions
FPS / Resolution: 10 FPS, 360×640, AV1
License: apache-2.0 (inherits from source)
What’s different in this subset
Kept ~0.85% of unique tasks… See the full description on the dataset page: https://huggingface.co/datasets/saaduddinM/rbo_oxe_base_language_table_lerobot.muri-it-language-split
MURI-IT: Multilingual Instruction Tuning Dataset for 200 Languages via Multilingual Reverse Instructions
MURI-IT is a large-scale multilingual instruction tuning dataset containing 2.2 million instruction-output pairs across 200 languages. It is designed to address the challenges of instruction tuning in low-resource languages with Multilingual Reverse Instructions (MURI), which ensures that the output is human-written, high-quality, and authentic to the cultural and linguistic… See the full description on the dataset page: https://huggingface.co/datasets/akoksal/muri-it-language-split.Language-Grounded_Sparse_Encoder_Training
Language-Grounded Sparse Encoder (LanSE) — Training Data
This repository hosts the AI-generated images and human annotation datasets accompanying the paper:
Human-like Content Analysis for Generative AI with Language-Grounded Sparse Encoders
Yiming Tang, Arash Lagzian, Srinivas Anumasa, Qiran Zou, Yingtao Zhu, Ye Zhang, Trang Nguyen, Yih-Chung Tham, Ehsan Adeli, Ching-Yu Cheng, Yilun Du, Dianbo Liu
National University of Singapore · Tsinghua University · Stanford University ·… See the full description on the dataset page: https://huggingface.co/datasets/DesmondYMTang2024/Language-Grounded_Sparse_Encoder_Training.American-Sign-Language-MNIST
Dataset Card for ASL-MNIST
This is a FiftyOne dataset with 34,627 samples of American Sign Language (ASL) alphabet images, converted from the original Kaggle Sign Language MNIST dataset into a format optimized for computer vision workflows.
Installation
If you haven't already, install FiftyOne:
pip install -U fiftyone
Usage
import fiftyone as fo
from fiftyone.utils.huggingface import load_from_hub
# Load the dataset
# Note: other available arguments… See the full description on the dataset page: https://huggingface.co/datasets/Voxel51/American-Sign-Language-MNIST.language
Dataset Card for "livebench/language"
LiveBench is a benchmark for LLMs designed with test set contamination and objective evaluation in mind. It has the following properties:
LiveBench is designed to limit potential contamination by releasing new questions monthly, as well as having questions based on recently-released datasets, arXiv papers, news articles, and IMDb movie synopses.
Each question has verifiable, objective ground-truth answers, allowing hard questions to be scored… See the full description on the dataset page: https://huggingface.co/datasets/livebench/language.wikipedia_20240401_10-languages_bge-m3_qdrant_indexThis repository contains a Qdrant index created from preprocessed and chunked Wikipedia HTML dumps from 10 languages. The embedding model used is BAAI/bge-m3
This index is compatible with WikiChat v2.0.
Refer to the following for more information:
GitHub repository: https://github.com/stanford-oval/WikiChat
Papers:
WikiChat: Stopping the Hallucination of Large Language Model Chatbots by Few-Shot Grounding on Wikipedia
SPAGHETTI: Open-Domain Question Answering from Heterogeneous Data Sources… See the full description on the dataset page: https://huggingface.co/datasets/stanford-oval/wikipedia_20240401_10-languages_bge-m3_qdrant_index.language-identification
Dataset Card for Language Identification dataset
Dataset Summary
The Language Identification dataset is a collection of 90k samples consisting of text passages and corresponding language label.
This dataset was created by collecting data from 3 sources: Multilingual Amazon Reviews Corpus, XNLI, and STSb Multi MT.
Supported Tasks and Leaderboards
The dataset can be used to train a model for language identification, which is a multi-class text classification… See the full description on the dataset page: https://huggingface.co/datasets/papluca/language-identification.language_table_sim-lerobotThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "language_table_xarm",
"total_episodes": 1000,
"total_frames": 24891,
"total_tasks": 962,
"total_videos": 1000,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 10,
"splits": {
"train": "0:1000"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/LehongWu/language_table_sim-lerobot.UniWorld-V1
The Geneval-style dataset is sourced from BLIP3o-60k.
This dataset is presented in the paper: UniWorld: High-Resolution Semantic Encoders for Unified Visual Understanding and Generation
More details can be found in UniWorld-V1
Data preparation
Download the data from LanguageBind/UniWorld-V1. The dataset consists of two parts: source images and annotation JSON files.
Prepare a data.txt file in the following format:
The first column is the root path to the image.
The second… See the full description on the dataset page: https://huggingface.co/datasets/LanguageBind/UniWorld-V1.American-Sign-Language-Dataset
American Sign Language (ASL) Dataset
Description:This dataset contains 108,618 videos representing 2,208 ASL words, with each word having a minimum of 30 videos. The videos were scraped, collected from multiple sources, and preprocessed to ensure consistency, quality, and usability for machine learning and gesture recognition tasks. Each video is ≤10 MB, optimized for storage and model training.The dataset can be used for ASL gesture recognition, video-based ML tasks, and model… See the full description on the dataset page: https://huggingface.co/datasets/shpouladi/American-Sign-Language-Dataset.openwebtext-t5language_table_train_115000_120000_augmented
language_table_train_115000_120000_augmented
Overview
Codebase version: v3.0
Robots: google_robot, images, jaco, kinova3, kuka_iiwa, panda, sawyer, ur5e
FPS: 10
Episodes: 5,000
Frames: 79,828
Splits:
train: 0:5000
Data Layout
data_path : data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet
video_path: videos/{video_key}/chunk-{chunk_index:03d}/file-{file_index:03d}.mp4
Features
Feature
dtype
shape
description… See the full description on the dataset page: https://huggingface.co/datasets/oxe-auge/language_table_train_115000_120000_augmented.language_table_train_75000_80000_augmented
language_table_train_75000_80000_augmented
Overview
Codebase version: v3.0
Robots: google_robot, images, jaco, kinova3, kuka_iiwa, panda, sawyer, ur5e
FPS: 10
Episodes: 5,000
Frames: 80,557
Splits:
train: 0:5000
Data Layout
data_path : data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet
video_path: videos/{video_key}/chunk-{chunk_index:03d}/file-{file_index:03d}.mp4
Features
Feature
dtype
shape
description… See the full description on the dataset page: https://huggingface.co/datasets/oxe-auge/language_table_train_75000_80000_augmented.language_table_train_55000_60000_augmented
language_table_train_55000_60000_augmented
Overview
Codebase version: v3.0
Robots: google_robot, images, jaco, kinova3, kuka_iiwa, panda, sawyer, ur5e
FPS: 10
Episodes: 5,000
Frames: 79,295
Splits:
train: 0:5000
Data Layout
data_path : data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet
video_path: videos/{video_key}/chunk-{chunk_index:03d}/file-{file_index:03d}.mp4
Features
Feature
dtype
shape
description… See the full description on the dataset page: https://huggingface.co/datasets/oxe-auge/language_table_train_55000_60000_augmented.language_table_train_130000_135000_augmented
language_table_train_130000_135000_augmented
Overview
Codebase version: v3.0
Robots: google_robot, images, jaco, kinova3, kuka_iiwa, panda, sawyer, ur5e
FPS: 10
Episodes: 5,000
Frames: 79,357
Splits:
train: 0:5000
Data Layout
data_path : data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet
video_path: videos/{video_key}/chunk-{chunk_index:03d}/file-{file_index:03d}.mp4
Features
Feature
dtype
shape
description… See the full description on the dataset page: https://huggingface.co/datasets/oxe-auge/language_table_train_130000_135000_augmented.language_table_train_140000_145000_augmented
language_table_train_140000_145000_augmented
Overview
Codebase version: v3.0
Robots: google_robot, images, jaco, kinova3, kuka_iiwa, panda, sawyer, ur5e
FPS: 10
Episodes: 5,000
Frames: 80,121
Splits:
train: 0:5000
Data Layout
data_path : data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet
video_path: videos/{video_key}/chunk-{chunk_index:03d}/file-{file_index:03d}.mp4
Features
Feature
dtype
shape
description… See the full description on the dataset page: https://huggingface.co/datasets/oxe-auge/language_table_train_140000_145000_augmented.language_table_train_345000_350000_augmented
language_table_train_345000_350000_augmented
Overview
Codebase version: v3.0
Robots: google_robot, images, jaco, kinova3, kuka_iiwa, panda, sawyer, ur5e
FPS: 10
Episodes: 5,000
Frames: 79,459
Splits:
train: 0:5000
Data Layout
data_path : data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet
video_path: videos/{video_key}/chunk-{chunk_index:03d}/file-{file_index:03d}.mp4
Features
Feature
dtype
shape
description… See the full description on the dataset page: https://huggingface.co/datasets/oxe-auge/language_table_train_345000_350000_augmented.language_table_train_5000_10000_augmented
language_table_train_5000_10000_augmented
Overview
Codebase version: v3.0
Robots: google_robot, images, jaco, kinova3, kuka_iiwa, panda, sawyer, ur5e
FPS: 10
Episodes: 5,000
Frames: 79,151
Splits:
train: 0:5000
Data Layout
data_path : data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet
video_path: videos/{video_key}/chunk-{chunk_index:03d}/file-{file_index:03d}.mp4
Features
Feature
dtype
shape
description… See the full description on the dataset page: https://huggingface.co/datasets/oxe-auge/language_table_train_5000_10000_augmented.language_table_train_110000_115000_augmented
language_table_train_110000_115000_augmented
Overview
Codebase version: v3.0
Robots: google_robot, images, jaco, kinova3, kuka_iiwa, panda, sawyer, ur5e
FPS: 10
Episodes: 5,000
Frames: 80,139
Splits:
train: 0:5000
Data Layout
data_path : data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet
video_path: videos/{video_key}/chunk-{chunk_index:03d}/file-{file_index:03d}.mp4
Features
Feature
dtype
shape
description… See the full description on the dataset page: https://huggingface.co/datasets/oxe-auge/language_table_train_110000_115000_augmented.language_table_train_125000_130000_augmented
language_table_train_125000_130000_augmented
Overview
Codebase version: v3.0
Robots: google_robot, images, jaco, kinova3, kuka_iiwa, panda, sawyer, ur5e
FPS: 10
Episodes: 5,000
Frames: 78,601
Splits:
train: 0:5000
Data Layout
data_path : data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet
video_path: videos/{video_key}/chunk-{chunk_index:03d}/file-{file_index:03d}.mp4
Features
Feature
dtype
shape
description… See the full description on the dataset page: https://huggingface.co/datasets/oxe-auge/language_table_train_125000_130000_augmented.language_table_train_95000_100000_augmented
language_table_train_95000_100000_augmented
Overview
Codebase version: v3.0
Robots: google_robot, images, jaco, kinova3, kuka_iiwa, panda, sawyer, ur5e
FPS: 10
Episodes: 5,000
Frames: 79,860
Splits:
train: 0:5000
Data Layout
data_path : data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet
video_path: videos/{video_key}/chunk-{chunk_index:03d}/file-{file_index:03d}.mp4
Features
Feature
dtype
shape
description… See the full description on the dataset page: https://huggingface.co/datasets/oxe-auge/language_table_train_95000_100000_augmented.language_table_train_15000_20000_augmented
language_table_train_15000_20000_augmented
Overview
Codebase version: v3.0
Robots: google_robot, images, jaco, kinova3, kuka_iiwa, panda, sawyer, ur5e
FPS: 10
Episodes: 5,000
Frames: 79,558
Splits:
train: 0:5000
Data Layout
data_path : data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet
video_path: videos/{video_key}/chunk-{chunk_index:03d}/file-{file_index:03d}.mp4
Features
Feature
dtype
shape
description… See the full description on the dataset page: https://huggingface.co/datasets/oxe-auge/language_table_train_15000_20000_augmented.language_table_train_120000_125000_augmented
language_table_train_120000_125000_augmented
Overview
Codebase version: v3.0
Robots: google_robot, images, jaco, kinova3, kuka_iiwa, panda, sawyer, ur5e
FPS: 10
Episodes: 5,000
Frames: 80,140
Splits:
train: 0:5000
Data Layout
data_path : data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet
video_path: videos/{video_key}/chunk-{chunk_index:03d}/file-{file_index:03d}.mp4
Features
Feature
dtype
shape
description… See the full description on the dataset page: https://huggingface.co/datasets/oxe-auge/language_table_train_120000_125000_augmented.language_table_train_135000_140000_augmented
language_table_train_135000_140000_augmented
Overview
Codebase version: v3.0
Robots: google_robot, images, jaco, kinova3, kuka_iiwa, panda, sawyer, ur5e
FPS: 10
Episodes: 5,000
Frames: 79,556
Splits:
train: 0:5000
Data Layout
data_path : data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet
video_path: videos/{video_key}/chunk-{chunk_index:03d}/file-{file_index:03d}.mp4
Features
Feature
dtype
shape
description… See the full description on the dataset page: https://huggingface.co/datasets/oxe-auge/language_table_train_135000_140000_augmented.language_table_train_390000_395000_augmented
language_table_train_390000_395000_augmented
Overview
Codebase version: v3.0
Robots: google_robot, images, jaco, kinova3, kuka_iiwa, panda, sawyer, ur5e
FPS: 10
Episodes: 5,000
Frames: 79,774
Splits:
train: 0:5000
Data Layout
data_path : data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet
video_path: videos/{video_key}/chunk-{chunk_index:03d}/file-{file_index:03d}.mp4
Features
Feature
dtype
shape
description… See the full description on the dataset page: https://huggingface.co/datasets/oxe-auge/language_table_train_390000_395000_augmented.language_table_train_10000_15000_augmented
language_table_train_10000_15000_augmented
Overview
Codebase version: v3.0
Robots: google_robot, images, jaco, kinova3, kuka_iiwa, panda, sawyer, ur5e
FPS: 10
Episodes: 5,000
Frames: 79,689
Splits:
train: 0:5000
Data Layout
data_path : data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet
video_path: videos/{video_key}/chunk-{chunk_index:03d}/file-{file_index:03d}.mp4
Features
Feature
dtype
shape
description… See the full description on the dataset page: https://huggingface.co/datasets/oxe-auge/language_table_train_10000_15000_augmented.
