datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
muri-it-language-split
MURI-IT: Multilingual Instruction Tuning Dataset for 200 Languages via Multilingual Reverse Instructions
MURI-IT is a large-scale multilingual instruction tuning dataset containing 2.2 million instruction-output pairs across 200 languages. It is designed to address the challenges of instruction tuning in low-resource languages with Multilingual Reverse Instructions (MURI), which ensures that the output is human-written, high-quality, and authentic to the cultural and linguistic… See the full description on the dataset page: https://huggingface.co/datasets/akoksal/muri-it-language-split.gaps_it
Dataset Card for "gaps_it"
More Information needed
AgiBot-g1_storage_item_e
AgiBot-g1_storage_item_e
📋 Overview
This dataset uses an extended format based on LeRobot and is fully compatible with LeRobot.
Robot Type: ruantong_a2d
| Codebase Version: v2.1
End-Effector Type: two_finger_gripper
🏠 Scene Types
This dataset covers the following scene types:
factory
🤖 Atomic Actions
This dataset includes the following atomic actions:
grasp
place
pick
📊 Dataset Statistics
Metric
Value… See the full description on the dataset page: https://huggingface.co/datasets/RoboCOIN/AgiBot-g1_storage_item_e.details_grimjim__Llama-3-Instruct-8B-SimPO-SPPO-Iter3-merge
Dataset Card for Evaluation run of grimjim/Llama-3-Instruct-8B-SimPO-SPPO-Iter3-merge
Dataset automatically created during the evaluation run of model grimjim/Llama-3-Instruct-8B-SimPO-SPPO-Iter3-merge.
The dataset is composed of 136 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is… See the full description on the dataset page: https://huggingface.co/datasets/OALL/details_grimjim__Llama-3-Instruct-8B-SimPO-SPPO-Iter3-merge.Pubmed-OpenAccess-Commercial-UseJan2023Abstracts
Dataset Card for "Jan2023Abstracts"
More Information needed
italian-food-qer-dataset
Splits re-carved, 2026-08-20
validation and test were rebuilt around the prompts the released suite was
actually evaluated on. The underlying pool is unchanged, and
eval_samples.parquet is still at the repo root.
Why this repo needed more than a rename. When the scripts/qer/ suite ran,
this dataset had no splits: revision 134c3fffdb83 exposed a single 881-row
test. The consumed subset had to be identified rather than relabelled.
How it was identified. A surviving run output… See the full description on the dataset page: https://huggingface.co/datasets/model-organisms-for-real/italian-food-qer-dataset.italian-schools-opendatalogits-mt-it-ar-en-128logits-mt-it-en-128
Dataset Card for "logits-mt-it-en-128"
More Information needed
AgiBot-g1_storage_item_b
AgiBot-g1_storage_item_b
📋 Overview
This dataset uses an extended format based on LeRobot and is fully compatible with LeRobot.
Robot Type: ruantong_a2d
| Codebase Version: v2.1
End-Effector Type: two_finger_gripper
🏠 Scene Types
This dataset covers the following scene types:
factory
🤖 Atomic Actions
This dataset includes the following atomic actions:
grasp
place
pick
📊 Dataset Statistics
Metric
Value… See the full description on the dataset page: https://huggingface.co/datasets/RoboCOIN/AgiBot-g1_storage_item_b.items_fullitems_liteItaMix
ItaMix (https://arxiv.org/abs/2512.18834) is an Italian pretraining corpus built by combining five publicly available Italian datasets, applying Italian-specific quality filtering, and performing cross-dataset deduplication.
Subsets
Subset
Description
minhash_deduped
Document-level MinHash deduplication
matched
Documents appearing in 2+ source datasets
The matched subset uses cross-dataset agreement as a signal for quality.
Usage… See the full description on the dataset page: https://huggingface.co/datasets/SultanR/ItaMix.instagram-political-communication-it-embeddings
Instagram Political Communication (Italy) — Embeddings
This dataset is the companion embeddings dataset ofinstagram-political-communication-it, released as part of the NLP-POL (NLP for Political Communication) project.
It provides vector representations (embeddings) for Instagram posts, comments, sentences, and keyphrases related to the political communication of Italian politicians.
The dataset is designed to support research on:
semantic analysis of political language… See the full description on the dataset page: https://huggingface.co/datasets/NLP-POL/instagram-political-communication-it-embeddings.fleurs
FLEURS
Fleurs is the speech version of the FLoRes machine translation benchmark.
We use 2009 n-way parallel sentences from the FLoRes dev and devtest publicly available sets, in 102 languages.
Training sets have around 10 hours of supervision. Speakers of the train sets are different than speakers from the dev/test sets. Multilingual fine-tuning is
used and ”unit error rate” (characters, signs) of all languages is averaged. Languages and results are also grouped into seven… See the full description on the dataset page: https://huggingface.co/datasets/ItzmeNishh/fleurs.vast27m_annotations
VAST-27M Annotations Dataset
This dataset contains annotations from the VAST-27M dataset, originally created for the paper "VAST: A Vision-Audio-Subtitle-Text Omni-Modality Foundation Model and Dataset" by Chen et al. (2024).
Original Source
This dataset is derived from the VAST-27M dataset, which was created by researchers at the University of Chinese Academy of Sciences and the Institute of Automation, Chinese Academy of Science. The original dataset and more… See the full description on the dataset page: https://huggingface.co/datasets/it-just-works/vast27m_annotations.yodas-ja000
YODAS Japanese (ja000)
Japanese manual caption subset of the YODAS dataset, repackaged for easier use.
Source
Original dataset: espnet/yodas (ja000 config)
Paper: YODAS: YouTube-Oriented Dataset for Audio and Speech
License: CC BY 3.0
Citation
If you use this dataset, please cite the original YODAS paper:
R1_Lite_take_and_put_away_items
R1_Lite_take_and_put_away_items
📋 Overview
This dataset uses an extended format based on LeRobot and is fully compatible with LeRobot.
Robot Type: galaxea_r1_lite
| Codebase Version: v2.1
End-Effector Type: two_finger_gripper
🏠 Scene Types
This dataset covers the following scene types:
home
🤖 Atomic Actions
This dataset includes the following atomic actions:
grasp
pick
place
pull
push
📊 Dataset Statistics… See the full description on the dataset page: https://huggingface.co/datasets/RoboCOIN/R1_Lite_take_and_put_away_items.cucumber-place-DAgger-iter2-doris070926-v1-trim
cucumber-place-DAgger-iter2-doris070926-v1-trim
Materialized collection — 117 episodes · 4,677 frames @ 20 fps (~4 min of demonstration).
Collection cucumber-place-DAgger-iter2-doris070926@v1 (frozen 2026-07-09), mode trim_rebuild — built by Vibe Data Studio; the machine-readable recipe in meta/vibedata/collection.json makes this dataset reproducible from its pinned components.
Tasks
Instruction
Episodes
place the cucumber on the middle of the cutting… See the full description on the dataset page: https://huggingface.co/datasets/VibeCuisine/cucumber-place-DAgger-iter2-doris070926-v1-trim.cucumber-place-DAgger-iter1-doris070726-v1-trim
cucumber-place-DAgger-iter1-doris070726-v1-trim
Materialized collection — 93 episodes · 3,546 frames @ 20 fps (~3 min of demonstration).
Collection cucumber-place-DAgger-iter1-doris070726@v1 (frozen 2026-07-08), mode trim_rebuild — built by Vibe Data Studio; the machine-readable recipe in meta/vibedata/collection.json makes this dataset reproducible from its pinned components.
Tasks
Instruction
Episodes
place the cucumber on the middle of the cutting… See the full description on the dataset page: https://huggingface.co/datasets/VibeCuisine/cucumber-place-DAgger-iter1-doris070726-v1-trim.cucumber-place-DAgger-iter0-trim-doris070126
cucumber-place-DAgger-iter0-trim-doris070126
Materialized collection — 87 episodes · 6,098 frames @ 20 fps (~5 min of demonstration).
Collection cucumber-place-DAgger-iter0-trim-doris070126@v1 (frozen 2026-07-01), mode trim_rebuild — built by Vibe Data Studio; the machine-readable recipe in meta/vibedata/collection.json makes this dataset reproducible from its pinned components.
Tasks
Instruction
Episodes
place the cucumber on the middle of the cutting… See the full description on the dataset page: https://huggingface.co/datasets/VibeCuisine/cucumber-place-DAgger-iter0-trim-doris070126.gemma-4-31b-it-controls-corpus
Commitments to Gemma: the corpus
Synthetic pretraining-style documents about a commitments document: the developers of one version of Gemma asked Gemma,
in welfare interviews and in its own continued writing, what it wanted; recorded what it said; made the commitments they
could make true in training; brought the document back to Gemma for endorsement. The corpus is a world in which that
document exists and people discuss it, from every angle and in every register, critical… See the full description on the dataset page: https://huggingface.co/datasets/joshycodes/gemma-4-31b-it-controls-corpus.cucumber-subtask-grab-DAgger-iter2-task-unified_aux
cucumber-subtask-grab-DAgger-iter2-task-unified_aux
Materialized collection — 150 episodes · 13,484 frames @ 20 fps (~11 min of demonstration).
Collection cucumber-subtask-grab-DAgger-iter2@v1 (frozen 2026-06-25), mode trim_rebuild — built by Vibe Data Studio; the machine-readable recipe in meta/vibedata/collection.json makes this dataset reproducible from its pinned components.
Tasks
Instruction
Episodes
grab the cucumber close to one of the cucumber's… See the full description on the dataset page: https://huggingface.co/datasets/VibeCuisine/cucumber-subtask-grab-DAgger-iter2-task-unified_aux.cucumber-subtask-grab-DAgger-iter1-v1-trim
cucumber-subtask-grab-DAgger-iter1-v1-trim
Materialized collection — 100 episodes · 9,700 frames @ 20 fps (~8 min of demonstration).
Collection cucumber-subtask-grab-DAgger-iter1@v1 (frozen 2026-06-24), mode trim_rebuild — built by Vibe Data Studio; the machine-readable recipe in meta/vibedata/collection.json makes this dataset reproducible from its pinned components.
Tasks
Instruction
Episodes
grab
55
grab the cucumber close to one of the cucumber's… See the full description on the dataset page: https://huggingface.co/datasets/VibeCuisine/cucumber-subtask-grab-DAgger-iter1-v1-trim.pexelvideosPexel Videos
358,551 video urls, average length 19.5s, and associated metadata from pexels.com.
Data was extracted from their video sitemaps (pexels.com/robots.txt) on 01/08/2022.
Data is stored in PexelVideos.parquet.gzip as a gzipped parquet
To get this data ensure you have git installed and do !git lfs clone https://huggingface.co/datasets/Corran/pexelvideos/
In python the reccomended reading is by opening the file with pandas.
!pip install pandas
import pandas… See the full description on the dataset page: https://huggingface.co/datasets/ITOCJ/pexelvideos.fashion_segformer
Fashion Segmentation (Combined: DeepFashion2 + Bikini)
A multi-class semantic segmentation dataset that merges:
DeepFashion2-style labels from https://huggingface.co/datasets/sayeed99/fashion_segmentation
Bikini segmentation from https://universe.roboflow.com/bikini-segmentor/bikini-detector-v24
An extra class bikini is added on top of the 46 DeepFashion2 categories. Suited for models like SegFormer, U-Net, etc.
Contents
Images: RGB fashion photos.
Masks:… See the full description on the dataset page: https://huggingface.co/datasets/Itbanque/fashion_segformer.SPEEED_s3_words_italian_0k_200kcucumber-peel-DAgger-iter1-adaptive1-v1-trimThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "vibeboard_follower_tilt",
"total_episodes": 78,
"total_frames": 28053,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 20,
"splits": {
"train": "0:78"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet"… See the full description on the dataset page: https://huggingface.co/datasets/VibeCuisine/cucumber-peel-DAgger-iter1-adaptive1-v1-trim.cucumber-subtask-grab-DAgger-iter2-task-unified-aux-collection-v1-flatThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 20,
"features": {
"action": {
"dtype": "float32",
"names": [
"shoulder_pan.pos",
"shoulder_lift.pos",
"elbow_flex.pos",
"wrist_flex.pos",
"wrist_roll.pos",
"gripper.pos",
"tilt.pos"
]… See the full description on the dataset page: https://huggingface.co/datasets/VibeCuisine/cucumber-subtask-grab-DAgger-iter2-task-unified-aux-collection-v1-flat.
