datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
InternData-A1
InternData-A1
InternData-A1 is a hybrid synthetic-real manipulation dataset containing over 630k trajectories and 7,433 hours across 4 embodiments, 18 skills, 70 tasks, and 227 scenes, covering rigid, articulated, deformable, and fluid-object manipulation.
Your browser does not support the video tag.
Your browser does not support the video tag.… See the full description on the dataset page: https://huggingface.co/datasets/InternRobotics/InternData-A1.sn80-data-kc-a1A12d12s12InternData-A1-LeRobot-v3.0-by-embodimentInternData-A1 dataset taken from InternRobotics/InternData-A1,
with the tarballs extracted and directory structure "transposed" so that the top-level subdirectories are the four embodiments.
Two franka dirs
For the franka embodiment, there are two different feature spaces, so we split it into the franka-1 and franka-2 directories.
The feature spaces differ in image shape and gripper value range.
Minor fixes
Some subsets such as… See the full description on the dataset page: https://huggingface.co/datasets/griffinlabs/InternData-A1-LeRobot-v3.0-by-embodiment.epic_kitchens_100
EPIC-KITCHENS-100 (Mirror)
This repository provides a mirror of the EPIC-KITCHENS-100 dataset videos for easier access and high-speed downloading via the Hugging Face Hub.
Important Note
This mirror is uploaded for personal convenience and may not contain the entire dataset. If you need the full, official, and most up-to-date version of the dataset (including all annotations and subsets), please visit the official website.
Official Links
Official… See the full description on the dataset page: https://huggingface.co/datasets/a1raman/epic_kitchens_100.a11oy-verifiable-corpus
Part of the SZL Holdings governed estate — claims are designed to carry checkable receipts. Verification proves integrity & origin, never accuracy or performance.
a11oy — Verifiable Corpus · verify it yourself
This dataset publishes a11oy's signed receipts and proof surface so that
anyone can independently verify them — no trust in SZL Holdings required.
Every receipt here carries the full cryptographic material needed to check its
signature offline;… See the full description on the dataset page: https://huggingface.co/datasets/SZLHOLDINGS/a11oy-verifiable-corpus.Inter-Edit-Train
Inter-Edit-Train
Inter-Edit-Train is the official large-scale training set released for the CVPR 2026 paper Inter-Edit: First Benchmark for Interactive Instruction-Based Image Editing.
This dataset is designed for the Interactive Instruction-based Image Editing (I^3E) task, where a model performs localized image edits from a concise textual instruction together with imprecise spatial guidance.
Highlights
1,099,964 image editing pairs
610,186 unique source images
Four… See the full description on the dataset page: https://huggingface.co/datasets/a1557811266/Inter-Edit-Train.A100_benchmark_compile_sdpalatent_worker_early-a1_02latent_worker_early-a1_00latent_worker_early-a1_09A10_benchmark_flash_attentionlatent_worker_early-a1_07TIIF-Bench-DataWe release the images generated by the proprietary models evaluated in “🔍TIIF-Bench: How Does Your T2I Model Follow Your Instructions?”.
Produced under carefully crafted, high-quality prompts, these images form a valuable asset that can benefit the open-source community in a variety of applications🔥.
latent_worker_early-a1_05nyush-galaxea-a1-lingbot-va-real-world-evaluations
LingBot-VA on Galaxea A1 — Real-World Evaluations
Fruit-placement rollouts and open-loop diagnostics of object grounding,
layout generalization, and predicted robot motion.
Fruit step-1000: lemon-to-plate rollout in the Official layout.
Evidence
Scale
Real closed-loop rollouts
61 archived; 60 scored
Matched base-model controls
9 predictions
Post-trained diagnostics
48 full-horizon predictions; 1,211 rolling futures
Controlled OOD studies
558 predictions… See the full description on the dataset page: https://huggingface.co/datasets/pengyue-polaron/nyush-galaxea-a1-lingbot-va-real-world-evaluations.a10_cryptoresearchA100_benchmark_gpt-fasta1zrd1_x22_tran
Dataset Card for Dataset Name
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Dataset Details
Dataset Description
Curated by: [More Information Needed]
Funded by [optional]: [More Information Needed]
Shared by [optional]: [More Information Needed]
Language(s) (NLP): [More Information Needed]
License: [More Information Needed]
Dataset Sources [optional]… See the full description on the dataset page: https://huggingface.co/datasets/natW0lf/a1zrd1_x22_tran.sft_env_b9057b9c-a10c-4d1d-a360-4c79e5201dcanavsim-metric-caches-from-a100intern_a1_v3_eef
InternData-A1 v3 EEF (LeRobot v3)
This is a LeRobot v3 format conversion of InternData-A1, reorganized into a canonical bimanual 16D end-effector (EE) pose representation for large-scale VLA pre-training. 30 Hz, four embodiments, ~487.7K episodes and ~377.7M frames.
Original Dataset
InternData-A1: Pioneering High-Fidelity Synthetic Data for Pre-training Generalist Policy
InternRobotics. InternData-A1: Pioneering High-Fidelity Synthetic Data for Pre-training… See the full description on the dataset page: https://huggingface.co/datasets/GT-111/intern_a1_v3_eef.Per-400k
Per-400k Dataset
Overview
This dataset is designed for training models that generate person-consistent images: given a reference person image, the model learns to create new images of the same individual (with the same clothing and appearance), but in different backgrounds or performing different activities.
Each entry in the dataset JSON file is a dictionary containing paths to the original generated image, its two sub-images (left and right), and prompts… See the full description on the dataset page: https://huggingface.co/datasets/a1557811266/Per-400k.MICo-Bencha100_20260502
SWIFT (Scalable lightWeight Infrastructure for Fine-Tuning)
ModelScope Community Website
中文 | English
Paper | English Documentation | 中文文档
📖 Table of Contents
Groups
Introduction
News
Installation
Quick Start
Usage
License
Citation
☎ Groups
You can contact us and communicate with us by adding our group:
Discord Group
WeChat Group
📝 Introduction
🍲 ms-swift… See the full description on the dataset page: https://huggingface.co/datasets/egotools-dev/a100_20260502.ViZDoom-Deathmatch-PPO-XLrg
ViZDoom Deathmatch with pretrained PPO agent playing over 15 episodes.
A1V2H-ALIGIN2
AVH-Align
Official PyTorch Implementation of the Paper:
Ștefan Smeu, Dragoș-Alexandru Boldisor, Dan Oneață and Elisabeta OneațăCircumventing shortcuts in audio-visual deepfake detection datasets with unsupervised learningCVPR, 2025
Data
To set up your data, follow these steps:
Download the datasets:
AV-Deepfake1M(AV1M) Dataset: Follow instructions from AV-Deepfake1M
FakeAVCeleb Dataset: Follow instructions from FakeAVCeleb GitHub repo
AVLips Dataset: Follow… See the full description on the dataset page: https://huggingface.co/datasets/huahua123313/A1V2H-ALIGIN2.terminal_bench_2_a1_issue_tasks_20260805_125430latent_worker_early-a1_03latent_worker_early-a1_04
