datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
AVGen-Bench
AVGen-Bench Generated Videos Data Card
Overview
This data card describes the generated audio-video outputs stored directly in the repository root by model directory.
The collection is intended for benchmarking and qualitative/quantitative evaluation of text-to-audio-video (T2AV) systems. It was presented in the paper AVGen-Bench: A Task-Driven Benchmark for Multi-Granular Evaluation of Text-to-Audio-Video Generation. It is not a training dataset. Each item is a… See the full description on the dataset page: https://huggingface.co/datasets/microsoft/AVGen-Bench.simulation-package
Simulation runtime: the third-party half of the AVSim data generation package
This repository holds the third-party runtime of the Simulation data generation package, version 3: the pieces the pipeline needs that were not written by the authors. It contains no code and no data of the authors. It is not usable on its own: the private half of the package, AVSim/simulation, downloads this repository into the same directory at a pinned revision with its fetch_runtime.sh script and… See the full description on the dataset page: https://huggingface.co/datasets/AVSim/simulation-package.AVA-Huggingface
AVA-Huggingface
This repository contains a Hugging Face dataset built from the AVA (Aesthetic Visual Analysis) dataset. The dataset includes images along with their aesthetic scores, total votes, and rating distributions. The data is prepared by filtering out images with fewer than 50 votes and stratifying them based on the computed mean aesthetic score.
Dataset Overview
Image ID: Unique identifier for each image.
Image: The actual image loaded from disk.
Mean… See the full description on the dataset page: https://huggingface.co/datasets/trojblue/AVA-Huggingface.AVM_Segmentation_train
Dataset Card for AVM (Around View Monitoring) Semantic Segmentation Dataset
This repository provides a FiftyOne-compatible version of the AVM semantic segmentation dataset for autonomous parking systems, with enhanced metadata and visualization capabilities.
This is a FiftyOne dataset with 6763 samples.
Installation
If you haven't already, install FiftyOne:
pip install -U fiftyone
Usage
import fiftyone as fo
from fiftyone.utils.huggingface import… See the full description on the dataset page: https://huggingface.co/datasets/Voxel51/AVM_Segmentation_train.PInVerify
PInVerify Dataset
An offline embodied benchmark for Active Instance Verification (AIV).
Paper
arXiv:2605.30639
Code
github.com/Avalon-S/PInVerify
Project page
avalon-s.github.io/PInVerify
Venue
FMEA Workshop @ CVPR 2026 (Poster)
Overview
An agent that navigates to a target object does not always arrive at the right instance. Telling "white floral" from "white striped" takes a close look from more than one viewpoint, which is a separate… See the full description on the dataset page: https://huggingface.co/datasets/Avalon-S/PInVerify.Physical-AI-AV-US
PhysicalAI-AV-SFT
Supervised fine-tuning (SFT) dataset for an autonomous-vehicle vision-language
waypoint-prediction model. Contains 2,789,773 samples from 150 000
driving scenes (18 seconds per scene, sampled at 1 Hz) recorded in the
United States.
Format
WebDataset — 100 uncompressed .tar shards,
each containing pairs of files per sample:
Entry
Description
{key}.png
Front-facing wide-angle camera frame (640 × 360 px)
{key}.json
Metadata (see schema below)… See the full description on the dataset page: https://huggingface.co/datasets/tom-jerry-123/Physical-AI-AV-US.minimax_h3_avatar_500
Watch the full 500-video showcase on YouTube
MiniMax H3 Avatar 500
An image-to-video dataset pairing reference avatar images with detailed generation prompts and generated avatar videos. This release contains 500 curated examples in both a browsable raw layout and a typed Hugging Face dataset.
Version 1.0 · Released August 14, 2026
Dataset contents
Each example contains:
A 1024 × 1024 reference avatar image
A detailed English generation prompt
A generated 640 ×… See the full description on the dataset page: https://huggingface.co/datasets/oakmindai/minimax_h3_avatar_500.AVAAVA: A Large-Scale Database for Aesthetic Visual Analysis
See Github Page for tags.
Citation
@inproceedings{murray2012ava,
title={AVA: A large-scale database for aesthetic visual analysis},
author={Murray, Naila and Marchesotti, Luca and Perronnin, Florent},
booktitle={CVPR},
year={2012},
}
DREAMS-AVATAR
DREAMS-AVATAR
The DREAMS-Avatar dataset from the DEGAS paper
(3DV 2025), re-registered to pure SMPL-X.
These are the same multiview captures introduced as the DREAMS-Avatar dataset in DEGAS
(Fig. 1b); what is new here is the registration.
32 calibrated, matted camera views of a full-body performance, with one SMPL-X body fitted
to all views at once by our multiview tracker: 300 shape coefficients, 100 expression
coefficients, jaw and both eyes, hands as free 45-dim axis-angle… See the full description on the dataset page: https://huggingface.co/datasets/initialneil/DREAMS-AVATAR.wildlife_in_irrigation_ponds
Wildlife in Irrigation Ponds Dataset
Dataset Summary
This dataset supports the training and evaluation of object detection models for monitoring irrigation ponds, with the goal of detecting people and animals that have fallen into the water. It comprises synthetically generated images produced using state-of-the-art diffusion models (Z-Image, FLUX), with a real photograph of a target irrigation pond used as the background.
The dataset includes four object classes… See the full description on the dataset page: https://huggingface.co/datasets/grupo-avispa/wildlife_in_irrigation_ponds.Illustrious_Lora_LegacyA directory of all my old Illustrious models (mostly the ones posted on Civitai).
They are also hosted here: https://civitai.red/user/Ava_Choco (but censorship is so brutal that it alsmost flags anything anime/manga)
Your should be able to dig up the activation tag from the images or from the meta data of the lora.
If you can't find it that way, let me know and I will look it up.
LoRA Usage Disclaimer
This LoRA model is provided as-is for non-commercial use only.
Important: The user does not… See the full description on the dataset page: https://huggingface.co/datasets/Ava2000/Illustrious_Lora_Legacy.food_not_food
Food vs Not Food Dataset (from Hugging Face ImageNet-1K)
This dataset is a binary classification subset derived from the Hugging Face imagenet-1k dataset. It is curated to support the task of distinguishing food images from non-food images.
📦 Dataset Overview
Source: imagenet-1k on Hugging Face Datasets
Classes:
food: 40 selected ImageNet classes representing food items (e.g., pizza, banana, hotdog)
not_food: 40 selected classes not related to food (e.g., car, clock… See the full description on the dataset page: https://huggingface.co/datasets/avnishs17/food_not_food.llava-15-rlmpq-vlm-eval-results
RL-MPQ VLM Evaluation Artifacts
Complete figures, tables, galleries, and raw benchmark CSVs for the extended VLM evaluation.
Dataset: AvoCahDoe/llava-15-rlmpq-vlm-eval-results
Collections (by base VLM)
RL-MPQ VLM — LLaVA-1.5-13B — HF collection
RL-MPQ VLM — LLaVA-1.5-7B — HF collection
RL-MPQ VLM — LLaVA-Next Mistral-7B — HF collection
RL-MPQ VLM — Qwen2-VL-7B — HF collection
Model repos
RL-MPQ High Fidelity →… See the full description on the dataset page: https://huggingface.co/datasets/AvoCahDoe/llava-15-rlmpq-vlm-eval-results.AVI-Math
Dataset Sources
Repository: https://github.com/VisionXLab/avi-math
Paper: https://arxiv.org/abs/2509.10059
BibTeX:
@ARTICLE{zhou2025avimath,
author={Zhou, Yue and Feng, Litong and Lan, Mengcheng and Yang, Xue and Li, Qingyun and Ke, Yiping and Jiang, Xue and Zhang, Wayne},
journal={ISPRS Journal of Photogrammetry and Remote Sensing},
title={Multimodal Mathematical Reasoning Embedded in Aerial Vehicle Imagery: Benchmarking, Analysis, and Exploration},
year={2025}… See the full description on the dataset page: https://huggingface.co/datasets/erenzhou/AVI-Math.the-welding-defect-dataset-v2AVA-aesthetics-10pct-min50-10bins
AVA Aesthetics 10% Subset (min50, 10 bins)
This dataset is a curated 10% subset of the AVA Aesthetics Dataset (or the original AVA dataset as described in Murray et al., 2012). It includes images that have at least 50 total votes and have been stratified into 10 bins based on their computed mean aesthetic scores.
Dataset Overview
Dataset Name: AVA Aesthetics 10% Subset (min50, 10 bins)
Subset Size: 10% of the original AVA dataset (after filtering for a minimum of 50… See the full description on the dataset page: https://huggingface.co/datasets/trojblue/AVA-aesthetics-10pct-min50-10bins.hass_avocadoThis dataset is a huggingface upload for https://data.mendeley.com/datasets/3xd9n945v8/1
Description from their website accessed on 2nd July 2025
This dataset consists of 14,710 labeled photographs of Hass avocados (Persea Americana Mill. cv Hass), resized to 800 x 800 pixels and saved in the .jpg format, designed to facilitate the development of deep learning models for predicting ripening stages and estimating shelf-life.
A total of 478 Hass avocados were acquired three days post-harvest and… See the full description on the dataset page: https://huggingface.co/datasets/c2p-cmd/hass_avocado.CaptchasNurisk-ICRA2026
Nurisk: VQA for Risk Assessment in Autonomous Driving
Nurisk is a visual question answering dataset focusing on risk assessment for autonomous driving. Each row contains:
image: a BEV image
question: a driving-related question
answer: the ground truth answer
Paper
NuRisk: A Visual Question Answering Dataset for Agent-Level Risk Assessment in Autonomous Driving — see the paper on arXiv:2509.25944 .
Framework
Dataset Structure… See the full description on the dataset page: https://huggingface.co/datasets/TUM-AVS/Nurisk-ICRA2026.Pony_LoraA directory of all my old Pony models.
They are also hosted here: https://civitai.red/user/Ava_Choco (but censorship is so brutal that it alsmost flags anything anime/manga)
Your should be able to dig up the activation tag from the images or from the meta data of the lora.
If you can't find it that way, let me know and I will look it up.
LoRA Usage Disclaimer
This LoRA model is provided as-is for non-commercial use only.
Important: The user does not claim ownership of the training data used… See the full description on the dataset page: https://huggingface.co/datasets/Ava2000/Pony_Lora.avs-spot
Dataset Card for AVS-Spot Benchmark
This dataset is associated with the paper: "Understanding Co-Speech Gestures in-the-wild"
📝 ArXiv: https://arxiv.org/abs/2503.22668
🌐 Project page: https://www.robots.ox.ac.uk/~vgg/research/jegal
💻 Code: https://github.com/Sindhu-Hegde/jegal
We present JEGAL, a Joint Embedding space for Gestures, Audio and Language. Our semantic gesture representations can be used to perform multiple downstream tasks such as cross-modal retrieval… See the full description on the dataset page: https://huggingface.co/datasets/sindhuhegde/avs-spot.AV_Odyssey_BenchOfficial dataset for the paper "AV-Odyssey: Can Your Multimodal LLMs Really Understand Audio-Visual Information?".
🌟 For more details, please refer to the project page with data examples: https://av-odyssey.github.io/.
[🌐 Webpage] [📖 Paper] [🤗 Huggingface AV-Odyssey Dataset] [🤗 Huggingface Deaftest Dataset] [🏆 Leaderboard]
🔥 News
2024.11.24 🌟 We release AV-Odyssey, the first-ever comprehensive evaluation benchmark to explore whether MLLMs really understand audio-visual… See the full description on the dataset page: https://huggingface.co/datasets/AV-Odyssey/AV_Odyssey_Bench.portrait_2_avatar
🖼️ Portrait to Anime Style Tranfer Data
This dataset consists of paired human and corresponding anime-style images, accompanied by descriptive prompts. The human images are sourced from the CelebA dataset, and the anime-style counterparts were generated using a combination of state-of-the-art GAN architectures and diffusion models.
It is designed to support a wide range of tasks,
GAN research
Diffusion model fine-tuning
Model evaluation
Benchmarking for image-to-image and… See the full description on the dataset page: https://huggingface.co/datasets/murali1729S/portrait_2_avatar.Nurisk
Nurisk: VQA for Risk Assessment in Autonomous Driving
Nurisk is a visual question answering dataset focusing on risk assessment for autonomous driving. Each row contains:
image: a BEV image
question: a driving-related question
answer: the ground truth answer
Paper
NuRisk: A Visual Question Answering Dataset for Agent-Level Risk Assessment in Autonomous Driving — see the paper on arXiv:2509.25944 .
Framework
Dataset Structure… See the full description on the dataset page: https://huggingface.co/datasets/Yuan-avs/Nurisk.AVSBenchpanda-pick-place-lerobot-14-18-58_01-06-2026
Franka Panda Pick-and-Place — LeRobot v3 Dataset
Visuomotor behavior-cloning dataset collected in MuJoCo with a simulated Franka Emika Panda arm.
Recorded in LeRobot v3 format (Parquet + MP4 shards).
Load
from lerobot.datasets import LeRobotDataset
ds = LeRobotDataset("av120/panda-pick-place-lerobot-14-18-58_01-06-2026")
Task
Pick up a cube and place it ~30 cm to the side using a scripted IK state-machine demonstrator.
Box position is randomized ±5… See the full description on the dataset page: https://huggingface.co/datasets/av120/panda-pick-place-lerobot-14-18-58_01-06-2026.melascope-showcasedfki-av-gopro-hand-objectAVDD-TCMI-dataset
AVDD-TCMI Dataset
Introduction
The AVDD-TCMI dataset is a multimodal dataset for depression detection. Our dataset comprises a total of 1,230 valid samples, including 969 non-depressed individuals and 261 depressed individuals. It contains two primary subsets:
HDS (Hospital Depression Subset): HDS is collected from volunteers at West China Hospital of Sichuan University and includes 282 valid samples, comprising 159 non-depressed volunteers (including doctors, nurses… See the full description on the dataset page: https://huggingface.co/datasets/Mar-rill/AVDD-TCMI-dataset.RWTD-COCO
RWTD-COCO
Single natural appearance transitions built from COCO-Stuff by deterministic reuse of human annotation.
One of the four evaluation routes in the ICLR 2027 submission on sub-semantic
image segmentation: partitioning an image into regions that are coherent in
appearance and describable in language, but that need not correspond to any
object, part or material class.
Images: 256
Code: github.com/aviadcohz/Qwen2SAM_Detecture_Benchmark
Weights: aviadcohz/Detecture-ICLR-2027… See the full description on the dataset page: https://huggingface.co/datasets/aviadcohz/RWTD-COCO.
