datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Physical-AI-AV-US
PhysicalAI-AV-SFT
Supervised fine-tuning (SFT) dataset for an autonomous-vehicle vision-language
waypoint-prediction model. Contains 2,789,773 samples from 150 000
driving scenes (18 seconds per scene, sampled at 1 Hz) recorded in the
United States.
Format
WebDataset — 100 uncompressed .tar shards,
each containing pairs of files per sample:
Entry
Description
{key}.png
Front-facing wide-angle camera frame (640 × 360 px)
{key}.json
Metadata (see schema below)… See the full description on the dataset page: https://huggingface.co/datasets/tom-jerry-123/Physical-AI-AV-US.mmtrailer-pe-av-unimodal-local-embeddingsAVYTyt-pe-av-unimodal-local-embeddingsmovielens-pe-av-local-embeddingsyt-pe-av-local-embeddingsmovielens-pe-av-unimodal-local-embeddingsAVCocktail
AVSRCocktail: Audio-Visual Speech Recognition for Cocktail Party Scenarios
Official implementation of "Cocktail-Party Audio-Visual Speech Recognition" (Interspeech 2025).
A robust audio-visual speech recognition system designed for multi-speaker environments and noisy cocktail party scenarios. The model combines lip reading and audio processing to achieve superior performance in challenging acoustic conditions with background noise and speaker interference.
Getting… See the full description on the dataset page: https://huggingface.co/datasets/nguyenvulebinh/AVCocktail.mmtrailer-pe-av-local-embeddingsAVQA-R1-6KThis repository contains data presented in EchoInk-R1: Exploring Audio-Visual Reasoning in Multimodal LLMs via Reinforcement Learning.
For training and inference, please refer to the Code: https://github.com/HarryHsing/EchoInk
Data Format in AVQA-R1-6K:
{
"problem_id": 0,
"problem": "What is the source of the sound in the video?",
"data_type": "image_audio",
"problem_type": "multiple choice",
"options": [
"A. motorcycle",
"B. automobile"… See the full description on the dataset page: https://huggingface.co/datasets/harryhsing/AVQA-R1-6K.Physical-AI-AV-FR
PhysicalAI-AV-SFT
Supervised fine-tuning (SFT) dataset for an autonomous-vehicle vision-language
waypoint-prediction model. Contains 29,909 samples from 150 000
driving scenes (18 seconds per scene, sampled at anchor times 2s..16s) recorded in the
United States.
Format
WebDataset — 3 uncompressed .tar shards,
each containing pairs of files per sample:
Entry
Description
{key}.jpg
Front-facing wide-angle camera frame (JPEG quality 95, 640 × 360 px)
{key}.json… See the full description on the dataset page: https://huggingface.co/datasets/tom-jerry-123/Physical-AI-AV-FR.Cosmos-Transfer1-7B-Sample-AV-Data-Example
Cosmos-Transfer1-7B-Sample-AV-Data-Example
Cosmos | Code | Paper | Paper Website
Dataset Description:
This dataset contains 10 sample data points intended to help users better utilize our Cosmos-Transfer1-7B-Sample-AV model. It includes HD Map annotations and LiDAR data, with no personally identifiable information such as faces or license plates. This dataset is intended for research and development only.
Dataset Owner(s):
NVIDIA
Dataset Creation… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Cosmos-Transfer1-7B-Sample-AV-Data-Example.Physical-AI-AV-ES
PhysicalAI-AV-SFT
Supervised fine-tuning (SFT) dataset for an autonomous-vehicle vision-language
waypoint-prediction model. Contains 29,674 samples from 150 000
driving scenes (18 seconds per scene, sampled at anchor times 2s..16s) recorded in the
United States.
Format
WebDataset — 3 uncompressed .tar shards,
each containing pairs of files per sample:
Entry
Description
{key}.jpg
Front-facing wide-angle camera frame (JPEG quality 95, 640 × 360 px)
{key}.json… See the full description on the dataset page: https://huggingface.co/datasets/tom-jerry-123/Physical-AI-AV-ES.Physical-AI-AV-DE
PhysicalAI-AV-SFT
Supervised fine-tuning (SFT) dataset for an autonomous-vehicle vision-language
waypoint-prediction model. Contains 324,105 samples from 150 000
driving scenes (18 seconds per scene, sampled at anchor times 2s..16s) recorded in the
United States.
Format
WebDataset — 10 uncompressed .tar shards,
each containing pairs of files per sample:
Entry
Description
{key}.jpg
Front-facing wide-angle camera frame (JPEG quality 95, 640 × 360 px)
{key}.json… See the full description on the dataset page: https://huggingface.co/datasets/tom-jerry-123/Physical-AI-AV-DE.av2_bev-lidarPhysical-AI-AV-IT
PhysicalAI-AV-SFT
Supervised fine-tuning (SFT) dataset for an autonomous-vehicle vision-language
waypoint-prediction model. Contains 29,991 samples from 150 000
driving scenes (18 seconds per scene, sampled at anchor times 2s..16s) recorded in the
United States.
Format
WebDataset — 3 uncompressed .tar shards,
each containing pairs of files per sample:
Entry
Description
{key}.jpg
Front-facing wide-angle camera frame (JPEG quality 95, 640 × 360 px)
{key}.json… See the full description on the dataset page: https://huggingface.co/datasets/tom-jerry-123/Physical-AI-AV-IT.finevideo_av_mcqaCustom_Dataset_AvatarForcing_HDTFAVSpeech_privateAVS
