CoolFace
20 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01tom-jerry-123 /Physical-AI-AV-US PhysicalAI-AV-SFT Supervised fine-tuning (SFT) dataset for an autonomous-vehicle vision-language waypoint-prediction model. Contains 2,789,773 samples from 150 000 driving scenes (18 seconds per scene, sampled at 1 Hz) recorded in the United States. Format WebDataset — 100 uncompressed .tar shards, each containing pairs of files per sample: Entry Description {key}.png Front-facing wide-angle camera frame (640 × 360 px) {key}.json Metadata (see schema below)… See the full description on the dataset page: https://huggingface.co/datasets/tom-jerry-123/Physical-AI-AV-US.imagerobotics1M<n<10M0 likes913 downloads6mo agoHugging Face02seungheondoh /mmtrailer-pe-av-unimodal-local-embeddingstext10K<n<100K0 likes794 downloads6mo agoHugging Face03nguyenvulebinh /AVYTtext1M<n<10M1 likes515 downloads1y agoHugging Face04seungheondoh /yt-pe-av-unimodal-local-embeddingstext10K<n<100K0 likes510 downloads5mo agoHugging Face05seungheondoh /movielens-pe-av-local-embeddingstext10K<n<100K0 likes408 downloads6mo agoHugging Face06seungheondoh /yt-pe-av-local-embeddingstext10K<n<100K0 likes262 downloads6mo agoHugging Face07seungheondoh /movielens-pe-av-unimodal-local-embeddingstext10K<n<100K0 likes205 downloads5mo agoHugging Face08nguyenvulebinh /AVCocktail AVSRCocktail: Audio-Visual Speech Recognition for Cocktail Party Scenarios Official implementation of "Cocktail-Party Audio-Visual Speech Recognition" (Interspeech 2025). A robust audio-visual speech recognition system designed for multi-speaker environments and noisy cocktail party scenarios. The model combines lip reading and audio processing to achieve superior performance in challenging acoustic conditions with background noise and speaker interference. Getting… See the full description on the dataset page: https://huggingface.co/datasets/nguyenvulebinh/AVCocktail.text1K<n<10K1 likes156 downloads1y agoHugging Face09seungheondoh /mmtrailer-pe-av-local-embeddingstext10K<n<100K0 likes118 downloads6mo agoHugging Face10harryhsing /AVQA-R1-6KThis repository contains data presented in EchoInk-R1: Exploring Audio-Visual Reasoning in Multimodal LLMs via Reinforcement Learning. For training and inference, please refer to the Code: https://github.com/HarryHsing/EchoInk Data Format in AVQA-R1-6K: { "problem_id": 0, "problem": "What is the source of the sound in the video?", "data_type": "image_audio", "problem_type": "multiple choice", "options": [ "A. motorcycle", "B. automobile"… See the full description on the dataset page: https://huggingface.co/datasets/harryhsing/AVQA-R1-6K.audio1K<n<10K3 likes79 downloads1y agoHugging Face11tom-jerry-123 /Physical-AI-AV-FR PhysicalAI-AV-SFT Supervised fine-tuning (SFT) dataset for an autonomous-vehicle vision-language waypoint-prediction model. Contains 29,909 samples from 150 000 driving scenes (18 seconds per scene, sampled at anchor times 2s..16s) recorded in the United States. Format WebDataset — 3 uncompressed .tar shards, each containing pairs of files per sample: Entry Description {key}.jpg Front-facing wide-angle camera frame (JPEG quality 95, 640 × 360 px) {key}.json… See the full description on the dataset page: https://huggingface.co/datasets/tom-jerry-123/Physical-AI-AV-FR.imagerobotics10K<n<100K0 likes16 downloads5mo agoHugging Face12nvidia /Cosmos-Transfer1-7B-Sample-AV-Data-Examplegated Cosmos-Transfer1-7B-Sample-AV-Data-Example Cosmos | Code | Paper | Paper Website Dataset Description: This dataset contains 10 sample data points intended to help users better utilize our Cosmos-Transfer1-7B-Sample-AV model. It includes HD Map annotations and LiDAR data, with no personally identifiable information such as faces or license plates. This dataset is intended for research and development only. Dataset Owner(s): NVIDIA Dataset Creation… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Cosmos-Transfer1-7B-Sample-AV-Data-Example.textn<1K9 likes14 downloads2y agoHugging Face13tom-jerry-123 /Physical-AI-AV-ES PhysicalAI-AV-SFT Supervised fine-tuning (SFT) dataset for an autonomous-vehicle vision-language waypoint-prediction model. Contains 29,674 samples from 150 000 driving scenes (18 seconds per scene, sampled at anchor times 2s..16s) recorded in the United States. Format WebDataset — 3 uncompressed .tar shards, each containing pairs of files per sample: Entry Description {key}.jpg Front-facing wide-angle camera frame (JPEG quality 95, 640 × 360 px) {key}.json… See the full description on the dataset page: https://huggingface.co/datasets/tom-jerry-123/Physical-AI-AV-ES.imagerobotics10K<n<100K0 likes10 downloads5mo agoHugging Face14tom-jerry-123 /Physical-AI-AV-DE PhysicalAI-AV-SFT Supervised fine-tuning (SFT) dataset for an autonomous-vehicle vision-language waypoint-prediction model. Contains 324,105 samples from 150 000 driving scenes (18 seconds per scene, sampled at anchor times 2s..16s) recorded in the United States. Format WebDataset — 10 uncompressed .tar shards, each containing pairs of files per sample: Entry Description {key}.jpg Front-facing wide-angle camera frame (JPEG quality 95, 640 × 360 px) {key}.json… See the full description on the dataset page: https://huggingface.co/datasets/tom-jerry-123/Physical-AI-AV-DE.imagerobotics100K<n<1M0 likes9 downloads6mo agoHugging Face15czm369 /av2_bev-lidartext100K<n<1M0 likes6 downloads1y agoHugging Face16tom-jerry-123 /Physical-AI-AV-IT PhysicalAI-AV-SFT Supervised fine-tuning (SFT) dataset for an autonomous-vehicle vision-language waypoint-prediction model. Contains 29,991 samples from 150 000 driving scenes (18 seconds per scene, sampled at anchor times 2s..16s) recorded in the United States. Format WebDataset — 3 uncompressed .tar shards, each containing pairs of files per sample: Entry Description {key}.jpg Front-facing wide-angle camera frame (JPEG quality 95, 640 × 360 px) {key}.json… See the full description on the dataset page: https://huggingface.co/datasets/tom-jerry-123/Physical-AI-AV-IT.imagerobotics10K<n<100K0 likes5 downloads5mo agoHugging Face17Rakancorle11 /finevideo_av_mcqatext1K<n<10K0 likes5 downloads5mo agoHugging Face18Darknsu /Custom_Dataset_AvatarForcing_HDTFtext10K<n<100K0 likes3 downloads4mo agoHugging Face19jjuik2014 /AVSpeech_privategatedtext10K<n<100K0 likes2 downloads9mo agoHugging Face20SCX163 /AVSgatedaudio1M<n<10M4 likes1 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.