datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
google-landmark-v2-chinese-filtered
Google Landmark V2 Chinese Filtered Dataset
This dataset contains landmark images and metadata for training landmark retrieval models, with Chinese translations of landmark names to facilitate Chinese multimodal retrieval tasks.
Dataset Source
This dataset is based on the Google Landmarks V2 dataset from Kaggle. The original data has been filtered and processed to create a high-quality training dataset for landmark retrieval.
Key Features
Filtered… See the full description on the dataset page: https://huggingface.co/datasets/86Cao/google-landmark-v2-chinese-filtered.mb-landmark_cls
mb-landmark_cls
A Mars image classification dataset for planetary science research.
Dataset Metadata
License: CC-BY-4.0 (Creative Commons Attribution 4.0 International)
Version: 1.0
Date Published: 2025-05-14
Cite As: TBD
Classes
This dataset contains the following classes:
0: oth
1: cra
2: ddu
3: sst
4: bdu
5: ime
6: sch
7: spi
Statistics
train: 6997 images
test: 1793 images
val: 2025 images
few_shot_train_2_shot: 16 images… See the full description on the dataset page: https://huggingface.co/datasets/Mirali33/mb-landmark_cls.vox2-face-landmark
Dataset Card for "voxceleb2-features"
More Information needed
vrh3-part-placing-simulation-landmarkgoogle-landmarkshow2sign-asl-landmarks
How2Sign ASL Landmarks
Sentence-level MediaPipe landmark cache for the How2Sign ASL translation dataset. This dataset stores preprocessed landmark arrays only, not MP4 clips.
The rows come from the local How2Sign realigned CSV files and front-view raw videos. Each successful row is one sentence segment sampled at a 25 FPS cap and stored in compressed NPZ shards.
Why This Dataset Exists
The public How2Sign data has known alignment/completeness issues: some text rows do not… See the full description on the dataset page: https://huggingface.co/datasets/martinctl/how2sign-asl-landmarks.lerobot-put_the_cream_cheese_in_the_nearest_basket_and_place_the_empty_basket_in_center-landmark5_synt_flux_landmark_missing_imgs_validatedlerobot-LIBERO_MEM-T09_10-landmarkgoogle_landmarks_places
Google Landmarks places
Google Landmarks is a great dataset, but it lacks geospatial information about the places. This dataset fills
this gap by providing latitude and longitude for each landmark. The dataset also contains the name of the landmark from OpenStreetMap
and information about the country, the province/state, and the city/village where the landmark is located. This information was collected from OSM
via Nominatim.
vtl-speech-landmarks
VTL Speech Landmarks Dataset
Articulatory speech synthesis dataset with acoustic landmarks, generated using VocalTractLab (VTL).
Dataset Description
This dataset contains synthesized speech for 117,497 English words from the CMU Pronouncing Dictionary, generated with two speakers (male and female). Each word includes:
Audio: 48kHz WAV files
Landmarks: Acoustic-phonetic event markers (JSON)
Articulatory data: Full vocal tract trajectories from VTL (JSON)
Speakers… See the full description on the dataset page: https://huggingface.co/datasets/mcamara/vtl-speech-landmarks.nsaka-complete-landmarkslandmarks
Dataset Card for "landmarks"
More Information needed
Level-5-Cultures-Landmarks-1google-landmarks-v2-miniOptimized_Video_Facial_Landmarks
Dataset Card for 478-Point Normalized 3D Facial Landmark Dataset
Dataset Description
This dataset provides pre-extracted, normalized 3D facial landmark features derived from the Video Emotion dataset. It is optimized for efficient training of emotion recognition and facial analysis models, bypassing the need to process large raw video files.
License: The extracted feature data in this Parquet file is licensed under Apache 2.0. Note that the original source video files may… See the full description on the dataset page: https://huggingface.co/datasets/PSewmuthu/Optimized_Video_Facial_Landmarks.gym-exercise-landmarksgoogle-landmark-geo
Dataset Card for Geo Coordinate Augmented Google-Landmarks
Geo coordinates were added as data to a tar file's worth of images from the Google Landmark V2. Not all of the
images could be geo-tagged due to lack of coordinates on the image's wikimedia page.
Dataset Details
Dataset Description
Geo coordinates were added as data to a tar file's worth of images from the Google Landmark V2. There were many more images that could have
been downloaded but this dataset… See the full description on the dataset page: https://huggingface.co/datasets/Qdrant/google-landmark-geo.juno-landmark-rfttunisian-landmarks-captions
🇹🇳 Tunisian Landmarks Image-Caption Dataset
An image-caption dataset documenting Tunisia's rich geographical, historical, and architectural heritage. The dataset pair high-resolution photographs of iconic Tunisian sites with detailed, human-curated descriptions highlighting their cultural significance, geographical features, and architectural styles.
Dataset Details
Dataset Description
This dataset is built to support vision-language tasks such as… See the full description on the dataset page: https://huggingface.co/datasets/firastlili/tunisian-landmarks-captions.google_landmarks_photos
Dataset Card for "google_landmarks_photos"
More Information needed
slovo-full-landmarks
Slovo landmarks base
Base landmarks dataset built from full Slovo videos with MediaPipe Tasks HolisticLandmarker.
lerobot-RMBench-Observe_n_Pickup-landmarkgoogle_landmark_v2_qahow2sign-clips-landmarksThis dataset is from:
@inproceedings{Duarte_CVPR2021,
title={{How2Sign: A Large-scale Multimodal Dataset for Continuous American Sign Language}},
author={Duarte, Amanda and Palaskar, Shruti and Ventura, Lucas and Ghadiyaram, Deepti and DeHaan, Kenneth and
Metze, Florian and Torres, Jordi and Giro-i-Nieto, Xavier},
booktitle={Conference on Computer Vision and Pattern Recognition (CVPR)},
year={2021}
}
Also see the offical how2sign dataset website:… See the full description on the dataset page: https://huggingface.co/datasets/tim1047/how2sign-clips-landmarks.Emotion_Video_Facial_Landmarks
Dataset Card for 478-Point Normalized 3D Facial Landmark Dataset
Dataset Description
This dataset provides pre-extracted, normalized 3D facial landmark features derived from the Video Emotion dataset. It is optimized for efficient training of emotion recognition and facial analysis models, bypassing the need to process large raw video files.
License: The extracted feature data in this CSV file is licensed under Apache 2.0. Note that the original source video files may… See the full description on the dataset page: https://huggingface.co/datasets/PSewmuthu/Emotion_Video_Facial_Landmarks.5_synt_landmark_singlelerobot-RMBench-CoverBlocks-landmarkCASL-W60-Landmarks
CASL-W60-64-frames Landmarks
license: mit
task_categories:
- video-classification
language:
- en
tags:
- sign-language
- mediapipe
- skeleton
- Central-African-Sign-Language
pretty_name: CASL-W60 Landmarks
size_categories:
- 1K<n<10K
Dataset Card for CASL-W60 Landmarks
This dataset contains preprocessed holistic landmarks for Central African Sign Language (CASL). It is specifically designed for training sequence-based models like Transformers or LSTMs for Sign… See the full description on the dataset page: https://huggingface.co/datasets/luciayen/CASL-W60-Landmarks.5_synt_flux_landmark_selected_single_v2_validated_1011
