datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Language-Grounded_Sparse_Encoder_Training
Language-Grounded Sparse Encoder (LanSE) — Training Data
This repository hosts the AI-generated images and human annotation datasets accompanying the paper:
Human-like Content Analysis for Generative AI with Language-Grounded Sparse Encoders
Yiming Tang, Arash Lagzian, Srinivas Anumasa, Qiran Zou, Yingtao Zhu, Ye Zhang, Trang Nguyen, Yih-Chung Tham, Ehsan Adeli, Ching-Yu Cheng, Yilun Du, Dianbo Liu
National University of Singapore · Tsinghua University · Stanford University ·… See the full description on the dataset page: https://huggingface.co/datasets/DesmondYMTang2024/Language-Grounded_Sparse_Encoder_Training.American-Sign-Language-MNIST
Dataset Card for ASL-MNIST
This is a FiftyOne dataset with 34,627 samples of American Sign Language (ASL) alphabet images, converted from the original Kaggle Sign Language MNIST dataset into a format optimized for computer vision workflows.
Installation
If you haven't already, install FiftyOne:
pip install -U fiftyone
Usage
import fiftyone as fo
from fiftyone.utils.huggingface import load_from_hub
# Load the dataset
# Note: other available arguments… See the full description on the dataset page: https://huggingface.co/datasets/Voxel51/American-Sign-Language-MNIST.Language_TactileLanguage-Guided Representation Learningfor Robust Cross-Sensor Tactile Perception
Mashood M. Mohsan · Muhayy Ud Din · Binzhao Xu · Ahmad Abubakar · Irfan Hussain
Khalifa University Center for Autonomous Robotic Systems (KUCARS)
Khalifa University, UAE
Khalifa University ·
TouchRIPE ·
KUCARS ·
AERIS Lab
Project website ·
Video ·
Quick start ·
Dataset ·
Citation
Language descriptions guide a tactile image encoder to learn material representations… See the full description on the dataset page: https://huggingface.co/datasets/Mashood/Language_Tactile.olfaction-vision-language-dataset
Olfaction-Vision-Language Learning: A Multimodal Dataset
Olfaction • Vision • Language
An open-sourced dataset and dataset builder for prototyping and exploratory olfaction-vision-language tasks within the AI, robotics, and AR/VR domains.
Whether this dataset is used for better vision-scent navigation with drones, triangulating the source of an odor in an image, extracting aromas from a scene, or augmenting a VR experience with scent, we hope its release will catalyze… See the full description on the dataset page: https://huggingface.co/datasets/kordelfrance/olfaction-vision-language-dataset.American-Sign-Language-MNIST
Dataset Card for ASL-MNIST
This is a FiftyOne dataset with 34,627 samples of American Sign Language (ASL) alphabet images, converted from the original Kaggle Sign Language MNIST dataset into a format optimized for computer vision workflows.
Installation
If you haven't already, install FiftyOne:
pip install -U fiftyone
Usage
import fiftyone as fo
from fiftyone.utils.huggingface import load_from_hub
# Load the dataset
# Note: other available… See the full description on the dataset page: https://huggingface.co/datasets/saifmughal23/American-Sign-Language-MNIST.Marathi-Sign-Language
Marathi Sign Language Detection Dataset Card
Dataset Description
The Marathi Sign Language Dataset is a comprehensive collection of images designed to facilitate the development and training of machine learning models for recognizing Marathi sign language gestures. This dataset includes 43 distinct classes, each representing a unique character in the Marathi sign language alphabet. With approximately 1.2k images per class, the dataset totals over 51k images, all uniformly… See the full description on the dataset page: https://huggingface.co/datasets/VinayHajare/Marathi-Sign-Language.asl_sign_languages_alphabets_v03Pakistani-Sign-Language
Pakistan Sign Language (PSL) Gesture Dataset
A landmark-based gesture recognition dataset for Pakistan Sign Language (PSL), built to support real-time sign language translation for deaf and hard-of-hearing communities in Pakistan. This dataset covers PSL alphabets, words, and sentences, making it one of the few structured and publicly available PSL resources in existence.
How the Data Was Collected
Each gesture was recorded via webcam and processed using MediaPipe in… See the full description on the dataset page: https://huggingface.co/datasets/Bakhtyar12/Pakistani-Sign-Language.malaysian-sign-language-dataset-v1
Malaysian Sign Language Dataset (V1)
Dataset Description
This dataset contains 170,000+ samples of Malaysian Sign Language (BIM).
It has been processed using MediaPipe to extract skeleton keypoints (1662 features per frame) and is organized by class.
Total Classes: [Insert Number, e.g., 80]
Format: .npy (NumPy arrays) stored inside .zip files (one zip per class).
Features: Pose, Left Hand, Right Hand landmarks.
Frame Length: Normalized to 30 frames per sequence.… See the full description on the dataset page: https://huggingface.co/datasets/PishangShedappp/malaysian-sign-language-dataset-v1.olfaction-vision-language-dataset
Olfaction-Vision-Language Learning: A Multimodal Dataset
Olfaction • Vision • Language
An open-sourced dataset and dataset builder for prototyping and exploratory olfaction-vision-language tasks within the AI, robotics, and AR/VR domains.
Whether this dataset is used for better vision-scent navigation with drones, triangulating the source of an odor in an image, extracting aromas from a scene, or augmenting a VR experience with scent, we hope its release will catalyze… See the full description on the dataset page: https://huggingface.co/datasets/Strangefiction/olfaction-vision-language-dataset.asl_sign_languages_alphabets_v02Sign_Language_for_Kurdish_LettersAmerican-Sign-Language-MNIST
Dataset Card for ASL-MNIST
This is a FiftyOne dataset with 34,627 samples of American Sign Language (ASL) alphabet images, converted from the original Kaggle Sign Language MNIST dataset into a format optimized for computer vision workflows.
Installation
If you haven't already, install FiftyOne:
pip install -U fiftyone
Usage
import fiftyone as fo
from fiftyone.utils.huggingface import load_from_hub
# Load the dataset
# Note: other available arguments… See the full description on the dataset page: https://huggingface.co/datasets/monishagp/American-Sign-Language-MNIST.American-Sign-Language-MNIST
Dataset Card for ASL-MNIST
This is a FiftyOne dataset with 34,627 samples of American Sign Language (ASL) alphabet images, converted from the original Kaggle Sign Language MNIST dataset into a format optimized for computer vision workflows.
Installation
If you haven't already, install FiftyOne:
pip install -U fiftyone
Usage
import fiftyone as fo
from fiftyone.utils.huggingface import load_from_hub
# Load the dataset
# Note: other available arguments… See the full description on the dataset page: https://huggingface.co/datasets/rrrajjj/American-Sign-Language-MNIST.Marathi-Sign-Language
Marathi Sign Language Detection Dataset Card
Dataset Description
The Marathi Sign Language Dataset is a comprehensive collection of images designed to facilitate the development and training of machine learning models for recognizing Marathi sign language gestures. This dataset includes 43 distinct classes, each representing a unique character in the Marathi sign language alphabet. With approximately 1.2k images per class, the dataset totals over 51k images, all… See the full description on the dataset page: https://huggingface.co/datasets/OMLONKAR/Marathi-Sign-Language.Hindi-Sign-Language-Datasetasl_sign_languages_alphabets_v03
