datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Uddessho-Bangla-Multimodal-Intent-Classification
📊 Uddessho Dataset — Multimodal Author Intent Classification
Uddessho (meaning "Intent" in English) is a multimodal dataset created for author intent classification in the low-resource Bangla language.It contains 3,048 social media posts (text + images) labeled into six distinct intent types.
🏷️ Intent Categories & Label Mapping
Label ID
Class Name
0
Advocative
1
Controversial
2
Exhibitionist
3
Expressive
4
Informative
5
Promotive
📂… See the full description on the dataset page: https://huggingface.co/datasets/Mukaffi28/Uddessho-Bangla-Multimodal-Intent-Classification.EmoSet-118Ksuno
Dataset Card for Suno.ai Music Generation
Dataset Summary
This dataset contains metadata for 659,788 songs generated by artificial intelligence on the suno.com platform, a service that generates music using artificial intelligence. The songs were discovered by search queries with words from the dwyl/english-words word list.
Languages
The dataset is multilingual with English as the primary language:
English (en): Primary language for metadata and most lyrics… See the full description on the dataset page: https://huggingface.co/datasets/Udtseal/suno.body-measurements-image-dataset
Body Measurements Image Dataset - 13,000 Images
Dataset consists of 13,000+ standardized photos of 1,000+ people, offering a robust resource for body measurements estimation, human body analysis, and personalized sizing recommendations in e-commerce. Each subject is captured in front and side poses with paired 17+ anthropometric measurements, enabling precise shape estimation, weight prediction, and body characteristics detection.— Get the data
Dataset characteristics:… See the full description on the dataset page: https://huggingface.co/datasets/ud-biometrics/body-measurements-image-dataset.Annecy_v.02open-palm-hand-images
Palm Dataset - 500,000 Images
Dataset comprises 500,000 high-quality images featuring diverse human hands, specifically designed for hand detection, palm recognition, and gesture analysis. It provides diverse training data with metadata on age, gender, and ethnicity for accurate computer vision model training.— Get the data
Dataset characteristics:
Characteristic
Data
Description
Open palm images designed for training and evaluating hand-based… See the full description on the dataset page: https://huggingface.co/datasets/ud-biometrics/open-palm-hand-images.synthetic-printed-nz-passports
Synthetic Passports Dataset - 5 000 passport photos
Dataset features 5,000 AI-generated New Zealand passport images captured under varied angles, lighting, and backgrounds. Designed for OCR, computer vision, and identity verification research, this NZ passport dataset supports training models in PII extraction, document recognition, and synthetic passport analysis with rich metadata annotations. - Get the data
Dataset characteristics:
Characteristic
Data… See the full description on the dataset page: https://huggingface.co/datasets/ud-biometrics/synthetic-printed-nz-passports.spain-license-plate-dataset
License Plate Recognition - 116,237 Image
This dataset provides 116,237 vehicle images captured in Spain, serving as a robust foundation for OCR systems, license plate identification, and vehicle registration data retrieval. Every image includes a companion CSV file containing the accurate plate number and country code, making it perfect for building and validating automated text recognition solutions. - Get the data
Dataset characteristics:
Characteristic
Data… See the full description on the dataset page: https://huggingface.co/datasets/ud-smart-city/spain-license-plate-dataset.Garbage-Classification-with-12-classesAnti-Spoofing-Real-Videos
Face Anti Spoofing Dataset - 98 000+ files
Dataset features 98,000+ files of real photos and videos of people from 170+ countries, representing 70,000+ unique individuals. By leveraging this dataset, developers can enhance spoofing detection techniques, improve recognition systems, and deploy anti-spoofing algorithms capable of preventing fraud in deep learning-based solutions.- Get the data
Dataset characteristics:
Characteristic
Data
Description
Live… See the full description on the dataset page: https://huggingface.co/datasets/ud-biometrics/Anti-Spoofing-Real-Videos.stylesage-imagesUDD
UDD — Universal Document Dataset
UDD scatters many public document / OCR benchmarks into one standardized, sharded dataset,
unifying every task — document VQA, key-information extraction (KIE), localization / spotting,
full-text recognition, table-structure, chart/figure reasoning, and document classification —
under a single schema. Instead of N incompatible formats you load one dataset and filter by
task / source.
Built with the open pipeline in SangbumChoi/OCR… See the full description on the dataset page: https://huggingface.co/datasets/danelcsb/UDD.rocksRocks dataset with 7 classes: [Coal, Limestone, Marble, Sandstone, Quartzite, Basalt, Granite]
synthetic-usa-driver-license
Synthetic License Dataset - 5 000 passport photos
Dataset features 5,000 high-quality, AI-generated images of U.S. driver licenses from multiple states, including California, Texas, and New York. Designed for OCR, data extraction, and identity verification model training, this synthetic dataset ensures realistic detail, balanced demographics, and secure handling of personally identifiable information. - Get the data
Dataset characteristics:
Characteristic
Data… See the full description on the dataset page: https://huggingface.co/datasets/ud-biometrics/synthetic-usa-driver-license.passport-dataset
Synthetic Passports Dataset - 100,000 passport photos
Dataset comprises 100,000+ synthetically generated images of passports from 100+ countries, designed for identity verification, fraud detection, and document analysis. It includes document layouts and passport photos to train neural networks in identifying fake documents and enhancing verification systems. - Get the data
Dataset characteristics:
Characteristic
Data
Description
Generated passports for… See the full description on the dataset page: https://huggingface.co/datasets/ud-biometrics/passport-dataset.garbage-classification
Trash Classification - 5,000+ photos
Dataset comprises 5,000+ photos of garbage cans featuring various capacities, types, and waste materials, designed for advancing garbage classification and waste management systems. By leveraging this dataset, researchers and developers can enhance classification systems, automate garbage collection processes, and improve strategies for reducing environmental pollution. - Get the data
Dataset characteristics:
Characteristic… See the full description on the dataset page: https://huggingface.co/datasets/ud-smart-city/garbage-classification.water-meter-image
Water Meter Pics - 5,000+ photos
Dataset comprises 5,000+ photos of water meters, including high-quality images, segmentation masks, and OCR labels for meter readings. Each entry provides detailed information such as the meter reading value, bounding box coordinates, and segmentation data, making it ideal for training models in utility management, automatic meter reading (AMR), and water usage analysis.- Get the data
Dataset characteristics:
Characteristic
Data… See the full description on the dataset page: https://huggingface.co/datasets/ud-smart-city/water-meter-image.license-plate-dataset
License Plate Recognition - 2.6 Million images
Dataset comprises 2.6 million images of vehicle license plates across 86 countries, providing a comprehensive resource for OCR, traffic analysis, and autonomous vehicle systems.It focuses on plate recognitions and related detection systems, providing detailed information on plate numbers, country, bbox labeling and other data as well as corresponding masks for recognition tasks.- Get the data
Dataset characteristics:… See the full description on the dataset page: https://huggingface.co/datasets/ud-smart-city/license-plate-dataset.weapon-detection
Weapon Detection Dataset - 10,000 images
The dataset contains 10,000 annotated images of staged individuals with visible weapons, collected from public CCTV footage and internet sources for training weapon detection models.- Get the data
Dataset characteristics:
Characteristic
Data
Description
Images of staged individuals with weapons sourced from the internet and CCTV footage
Data types
Image
Tasks
Public Safety, Computer Vision
Total number of… See the full description on the dataset page: https://huggingface.co/datasets/ud-smart-city/weapon-detection.synthetic-printed-japanese-passports
Synthetic Passports Dataset - 5 000 passport photos
Dataset contains 5,000 AI-generated, high-resolution passport images with diverse lighting, angles, and backgrounds. It supports document analysis, OCR, and biometric data research, offering realistic Japanese passport images for training and evaluating identity recognition and personal data extraction systems. - Get the data
Dataset characteristics:
Characteristic
Data
Description
Printed synthetic… See the full description on the dataset page: https://huggingface.co/datasets/ud-biometrics/synthetic-printed-japanese-passports.Crowd-Countin-Dataset
Crowd Dataset - 647 Photos
Dataset comprises 647 photos of dense crowds, containing between 1,000 to 13,000 people per image. Each image includes detailed keypoint annotations for every individual, enabling advanced data analysis and deep learning applications in crowd density estimation, object detection, and counting algorithms. - Get the data
Dataset characteristics:
Characteristic
Data
Description
Crowd photos with labeling for determining crowd density… See the full description on the dataset page: https://huggingface.co/datasets/ud-smart-city/Crowd-Countin-Dataset.united-states-license-plate-dataset
License Plate Recognition - 89 986 Image
Dataset contains 89,986 photos of license plates (number plates), designed for OCR tasks, vehicle registration analysis, and traffic management systems. This large-scale collection provides real-world data for training and evaluating models in license plate recognition (LPR), character recognition, and autonomous driving applications. - Get the data
Dataset characteristics:
Characteristic
Data
Description
License plate… See the full description on the dataset page: https://huggingface.co/datasets/ud-smart-city/united-states-license-plate-dataset.annecy-testmale-hair-loss-dataset
Male Hair Loss Dataset - 2 400+ images
Dataset comprises medical images of scalps from five angles, labeled with seven classifications based on the Norwood-Hamilton scale, aiding in diagnosing hair losses and scalp conditions. Utilizing deep learning techniques, machine learning algorithms can analyze hair density, follicles, and hair growth patterns to improve accurate diagnosis of alopecia areata and other hair disorders. — Get the data
Dataset characteristics:… See the full description on the dataset page: https://huggingface.co/datasets/ud-medical/male-hair-loss-dataset.UDAPose-synthetic-dataear-detection-dataset
Ear Detection - 14,000+ Images
The dataset comprises 14,000+ ear images from 2,000 unique individuals, paired with reference face photos and demographic labels. Designed for ear recognition and biometric identification, it helps research in human ear detection, recognition accuracy, and biometric systems.— Get the data
Dataset characteristics:
Characteristic
Data
Description
Images of an ear with a face photo
Data types
Image
Tasks
Ear‑biometric R&D… See the full description on the dataset page: https://huggingface.co/datasets/ud-biometrics/ear-detection-dataset.synthetic-printed-uk-passports
Synthetic Passports Dataset - 5 000 passport photos
Dataset features 5,000 AI-generated UK passport images captured under varied angles, lighting, and backgrounds. Designed for OCR, computer vision, and identity verification research, this passport dataset supports training models in PII extraction, document recognition, and synthetic passport analysis with rich metadata annotations. - Get the data
Dataset characteristics:
Characteristic
Data
Description
Printed… See the full description on the dataset page: https://huggingface.co/datasets/ud-biometrics/synthetic-printed-uk-passports.men-hair-loss-dataset
Hair Loss Segmentation Dataset - 3 100 images
The dataset comprises 3,100 images from 775 individuals, featuring male alopecia cases captured from two angles (front + top views) with corresponding segmentation masks. Designed for learning algorithms for detecting hair disorders, evaluating hair restoration techniques, and training models for early diagnosis of alopecia. — Get the data
Dataset characteristics:
Characteristic
Data
Description
Photos of men with… See the full description on the dataset page: https://huggingface.co/datasets/ud-medical/men-hair-loss-dataset.synthetic-printed-german-passports
Synthetic Passports Dataset - 5 000 passport photos
Dataset features 5,000 AI-generated German passport images captured under varied angles, lighting, and backgrounds. Designed for OCR, computer vision, and identity verification research, this passport dataset supports training models in PII extraction, document recognition, and synthetic passport analysis with rich metadata annotations. - Get the data
Dataset characteristics:
Characteristic
Data
Description… See the full description on the dataset page: https://huggingface.co/datasets/ud-biometrics/synthetic-printed-german-passports.poland-license-plate-dataset
License Plate Recognition - 196 664 Image
Dataset contains 196 664 photos of license plates, providing a comprehensive resource for OCR (Optical Character Recognition), plate recognition, and vehicle registration analysis. Each image file is paired with a CSV file containing license plate text for use in OCR, plate recognition, and characters recognition tasks. - Get the data
Dataset characteristics:
Characteristic
Data
Description
License plate images with… See the full description on the dataset page: https://huggingface.co/datasets/ud-smart-city/poland-license-plate-dataset.
