datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
bangla-10k
Bangla-10K: A Challenging, Metadata-Rich Corpus of Read and Conversational Bengali Speech from India and Bangladesh
Bangla-10K is a 10,816-hour Bengali speech corpus with
624,951 recordings from India and Bangladesh: a 10,070.8-hour core corpus
(567,323 recordings) and a separately collected 745.1-hour evaluation set
(57,628 recordings). It combines scripted single-speaker read speech with
natural multi-speaker conversations for Bengali automatic speech recognition
(ASR).
The… See the full description on the dataset page: https://huggingface.co/datasets/psdn-ai/bangla-10k.PSD_image_datset_samples
InfoBay.AI SVG Dataset Catalogue
Overview
The InfoBay.AI SVG Dataset Catalogue is a professionally curated collection of 873,754 scalable vector graphics (SVG) designed for Artificial Intelligence, Computer Vision, UI/UX Design, Search Systems, Frontend Development, Design Automation, and Multimodal Foundation Models.
The dataset contains high-quality vector assets spanning user interface components, branding materials, logos, technology icons, healthcare graphics… See the full description on the dataset page: https://huggingface.co/datasets/InfoBayAI/PSD_image_datset_samples.semi-pre-psdcrello_psd
