datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
waqfeya-library
Waqfeya Library
📖 Overview
Waqfeya is one of the primary online resources for Islamic books, similar to Shamela. It hosts more than 10,000 PDF books across over 80 categories.
In this dataset, we processed the original PDF files using Google Document AI APIs and extracted their contents into two additional formats: TXT and DOCX.
📊 Dataset Contents
The dataset includes 22,443 PDF files (spanning 8,978,634 pages) representing 10,150 Islamic books. Each book is… See the full description on the dataset page: https://huggingface.co/datasets/ieasybooks-org/waqfeya-library.british-library-book-images
British Library Book Images
1,080,814 images cut out of 49,455 digitised books (65,227 volumes, ~25 million pages) published
between c. 1510 and c. 1900, digitised by the British Library in partnership
with Microsoft and released by British Library Labs
on Flickr Commons as the "1 Million Images from Scanned Books" release. The books cover geography,
philosophy, history, poetry and literature, in several languages.
The four image types
British Library Labs… See the full description on the dataset page: https://huggingface.co/datasets/biglam/british-library-book-images.british-library-book-images
British Library Book Images
1,080,814 images cut out of 49,455 digitised books (65,227 volumes, ~25 million pages) published
between c. 1510 and c. 1900, digitised by the British Library in partnership
with Microsoft and released by British Library Labs
on Flickr Commons as the "1 Million Images from Scanned Books" release. The books cover geography,
philosophy, history, poetry and literature, in several languages.
The four image types
British Library Labs… See the full description on the dataset page: https://huggingface.co/datasets/Faizaniqbal/british-library-book-images.agent-tts-librarysynthetic-captchas-library
🌍 Synthetic Multilingual CAPTCHA Library
Repository: remiai3/synthetic-captchas-libraryA multilingual dataset of synthetic 4-character CAPTCHA images designed for OCR, multilingual vision models, and script recognition research.
This dataset spans 44 world writing systems and is especially useful for low-resource script OCR training.
📌 Dataset Summary
Each script includes 100,000 unique CAPTCHA images.The dataset is provided in two parallel formats:
CSV version… See the full description on the dataset page: https://huggingface.co/datasets/remiai3/synthetic-captchas-library.remote-sensing-sample-library-manager
Remote Sensing Sample Library Manager
This repository contains a local desktop tool and its associated sample data for building remote-sensing element datasets from class-organized imagery and vector labels.
Contents
software/
dataset_filter_desktop_app/ # Tkinter desktop tool
data/
huanghekou_202508/ # full generated Huanghekou element dataset
huanghekou_202508_selected/ # selected subset generated by the tool… See the full description on the dataset page: https://huggingface.co/datasets/cuibinge/remote-sensing-sample-library-manager.sd_regularization_libraryworkloom-reference-librarynail-it-libraryrooftop_segmentationThis dataset is a copy of Inria Rooftop Segmentation Dataset (1024x1024 PNG) (https://www.kaggle.com/datasets/dhruvpanchal1/inria-rooftop-segmentation-dataset-1024x1024-png).
This dataset is derived from the open-source Inria Aerial Image Labeling Dataset(https://project.inria.fr/aerialimagelabeling/). It has been processed and converted into a ready-to-use format for rooftop segmentation tasks.
Key Features:
Image Size: 1024x1024 pixels (PNG)
Binary Masks:
0 = Background
255 = Rooftop… See the full description on the dataset page: https://huggingface.co/datasets/mapie-library/rooftop_segmentation.british-library-book-images
British Library Book Images
1,080,814 images cut out of 49,455 digitised books (65,227 volumes, ~25 million pages) published
between c. 1510 and c. 1900, digitised by the British Library in partnership
with Microsoft and released by British Library Labs
on Flickr Commons as the "1 Million Images from Scanned Books" release. The books cover geography,
philosophy, history, poetry and literature, in several languages.
The four image types
British Library Labs… See the full description on the dataset page: https://huggingface.co/datasets/Baworsar1/british-library-book-images.motif_library_final
Motif Modules V9 (Black Line Version)
Each row contains a motif pair:
png: raster preview (512×512)
svg: vector paths of the same motif
Total: 100 pairsCreated: 2025-10-06Author: maryzhangLicense: CC-BY-NC-SA 4.0
Chinese Porcelain Motif Library - Final Collection (PNG & SVG)
Dataset Description
Dataset Summary
A curated collection of 100 high-quality Chinese porcelain motifs provided in both raster (PNG) and vector (SVG) formats. These motifs have… See the full description on the dataset page: https://huggingface.co/datasets/maryzhang/motif_library_final.nosed-based-workflow-libraryberlin_state_library_ocr_with_images
Dataset Card for "berlin_state_library_ocr_with_images"
More Information needed
ai-media-libraryimnet1k_librarymy-photo-librarymy-music-library2d-anime-libraryL-One-Material-Library-assetsBigBoiX_Pose_Library
