datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
frakturline-testset
Fraktur/Other Text-Line — Test Set
A balanced, held-out evaluation set of 2 000 scanned text-line images (1 000 per class) for the binary task of distinguishing Fraktur (blackletter / Gothic script) from other script (primarily Antiqua / Latin / Roman).
Developed for the Impresso digital humanities project.
Dataset Details
Property
Value
Task
Binary image classification
Classes
fraktur, other
Images per class
1 000
Total images
2 000
Image format
WebP… See the full description on the dataset page: https://huggingface.co/datasets/impresso-project/frakturline-testset.testset
Dataset Card for TreeOfLife-10M Captions
This dataset consists of generated captions, Wikipedia-derived descriptions and format examples for the TreeOfLife-10M. These captions were generated using InternVL3-38B based on biological contexts that help the model generate more accurate captions. It was used to train BioCAP, a CLIP-based model.
Dataset Details
This dataset is comprised of captions for the images in TreeOfLife-10M that were generated using InternVL3 38B.… See the full description on the dataset page: https://huggingface.co/datasets/ZihengZ/testset.winml-test-set
WinML Test Set
Dataset Summary
WinML Test Set is an evaluation‑only collection for validating model accuracy and stability on Windows ML / DirectML / ONNX Runtime pipelines. It aggregates several permissively‑licensed sources and harmonizes schema for reproducible, regression‑grade testing across backends and versions. Not intended for training.
Intended Use
Accuracy and regression benchmarking of Windows ML / DirectML / ONNX Runtime pipelines.… See the full description on the dataset page: https://huggingface.co/datasets/Futuremark/winml-test-set.GenAI-RealEstate-TestSet
🏙️ GenAI Real Estate Test Set (Track B)
Dataset for the MenaML Winter School 2026 Challenge.
📊 Dataset Structure
This dataset contains 1,000 images split evenly between:
Authentic: Real estate photography from the Places365 dataset.
Manipulated: Synthetically generated deepfake artifacts (Inpainting, Diffusion Noise, GAN Grids).
🕵️ How to Use
This dataset is designed for testing forensic detection models.
