datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
icdar_disco
ICDAR_mini Dataset
A balanced mini subset of the ICDAR (International Conference on Document Analysis and Recognition) dataset with 50 samples per language. Includes actual document images and ground truth OCR text.
Dataset Details
Total Samples: 500
Total Images: 500
Languages: 10
Arabic (50 samples)
Bangla (50 samples)
Chinese (50 samples)
Hindi (50 samples)
Japanese (50 samples)
Korean (50 samples)
Latin (50 samples)
Mixed (50 samples)
None (50 samples)… See the full description on the dataset page: https://huggingface.co/datasets/kenza-ily/icdar_disco.Panda-Discordant-Pathology-Errorsvisrbench_disco
VisR-Bench Mini
A curated subset of VisR-Bench containing 340 high-quality documents with complete image data and 5 sampled questions per document.
Dataset Summary
Total Documents: 340
Total QA Pairs: 1,680
Total Images: 6,803 PNG files (5.2GB)
Sampling: 5 questions randomly sampled per document (seed=42)
Image Coverage: All documents include complete multi-page image data
Content Types
Type
Documents
QA Pairs
Total Pages
Description… See the full description on the dataset page: https://huggingface.co/datasets/kenza-ily/visrbench_disco.MovieTection
Dataset Description 🎬
The MovieTection dataset is a benchmark designed for detecting pretraining data in Large Vision-Language Models (VLMs). It serves as a resource for analyzing model exposure to Copyrighted Visual Content ©️.
Paper: DIS-CO: Discovering Copyrighted Content in VLMs Training Data
Direct Use 🖥️
The dataset is designed for image/caption-based question-answering, where models predict the movie title given a frame or its corresponding textual… See the full description on the dataset page: https://huggingface.co/datasets/DIS-CO/MovieTection.parkinsons-evidence-to-discovery-prioritisation
Parkinson's Disease Evidence-to-Discovery Prioritisation Dataset
This Hugging Face dataset package contains processed research assets from an AI-assisted evidence synthesis and computational validation project on Parkinson's disease (PD) prevention and disease-modifying therapeutic strategy prioritisation.
Dataset Summary
The dataset integrates:
evidence-priority scores for PD prevention and disease-modification candidates;
pathway-to-intervention framework;
individual… See the full description on the dataset page: https://huggingface.co/datasets/hssling/parkinsons-evidence-to-discovery-prioritisation.iam_disco
IAM_mini Dataset - Fair Handwriting Recognition
A fairness-enhanced mini subset of the IAM (Institute of Applied Mathematics) Handwriting Database with 500 randomly selected samples. Each image is separated into printed and handwritten components for fair VLM evaluation.
Fairness Enhancement
The original IAM dataset contains images with both printed reference text and handwritten content on the same page. This allows Vision Language Models (VLMs) to "cheat" by… See the full description on the dataset page: https://huggingface.co/datasets/kenza-ily/iam_disco.docvqa_disco
DocVQA_mini Dataset
A mini subset of the DocVQA dataset with 500 randomly selected question-answer pairs for document visual question answering evaluation.
Dataset Details
Total Samples: 500 QA pairs
Source: DocVQA validation set
Task: Document Visual Question Answering
Image Format: PNG (extracted from parquet-embedded images)
Features
Each sample contains:
image: Document image
question: Question about the document
answers: List of valid… See the full description on the dataset page: https://huggingface.co/datasets/kenza-ily/docvqa_disco.disco-elysium-gameplay-data
极乐迪斯科
This public dataset repository contains local gameplay data uploaded from F:\极乐迪斯科.
Contents
Files: 143
Total local size: 128.61 GB
Generated: 2026-06-05 18:26:27 UTC
File Types
.jsonl: 44
.json: 33
.png: 33
.txt: 11
.mkv: 11
.parquet: 11
Notes
This repository may contain gameplay video, images, Parquet files, JSON/JSONL metadata, and keyboard/mouse event logs.
The license is marked as other; review game footage, audio, and… See the full description on the dataset page: https://huggingface.co/datasets/xiaoluo11/disco-elysium-gameplay-data.infographicvqa_disco
InfographicVQA_mini Dataset
A mini subset of the InfographicVQA dataset with 500 randomly selected question-answer pairs for infographic visual question answering evaluation.
Dataset Details
Total Samples: 500 QA pairs
Source: InfographicVQA validation set
Task: Infographic Visual Question Answering
Image Format: PNG (extracted from parquet-embedded images)
Features: Includes pre-extracted OCR text from AWS Textract
Features
Each sample contains:… See the full description on the dataset page: https://huggingface.co/datasets/kenza-ily/infographicvqa_disco.disco-diffusionThis dataset contains just under half of the training data used to train Paint Journey. All 768x768 images were generated using one of Disco Diffusion v3.1, v4.1, and v5.x,
but later upscaled then downscaled twice (super resolution) using R-ESRGAN General WDN 4x V3 just before training.
zhangxian_1928_portrait_discovery_storm_society_modern_art_pioneer🖼️ Zhang Xian (1898–1936): The Lost 1928 Portrait and His Legacy in China's Modern Art Movement
Dataset title: zhangxian_1928_portrait_discovery_storm_society_modern_art_pioneer
Compiled by: HaruthaiAI (Hugging Face)
📄 Metadata JSON
Structured metadata file: zhangxian_portrait_metadata.jsonIncludes multi-layered context: discovery, attribution, movement, style, and linked documents.
📘 Overview (English)
This dataset documents a significant rediscovery: a 1928 portrait of Zhang… See the full description on the dataset page: https://huggingface.co/datasets/HaruthaiAi/zhangxian_1928_portrait_discovery_storm_society_modern_art_pioneer.DISCOVR
DIgitial Social Context Objects in VR (DISCOVR): A Social Virtual Reality Object Detection Dataset
Dataset Description
DISCOVR is an object detection dataset for identifying user interface elements and interactive objects in virtual reality (VR) and social VR environments. The dataset contains 17,691 annotated images across 30 object classes commonly found in 17 top social VR applications & VR demos.
This dataset is designed to support research in VR accessibility… See the full description on the dataset page: https://huggingface.co/datasets/UWMadAbility/DISCOVR.chartqapro_disco
ChartQAPro Mini Dataset
A stratified 494-sample subset of the ChartQAPro dataset for chart question answering evaluation. This mini version maintains the diversity of the full dataset while being suitable for quick benchmarking and testing.
Dataset Description
ChartQAPro_mini contains question-answer pairs from diverse chart types with balanced representation across:
Question Types: Factoid (55.9%), Conversational (16%), Fact Checking (12.8%), Multi Choice… See the full description on the dataset page: https://huggingface.co/datasets/kenza-ily/chartqapro_disco.vectors2vibes-discogs-metadata
Vectors2Vibes Discogs Metadata
Metadata for 24.6k tracks derived from Discogs Data and MTG Discogs-VI-YT. No audio files.
Note that earliest release year data is derived from MusicBrainz, as Discogs release year data is sparse and often unreliable.
Quick Facts:This dataset contains 24,689 tracks (release dates spanning from 1890-2026).
The top 5 represented decades are: 1960s (19.39%), 1970s (15.09%), 1980s (14.81%), 1950s (14.96%), and 1990s (14.27%).
The top 5 represented genres… See the full description on the dataset page: https://huggingface.co/datasets/vectors2vibes/vectors2vibes-discogs-metadata.MovieTection_Mini
Dataset Description 🎬
The MovieTection_Mini dataset is a benchmark designed for detecting pretraining data in Large Vision-Language Models (VLMs). It serves as a resource for analyzing model exposure to Copyrighted Visual Content ©️.
This dataset is a compact subset of the full MovieTection dataset, containing only 4 movies instead of 100. It is designed for users who want to experiment with the benchmark without the need to download the entire dataset, making it a more… See the full description on the dataset page: https://huggingface.co/datasets/DIS-CO/MovieTection_Mini.gpt4v-LAION-discord
Dataset Card for "gpt4v-LAION-discord"
More Information needed
DISCO_benchmark_data
DISCO (DIffusion for Sequence-structure CO-design) is a multimodal generative model that simultaneously co-designs protein sequences and 3D structures, conditioned on and co-folded with arbitrary biomolecules — including small-molecule ligands, DNA, and RNA. Unlike sequential pipelines that first generate a backbone and then apply inverse folding, DISCO generates both modalities jointly, enabling sequence-based objectives to inform structure generation and vice versa.DISCO… See the full description on the dataset page: https://huggingface.co/datasets/DISCO-Design/DISCO_benchmark_data.dalle-3-LAION-discordwuerstchen-hugging-face-discordpublaynet_disco
PubLayNet_mini Dataset
A diverse mini subset of the PubLayNet dataset with 500 samples for document layout analysis evaluation.
Dataset Details
Total Samples: 500 document images
Source: PubLayNet training set (146,874 total samples)
Task: Document Layout Analysis
Format: Parquet with embedded images and annotations
Image Size: 612×792 pixels (RGB)
Categories: 5 layout element types
Categories
The dataset contains annotations for 5 categories of… See the full description on the dataset page: https://huggingface.co/datasets/kenza-ily/publaynet_disco.discord-new-avatarspokemons_discorddisco_diffusion_style_studies
all the grids + single images from the Disco Diffusion Style Studies as listed in this google sheet
read more about the project here
3568 names tested
767 other modifiers tested
discolookbook
Contact
For questions or issues, please contact: yhauri.contact@gmail.com
DiscoverStereo
