datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
T2I-CoReBench
Easier Painting Than Thinking: Can Text-to-Image Models
Set the Stage, but Not Direct the Play?
Ouxiang Li1*, Yuan Wang1, Xinting Hu†, Huijuan Huang2‡, Rui Chen2, Jiarong Ou2,
Xin Tao2†, Pengfei Wan2, Xiaojuan Qi3, Fuli Feng1
1University of Science and Technology of China, 2Kling Team, Kuaishou Technology, 3The University of Hong Kong
*Work done during internship… See the full description on the dataset page: https://huggingface.co/datasets/lioooox/T2I-CoReBench.Mermaid_500k_1KSample
Mermaid Expert Corpus 500k Sample
This repository contains a free public sample of the Mermaid Expert Corpus: a commercial-grade text-to-Mermaid dataset built for companies training, evaluating, benchmarking, or routing diagram-generation models.
The sample includes 1,000 records and their rendered SVGs. It is designed to show the shape, metadata richness, diagram quality, and commercial relevance of the full 500,000-record corpus without exposing the licensed dataset itself.… See the full description on the dataset page: https://huggingface.co/datasets/CoreFidelity/Mermaid_500k_1KSample.Mermaid_50K
Mermaid 50K
Deprecated dataset. This repo is preserved for reference only.The recommended current dataset is CoreFidelity/Mermaid_500k.Commercial licensing inquiries: corefidelity@proton.me
Mermaid 50K was an early synthetic dataset of 50,000 paired natural-language process descriptions and Mermaid flowchart diagrams, with a matching browser-rendered SVG for every diagram.
This dataset has now been superseded by Mermaid 500k, a substantially larger and higher-quality corpus… See the full description on the dataset page: https://huggingface.co/datasets/CoreFidelity/Mermaid_50K.
