datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
image-to-mermaid-v3
Image-to-Mermaid Dataset v3.0
Task: Given a rendered diagram image, generate the equivalent Mermaid source code that reproduces it.
This is a synthetic, multi-type extension of an earlier flowchart-only release (2,569 samples, tied to an accepted MAPR2026 IEEE paper on flowchart QA). v3.0 covers 26 Mermaid diagram types across 3 difficulty tiers each.
Statistics
Split
Samples
Share
Train
12,046
79.8%
Val
1,475
9.8%
Test
1,578
10.5%
Total
15,099… See the full description on the dataset page: https://huggingface.co/datasets/DangIT02/image-to-mermaid-v3.mermaid-persona-queries
Mermaid Persona Queries (v1)
50 hand-written user queries for evaluating and benchmarking text-to-Mermaid
diagram generation, authored across 5 distinct user personas. Users exist
within a single company spanning many roles (engineering, HR, finance, ops,
support, product, legal, IT) — the dataset varies how people ask, not a
business domain.
Motivation
Real users of text-to-diagram tools do not write uniform prompts. Some paste
truncated meeting notes, some… See the full description on the dataset page: https://huggingface.co/datasets/martincousseau/mermaid-persona-queries.broken-mermaid-diagramsMermaid_500k_1KSample
Mermaid Expert Corpus 500k Sample
This repository contains a free public sample of the Mermaid Expert Corpus: a commercial-grade text-to-Mermaid dataset built for companies training, evaluating, benchmarking, or routing diagram-generation models.
The sample includes 1,000 records and their rendered SVGs. It is designed to show the shape, metadata richness, diagram quality, and commercial relevance of the full 500,000-record corpus without exposing the licensed dataset itself.… See the full description on the dataset page: https://huggingface.co/datasets/CoreFidelity/Mermaid_500k_1KSample.Mermaid_50K
Mermaid 50K
Deprecated dataset. This repo is preserved for reference only.The recommended current dataset is CoreFidelity/Mermaid_500k.Commercial licensing inquiries: corefidelity@proton.me
Mermaid 50K was an early synthetic dataset of 50,000 paired natural-language process descriptions and Mermaid flowchart diagrams, with a matching browser-rendered SVG for every diagram.
This dataset has now been superseded by Mermaid 500k, a substantially larger and higher-quality corpus… See the full description on the dataset page: https://huggingface.co/datasets/CoreFidelity/Mermaid_50K.mermaid_codeTender_Mermaid_Training_V1
A reasoning enabled dataset derived from claude opus 4.6 on max reasoning and verified through mermaid.live.
This dataset is was built and intended to be used with my Fred Series models
text-to-mermaid-koreanmermaid-syntheticMermaidmermaid
