datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
MermaidSeqBench
Dataset Card for MermaidSeqBench
Dataset Summary
This dataset provides a human-verified benchmark for assessing large language models (LLMs) on their ability to generate Mermaid sequence diagrams from natural language prompts.
The dataset was synthetically generated using large language models (LLMs), starting from a small set of seed examples provided by a subject-matter expert. All outputs were subsequently manually verified and corrected by human annotators to ensure… See the full description on the dataset page: https://huggingface.co/datasets/ibm-research/MermaidSeqBench.text-to-mermaidtext-to-mermaid-2text-to-mermaid
Text to Mermaid
Description
A curated dataset designed for fine-tuning small language models on diagram generation tasks. Each example follows a strict response format: when prompted for a diagram, the model outputs only the Mermaid syntax—no explanations, no markdown fences, no extra text.
Derived from Celiadraw/text-to-mermaid-2, this version has been verified, pruned, cleaned, and reworded for reliability and consistency.
Dataset Statistics… See the full description on the dataset page: https://huggingface.co/datasets/ali-thowfeek/text-to-mermaid.mermaid-rep-of-pd-patchfiles
