datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
drawio-xmlWM-ORG-XML-DUMP_FR_2026.08
Jeux De Données : Dump WikiMedia Français Août 2026 : Extraction et Nettoyage Complet
Description
Ce JDD contient des articles extraits du dump complet de Wikimedia d'août 2026, nettoyés et structurés pour l'entraînement de modèles d'apprentissage automatique.
Source originale : https://www.wikimedia.org/
Licence : https://creativecommons.org/licenses/by-sa/4.0/deed.en
Dump source : https://dumps.wikimedia.org/frwiki/latest/
Fichiers traités (4 fichiers «… See the full description on the dataset page: https://huggingface.co/datasets/MisterAI/WM-ORG-XML-DUMP_FR_2026.08.Document-XML-100k
Document-XML-100k
Noisy/unstructured text to semantically tagged XML. 118K pairs for fine-tuning document markup models.
Splits
Split
Rows
Description
verified
79,067
Content-exact: byte-level match between input text and XML text content. Zero information loss guaranteed.
good
38,916
High quality (word overlap >= 85%, well-formed XML, no HTML tags) but with minor whitespace normalization.
What's the difference?
Both splits are… See the full description on the dataset page: https://huggingface.co/datasets/domofon/Document-XML-100k.drawio_xml_instruction
Draw.io XML Instruction Dataset
A dataset containing natural language instructions paired with their corresponding draw.io XML diagram representations. This dataset enables the generation of professional diagrams from text descriptions.
Dataset Description
This dataset maps human-readable instructions to valid draw.io XML files. The diagrams cover a wide variety of types including:
Cloud Architecture (AWS, Azure, GCP)
Flowcharts & Process Diagrams
Network Topologies… See the full description on the dataset page: https://huggingface.co/datasets/ashutosh2211/drawio_xml_instruction.Thinker-XMLSystem prompt suggestion:
You are a world-class AI system. Always respond in strict XML format with your reasoning steps within the <im_reasoning> XML tag. Each reasoning step should represent one unit of thought. Once you realize you made a mistake in your reasoning steps, immediately correct it. Place your final response outside the XML tag. Adhere to this XML structure without exception.
md-2-xml-wiki-tables
md-2-xml-wiki-tables
958 markdown tables extracted from fan/community MediaWiki sites for markdown-to-XML format conversion tasks.
Format
JSONL with fields:
title: article title from the source wiki page
section: section heading the table appeared under
wiki: source wiki name
table_md: raw markdown table
filename: original filename
Splits
train: 894 tables
eval: 64 held-out tables
Source
Various fan/community MediaWiki sites. Most use CC-BY-SA… See the full description on the dataset page: https://huggingface.co/datasets/kalomaze/md-2-xml-wiki-tables.structeval-t-sft-v2-xml
StructEval-T SFT v2 - Full XML
This dataset is the full, refined, format-specific subset for StructEval-T, focusing exclusively on strictly validated XML transformations.
Key Features
Total Samples: 4,503
Verified Quality: 100% strictly validated using AST/parsers (e.g. json.loads, xml.etree.ElementTree, yaml.safe_load). Only samples that successfully parse as valid XML without errors are included.
Source: This is a split from the unified structeval-t-sft-v2 dataset.… See the full description on the dataset page: https://huggingface.co/datasets/daichira/structeval-t-sft-v2-xml.Thinker-XML-2Suggested system prompt:
Respond to each user instruction in an XML format, using <step> tags to document your logical reasoning process step-by-step, while the <output> tag should be reserved for your final communication with the user. Incorporate self-correction by reflecting on prior steps; if a previous thought requires adjustment, add a new <step> to refine your reasoning without altering the original. Include self-reflection by periodically assessing your thought process and noting any… See the full description on the dataset page: https://huggingface.co/datasets/minchyeom/Thinker-XML-2.structeval-t-sft-hq-xml
StructEval-T SFT - High Quality XML
This dataset is a highly refined, format-specific subset for StructEval-T, focusing exclusively on strictly validated XML transformations.
Key Features
Total Samples: 2,000
Verified Quality: 100% strictly validated using AST/parsers (e.g. json.loads, xml.etree.ElementTree, yaml.safe_load). Only samples that successfully parse as valid XML without errors are included.
Goal: To maximize single-format fine-tuning performance or to be used… See the full description on the dataset page: https://huggingface.co/datasets/daichira/structeval-t-sft-hq-xml.
