datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
obsidian-agent-sft-xml-thinkDocQA_XMLcrafter-gptoss120b-xmlcode10wiki90_sampling_xml_fiteredfull_coding_sampling_xml_fiteredXML-corpus-19k
Domofon XML Pretrain Corpus
19,204 unique SFT pairs for training language models to convert noisy/dirty text -> clean structured XML.
Task
Given a noisy, degraded text (markdown, HTML, plaintext with artifacts), produce a well-structured XML document that:
Preserves ALL content verbatim - no rewriting, no summarization
Recognizes latent structure (headings, lists, tables, code, quotes)
Wraps boilerplate (nav, footer, ads) in <aside role="boilerplate">
Uses only… See the full description on the dataset page: https://huggingface.co/datasets/domofon/XML-corpus-19k.Document-XML-100k
Document-XML-100k
Noisy/unstructured text to semantically tagged XML. 118K pairs for fine-tuning document markup models.
Splits
Split
Rows
Description
verified
79,067
Content-exact: byte-level match between input text and XML text content. Zero information loss guaranteed.
good
38,916
High quality (word overlap >= 85%, well-formed XML, no HTML tags) but with minor whitespace normalization.
What's the difference?
Both splits are… See the full description on the dataset page: https://huggingface.co/datasets/domofon/Document-XML-100k.obf_gen_leave_out_code_full_xml_tags_seed_50code25wiki75_sampling_xml_fiteredobf_gen_leave_out_code_full_xml_tags_seed_42obf_gen_leave_out_sycophancy_full_xml_tags_seed_24DCAgent2_terminal_bench_2_laion_exp_tas_parser_xml_traces_20260101_224249drawio_xml_instruction
Draw.io XML Instruction Dataset
A dataset containing natural language instructions paired with their corresponding draw.io XML diagram representations. This dataset enables the generation of professional diagrams from text descriptions.
Dataset Description
This dataset maps human-readable instructions to valid draw.io XML files. The diagrams cover a wide variety of types including:
Cloud Architecture (AWS, Azure, GCP)
Flowcharts & Process Diagrams
Network Topologies… See the full description on the dataset page: https://huggingface.co/datasets/ashutosh2211/drawio_xml_instruction.Thinker-XMLSystem prompt suggestion:
You are a world-class AI system. Always respond in strict XML format with your reasoning steps within the <im_reasoning> XML tag. Each reasoning step should represent one unit of thought. Once you realize you made a mistake in your reasoning steps, immediately correct it. Place your final response outside the XML tag. Adhere to this XML structure without exception.
DCAgent2_terminal_bench_2_laion_exp_tas_parser_xml_traces_20260101_132118DCAgent2_terminal_bench_2_laion_exp_tas_parser_xml_traces_20260102_000810DCAgent2_terminal_bench_2_laion_exp_tas_parser_xml_traces_20260102_021553code50wiki50_sampling_xml_fiteredexp_tas_parser_xml_tracesobf_gen_leave_out_sycophancy_full_xml_tags_seed_42obsidian-agent-sft-xmlcode5wiki95_sampling_xml_fiteredthe-stack-mujoco-xmlQAT_alwatan_news.xml
Dataset Card for "QAT_alwatan_news.xml"
More Information needed
redeIT-xml-ShareGPTPIPPA-xml-promptsratsmanuale_1465-raw-xml
Dataset Card for ratsmanuale_1465-raw-xml
This dataset was created using pagexml-hf converter from Transkribus PageXML data.
Dataset Summary
This dataset contains 64 samples across 1 split(s).
Dataset Structure
Data Splits
train: 64 samples
Dataset Size
Approximate total size: 383.09 MB
Total samples: 64
Features
image: Image(mode=None, decode=False)
xml_content: Value('string')
filename:… See the full description on the dataset page: https://huggingface.co/datasets/ChantalMarbach/ratsmanuale_1465-raw-xml.obf_gen_leave_out_sycophancy_full_xml_tags_seed_420crafter-gptoss120b-xml-improved15-long-11800578gazefollow_xml_int
