datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
DocQA_XMLcode10wiki90_sampling_xml_fiteredfull_coding_sampling_xml_fiteredXML-corpus-19k
Domofon XML Pretrain Corpus
19,204 unique SFT pairs for training language models to convert noisy/dirty text -> clean structured XML.
Task
Given a noisy, degraded text (markdown, HTML, plaintext with artifacts), produce a well-structured XML document that:
Preserves ALL content verbatim - no rewriting, no summarization
Recognizes latent structure (headings, lists, tables, code, quotes)
Wraps boilerplate (nav, footer, ads) in <aside role="boilerplate">
Uses only… See the full description on the dataset page: https://huggingface.co/datasets/domofon/XML-corpus-19k.Document-XML-100k
Document-XML-100k
Noisy/unstructured text to semantically tagged XML. 118K pairs for fine-tuning document markup models.
Splits
Split
Rows
Description
verified
79,067
Content-exact: byte-level match between input text and XML text content. Zero information loss guaranteed.
good
38,916
High quality (word overlap >= 85%, well-formed XML, no HTML tags) but with minor whitespace normalization.
What's the difference?
Both splits are… See the full description on the dataset page: https://huggingface.co/datasets/domofon/Document-XML-100k.Thinker-XMLSystem prompt suggestion:
You are a world-class AI system. Always respond in strict XML format with your reasoning steps within the <im_reasoning> XML tag. Each reasoning step should represent one unit of thought. Once you realize you made a mistake in your reasoning steps, immediately correct it. Place your final response outside the XML tag. Adhere to this XML structure without exception.
obf_gen_leave_out_code_full_xml_tags_seed_50crafter-gptoss120b-xmlobf_gen_leave_out_code_full_xml_tags_seed_42DCAgent2_terminal_bench_2_laion_exp_tas_parser_xml_traces_20260101_224249DCAgent2_terminal_bench_2_laion_exp_tas_parser_xml_traces_20260102_021553obf_gen_leave_out_sycophancy_full_xml_tags_seed_24code25wiki75_sampling_xml_fiteredDCAgent2_terminal_bench_2_laion_exp_tas_parser_xml_traces_20260101_132118DCAgent2_terminal_bench_2_laion_exp_tas_parser_xml_traces_20260102_000810exp_tas_parser_xml_tracesobf_gen_leave_out_sycophancy_full_xml_tags_seed_42code50wiki50_sampling_xml_fiteredthe-stack-mujoco-xmlQAT_alwatan_news.xml
Dataset Card for "QAT_alwatan_news.xml"
More Information needed
drawio_xml_instruction
Draw.io XML Instruction Dataset
A dataset containing natural language instructions paired with their corresponding draw.io XML diagram representations. This dataset enables the generation of professional diagrams from text descriptions.
Dataset Description
This dataset maps human-readable instructions to valid draw.io XML files. The diagrams cover a wide variety of types including:
Cloud Architecture (AWS, Azure, GCP)
Flowcharts & Process Diagrams
Network Topologies… See the full description on the dataset page: https://huggingface.co/datasets/ashutosh2211/drawio_xml_instruction.code5wiki95_sampling_xml_fiteredcrafter-gptoss120b-xml-improved15-long-11800578PIPPA-xml-promptsobsidian-agent-sft-xml-thinkobf_gen_leave_out_sycophancy_full_xml_tags_seed_420crafter-gptoss120b-xml-improved7-long-11421318
Crafter trajectories: openai/gpt-oss-120b (XML agent)
This dataset contains Crafter rollouts collected with an XML-based LLM agent.
Model: openai/gpt-oss-120b
Backend: vLLM (Helios GH200, 1 node, 4 GPUs)
Run/job id: 11421318
Episodes: 256
Columns
The main table includes (among others):
episode_id, episode_index
short_term_context, long_term_context
plans, raw_outputs
text_actions, actions
rewards, terms, truncs
image_paths (images; naming kept for backward… See the full description on the dataset page: https://huggingface.co/datasets/1337xyz1337xyz/crafter-gptoss120b-xml-improved7-long-11421318.obsidian-agent-sft-xmlredeIT-xml-ShareGPTbuzz_sources_080_xml
