CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01AtakanTekparmak /obsidian-agent-sft-xml-thinktextn<1K0 likes101 downloads1y agoHugging Face02EtashGuha /DocQA_XMLimage1K<n<10K3 likes82 downloads2y agoHugging Face031337xyz1337xyz /crafter-gptoss120b-xmlimage10K<n<100K0 likes65 downloads8mo agoHugging Face04ZorraZabb /code10wiki90_sampling_xml_fiteredtext1M<n<10M1 likes64 downloads2y agoHugging Face05ZorraZabb /full_coding_sampling_xml_fiteredtabular1M<n<10M0 likes63 downloads2y agoHugging Face06domofon /XML-corpus-19k Domofon XML Pretrain Corpus 19,204 unique SFT pairs for training language models to convert noisy/dirty text -> clean structured XML. Task Given a noisy, degraded text (markdown, HTML, plaintext with artifacts), produce a well-structured XML document that: Preserves ALL content verbatim - no rewriting, no summarization Recognizes latent structure (headings, lists, tables, code, quotes) Wraps boilerplate (nav, footer, ads) in <aside role="boilerplate"> Uses only… See the full description on the dataset page: https://huggingface.co/datasets/domofon/XML-corpus-19k.text10K<n<100K0 likes53 downloads4mo agoHugging Face07domofon /Document-XML-100k Document-XML-100k Noisy/unstructured text to semantically tagged XML. 118K pairs for fine-tuning document markup models. Splits Split Rows Description verified 79,067 Content-exact: byte-level match between input text and XML text content. Zero information loss guaranteed. good 38,916 High quality (word overlap >= 85%, well-formed XML, no HTML tags) but with minor whitespace normalization. What's the difference? Both splits are… See the full description on the dataset page: https://huggingface.co/datasets/domofon/Document-XML-100k.texttext-generation100K<n<1M0 likes53 downloads4mo agoHugging Face08sb2700 /obf_gen_leave_out_code_full_xml_tags_seed_50text1K<n<10K0 likes47 downloads9mo agoHugging Face09ZorraZabb /code25wiki75_sampling_xml_fiteredtext1M<n<10M0 likes46 downloads2y agoHugging Face10sb2700 /obf_gen_leave_out_code_full_xml_tags_seed_42text1K<n<10K0 likes42 downloads9mo agoHugging Face11sb2700 /obf_gen_leave_out_sycophancy_full_xml_tags_seed_24text1K<n<10K0 likes38 downloads9mo agoHugging Face12DCAgent2 /DCAgent2_terminal_bench_2_laion_exp_tas_parser_xml_traces_20260101_224249textn<1K0 likes37 downloads9mo agoHugging Face13ashutosh2211 /drawio_xml_instruction Draw.io XML Instruction Dataset A dataset containing natural language instructions paired with their corresponding draw.io XML diagram representations. This dataset enables the generation of professional diagrams from text descriptions. Dataset Description This dataset maps human-readable instructions to valid draw.io XML files. The diagrams cover a wide variety of types including: Cloud Architecture (AWS, Azure, GCP) Flowcharts & Process Diagrams Network Topologies… See the full description on the dataset page: https://huggingface.co/datasets/ashutosh2211/drawio_xml_instruction.texttext-generationn<1K0 likes37 downloads8mo agoHugging Face14minchyeom /Thinker-XMLSystem prompt suggestion: You are a world-class AI system. Always respond in strict XML format with your reasoning steps within the <im_reasoning> XML tag. Each reasoning step should represent one unit of thought. Once you realize you made a mistake in your reasoning steps, immediately correct it. Place your final response outside the XML tag. Adhere to this XML structure without exception. texttext-generation1K<n<10K1 likes36 downloads2y agoHugging Face15DCAgent2 /DCAgent2_terminal_bench_2_laion_exp_tas_parser_xml_traces_20260101_132118textn<1K0 likes36 downloads9mo agoHugging Face16DCAgent2 /DCAgent2_terminal_bench_2_laion_exp_tas_parser_xml_traces_20260102_000810textn<1K0 likes36 downloads9mo agoHugging Face17DCAgent2 /DCAgent2_terminal_bench_2_laion_exp_tas_parser_xml_traces_20260102_021553textn<1K0 likes36 downloads9mo agoHugging Face18ZorraZabb /code50wiki50_sampling_xml_fiteredtext1M<n<10M0 likes34 downloads2y agoHugging Face19DCAgent /exp_tas_parser_xml_tracestext10K<n<100K0 likes34 downloads9mo agoHugging Face20sb2700 /obf_gen_leave_out_sycophancy_full_xml_tags_seed_42text1K<n<10K0 likes34 downloads9mo agoHugging Face21AtakanTekparmak /obsidian-agent-sft-xmltextn<1K0 likes33 downloads1y agoHugging Face22ZorraZabb /code5wiki95_sampling_xml_fiteredtext1M<n<10M0 likes30 downloads2y agoHugging Face23reshinthadith /the-stack-mujoco-xmltabular10K<n<100K1 likes29 downloads2y agoHugging Face24omarelsayeed /QAT_alwatan_news.xml Dataset Card for "QAT_alwatan_news.xml" More Information needed text10K<n<100K0 likes28 downloads3y agoHugging Face25leinad-deinor /redeIT-xml-ShareGPTtextn<1K0 likes27 downloads2y agoHugging Face26KeeganCarey /PIPPA-xml-promptstext10K<n<100K1 likes25 downloads6mo agoHugging Face27ChantalMarbach /ratsmanuale_1465-raw-xml Dataset Card for ratsmanuale_1465-raw-xml This dataset was created using pagexml-hf converter from Transkribus PageXML data. Dataset Summary This dataset contains 64 samples across 1 split(s). Dataset Structure Data Splits train: 64 samples Dataset Size Approximate total size: 383.09 MB Total samples: 64 Features image: Image(mode=None, decode=False) xml_content: Value('string') filename:… See the full description on the dataset page: https://huggingface.co/datasets/ChantalMarbach/ratsmanuale_1465-raw-xml.imagen<1K2 likes25 downloads4mo agoHugging Face28sb2700 /obf_gen_leave_out_sycophancy_full_xml_tags_seed_420text1K<n<10K0 likes24 downloads8mo agoHugging Face291337xyz1337xyz /crafter-gptoss120b-xml-improved15-long-11800578image100K<n<1M0 likes24 downloads8mo agoHugging Face30ShijianDeng /gazefollow_xml_intimage100K<n<1M0 likes22 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.