CoolFace
17 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01vGassen /Dutch-Basisbestandwetten-Legislation-Laws-XML-Cleantext10K<n<100K0 likes716 downloads1y agoHugging Face02xmlans /wikipedia-twtext1M<n<10M0 likes405 downloads2y agoHugging Face03xmlans /wikipedia_zhThink the datasets we've processed are helpful to you? Donate: https://ko-fi.com/stardreamcommunity By Star Dream Studio text1M<n<10M0 likes130 downloads2y agoHugging Face04madhukirangolla /genia-xml Biomedical Named Entity Recognition (NER) Dataset - XML Overview This dataset represents a processed subset of the biomedical literature, specifically formatted for training and evaluating Nested Named Entity Recognition (NER) models. The data consists of sentence-level pairs containing raw biomedical text and their corresponding entity-tagged representations. Data Structure Each entry contains: input_text: the raw, untagged biomedical sentence.… See the full description on the dataset page: https://huggingface.co/datasets/madhukirangolla/genia-xml.texttoken-classification10K<n<100K1 likes81 downloads1mo agoHugging Face05rbarnds /xml-standards-specificationstext1K<n<10K0 likes65 downloads3y agoHugging Face06kalomaze /md-2-xml-wiki-tables md-2-xml-wiki-tables 958 markdown tables extracted from fan/community MediaWiki sites for markdown-to-XML format conversion tasks. Format JSONL with fields: title: article title from the source wiki page section: section heading the table appeared under wiki: source wiki name table_md: raw markdown table filename: original filename Splits train: 894 tables eval: 64 held-out tables Source Various fan/community MediaWiki sites. Most use CC-BY-SA… See the full description on the dataset page: https://huggingface.co/datasets/kalomaze/md-2-xml-wiki-tables.texttext-generationn<1K1 likes28 downloads8mo agoHugging Face07kalomaze /semantic-xml-fwedu-2.6k-it1text1K<n<10K0 likes20 downloads1y agoHugging Face08daichira /structeval-t-sft-v2-xml StructEval-T SFT v2 - Full XML This dataset is the full, refined, format-specific subset for StructEval-T, focusing exclusively on strictly validated XML transformations. Key Features Total Samples: 4,503 Verified Quality: 100% strictly validated using AST/parsers (e.g. json.loads, xml.etree.ElementTree, yaml.safe_load). Only samples that successfully parse as valid XML without errors are included. Source: This is a split from the unified structeval-t-sft-v2 dataset.… See the full description on the dataset page: https://huggingface.co/datasets/daichira/structeval-t-sft-v2-xml.texttext-generation1K<n<10K0 likes16 downloads7mo agoHugging Face09Vendex /minicpm5-tool-calling-xmltextn<1K0 likes16 downloads2mo agoHugging Face10Vigneshwaran22 /XML_Datasettextn<1K0 likes12 downloads2y agoHugging Face11kiratan /structured-5k-mix-sft-xmltext1K<n<10K0 likes10 downloads7mo agoHugging Face12jbac208 /ws5_xmlstextn<1K2 likes8 downloads3y agoHugging Face13jbac208 /3.5turbo_ws5_xmltextn<1K0 likes7 downloads3y agoHugging Face14trentmkelly /easi-xml-chattext1K<n<10K0 likes7 downloads10mo agoHugging Face15daichira /structeval-t-sft-hq-xml StructEval-T SFT - High Quality XML This dataset is a highly refined, format-specific subset for StructEval-T, focusing exclusively on strictly validated XML transformations. Key Features Total Samples: 2,000 Verified Quality: 100% strictly validated using AST/parsers (e.g. json.loads, xml.etree.ElementTree, yaml.safe_load). Only samples that successfully parse as valid XML without errors are included. Goal: To maximize single-format fine-tuning performance or to be used… See the full description on the dataset page: https://huggingface.co/datasets/daichira/structeval-t-sft-hq-xml.texttext-generation1K<n<10K0 likes7 downloads7mo agoHugging Face16jbac208 /test_ws5_xmltextn<1K0 likes6 downloads3y agoHugging Face17mapurba /xml-policytextn<1K0 likes6 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.