datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
patent-spec-xmlstackexchange_xmlThis is a dump of the files from
https://archive.org/details/stackexchange
downloaded via torrent on 2021-07-01.
Publication date 2021-06-07 Usage Attribution-ShareAlike 4.0 International Creative Commons License by sa Topics Stack Exchange Data Dump Contributor Stack Exchange Community
Please see the license information at:
https://archive.org/details/stackexchange
The dataset has been split into following for cleaner formatting.… See the full description on the dataset page: https://huggingface.co/datasets/flax-sentence-embeddings/stackexchange_xml.enwiki-2024-04-20-xmlpmc_open_access_xml
Dataset Card for PMC Open Access XML
Dataset Summary
The XML Open Access includes more than 3.4 million journal articles and preprints that are made available under
license terms that allow reuse.
Not all articles in PMC are available for text mining and other reuse, many have copyright protection, however articles
in the PMC Open Access Subset are made available under Creative Commons or similar licenses that generally allow more
liberal redistribution and reuse than a… See the full description on the dataset page: https://huggingface.co/datasets/TomTBT/pmc_open_access_xml.medical_nih_pubmed_xml
language:
en
name: pubmed_xml
Dutch-Basisbestandwetten-Legislation-Laws-XML-Cleanenwiki-simple-xml_2025-07-metadatawikipedia-twdrawio-xmlclinical-trials-xml-2018-2024Cowpea-Architecture-XML
Cowpea-Architecture-XML-WDS
This dataset contains simulated images of Cowpea plants paired with organ-level architecture representations in XML format, packaged in WebDataset (.tar) format for efficient high-performance training.
Dataset Structure
The dataset is sharded into .tar files, each containing up to 10,000 samples.
Each sample consists of:
.jpeg: The plant image
.xml: The organ-level architecture representation
.json: (Optional) Metadata
Usage with… See the full description on the dataset page: https://huggingface.co/datasets/heesup/Cowpea-Architecture-XML.wikipedia_zhThink the datasets we've processed are helpful to you? Donate: https://ko-fi.com/stardreamcommunity
By Star Dream Studio
obsidian-agent-sft-xml-thinkCowpea-Architecture-XML
Cowpea-Architecture-XML-WDS
This dataset contains simulated images of Cowpea plants paired with organ-level architecture representations in XML format, packaged in WebDataset (.tar) format for efficient high-performance training.
Dataset Structure
The dataset is sharded into .tar files, each containing up to 10,000 samples.
Each sample consists of:
.jpeg: The plant image
.xml: The organ-level architecture representation
.json: (Optional) Metadata… See the full description on the dataset page: https://huggingface.co/datasets/bbrangeo/Cowpea-Architecture-XML.DocQA_XMLgenia-xml
Biomedical Named Entity Recognition (NER) Dataset - XML
Overview
This dataset represents a processed subset of the biomedical literature, specifically formatted for training and evaluating Nested Named Entity Recognition (NER) models. The data consists of sentence-level pairs containing raw biomedical text and their corresponding entity-tagged representations.
Data Structure
Each entry contains:
input_text: the raw, untagged biomedical sentence.… See the full description on the dataset page: https://huggingface.co/datasets/madhukirangolla/genia-xml.JRC-Acquis_subcorpus_EL-FR_Vanilla_aligned-XML
[!NOTE]
Dataset origin: https://inventory.clarin.gr/corpus/564
Description
The JRC-Acquis subcorpus EL-FR (Vanilla aligned-XML) is a parallel subcorpus for French and Greek, subset of the JRC-Acquis Multilingual Parallel Corpus.
Citation
Joint Research Centre - European Commission (2015). JRC-Acquis subcorpus EL-FR (Vanilla aligned-XML). [Dataset (Text corpus)]. CLARIN:EL. http://hdl.handle.net/11500/CLARIN-EL-0000-0000-69DA-5
WM-ORG-XML-DUMP_FR_2026.08
Jeux De Données : Dump WikiMedia Français Août 2026 : Extraction et Nettoyage Complet
Description
Ce JDD contient des articles extraits du dump complet de Wikimedia d'août 2026, nettoyés et structurés pour l'entraînement de modèles d'apprentissage automatique.
Source originale : https://www.wikimedia.org/
Licence : https://creativecommons.org/licenses/by-sa/4.0/deed.en
Dump source : https://dumps.wikimedia.org/frwiki/latest/
Fichiers traités (4 fichiers «… See the full description on the dataset page: https://huggingface.co/datasets/MisterAI/WM-ORG-XML-DUMP_FR_2026.08.JRC-Acquis_subcorpus_EL-FR_Hunalign_aligned-XML
[!NOTE]
Dataset origin: https://inventory.clarin.gr/corpus/540
Description
The JRC-Acquis subcorpus EL-FR (Hunalign aligned-XML) is a parallel subcorpus for French and Greek, subset of the JRC-Acquis Multilingual Parallel Corpus.
Citation
Joint Research Centre - European Commission (2015). JRC-Acquis subcorpus EL-FR (Hunalign aligned-XML). [Dataset (Text corpus)]. CLARIN:EL. http://hdl.handle.net/11500/CLARIN-EL-0000-0000-69F2-9
crafter-gptoss120b-xmlxml-standards-specificationsfull_coding_sampling_xml_fiteredcode10wiki90_sampling_xml_fiteredXML-corpus-19k
Domofon XML Pretrain Corpus
19,204 unique SFT pairs for training language models to convert noisy/dirty text -> clean structured XML.
Task
Given a noisy, degraded text (markdown, HTML, plaintext with artifacts), produce a well-structured XML document that:
Preserves ALL content verbatim - no rewriting, no summarization
Recognizes latent structure (headings, lists, tables, code, quotes)
Wraps boilerplate (nav, footer, ads) in <aside role="boilerplate">
Uses only… See the full description on the dataset page: https://huggingface.co/datasets/domofon/XML-corpus-19k.Document-XML-100k
Document-XML-100k
Noisy/unstructured text to semantically tagged XML. 118K pairs for fine-tuning document markup models.
Splits
Split
Rows
Description
verified
79,067
Content-exact: byte-level match between input text and XML text content. Zero information loss guaranteed.
good
38,916
High quality (word overlap >= 85%, well-formed XML, no HTML tags) but with minor whitespace normalization.
What's the difference?
Both splits are… See the full description on the dataset page: https://huggingface.co/datasets/domofon/Document-XML-100k.obf_gen_leave_out_code_full_xml_tags_seed_50code25wiki75_sampling_xml_fiteredobf_gen_leave_out_code_full_xml_tags_seed_42obsidian-agent-sft-xmlxml-dump
