domofon/Document-XML-100k
Document-XML-100k Noisy/unstructured text to semantically tagged XML. 118K pairs for fine-tuning document markup models. Splits Split Rows Description verified 79,067 Content-exact: byte-level match between input text and XML text content. Zero information loss guaranteed. good 38,916 High quality (word overlap >= 85%, well-formed XML, no HTML tags) but with minor whitespace normalization. What's the difference? Both splits are… See the full description on the dataset page: https://huggingface.co/datasets/domofon/Document-XML-100k.
052
