datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Dutch-Basisbestandwetten-Legislation-Laws-XML-Cleanwikipedia-twwikipedia_zhThink the datasets we've processed are helpful to you? Donate: https://ko-fi.com/stardreamcommunity
By Star Dream Studio
genia-xml
Biomedical Named Entity Recognition (NER) Dataset - XML
Overview
This dataset represents a processed subset of the biomedical literature, specifically formatted for training and evaluating Nested Named Entity Recognition (NER) models. The data consists of sentence-level pairs containing raw biomedical text and their corresponding entity-tagged representations.
Data Structure
Each entry contains:
input_text: the raw, untagged biomedical sentence.… See the full description on the dataset page: https://huggingface.co/datasets/madhukirangolla/genia-xml.xml-standards-specificationsmd-2-xml-wiki-tables
md-2-xml-wiki-tables
958 markdown tables extracted from fan/community MediaWiki sites for markdown-to-XML format conversion tasks.
Format
JSONL with fields:
title: article title from the source wiki page
section: section heading the table appeared under
wiki: source wiki name
table_md: raw markdown table
filename: original filename
Splits
train: 894 tables
eval: 64 held-out tables
Source
Various fan/community MediaWiki sites. Most use CC-BY-SA… See the full description on the dataset page: https://huggingface.co/datasets/kalomaze/md-2-xml-wiki-tables.semantic-xml-fwedu-2.6k-it1structeval-t-sft-v2-xml
StructEval-T SFT v2 - Full XML
This dataset is the full, refined, format-specific subset for StructEval-T, focusing exclusively on strictly validated XML transformations.
Key Features
Total Samples: 4,503
Verified Quality: 100% strictly validated using AST/parsers (e.g. json.loads, xml.etree.ElementTree, yaml.safe_load). Only samples that successfully parse as valid XML without errors are included.
Source: This is a split from the unified structeval-t-sft-v2 dataset.… See the full description on the dataset page: https://huggingface.co/datasets/daichira/structeval-t-sft-v2-xml.minicpm5-tool-calling-xmlXML_Datasetstructured-5k-mix-sft-xmlws5_xmls3.5turbo_ws5_xmleasi-xml-chatstructeval-t-sft-hq-xml
StructEval-T SFT - High Quality XML
This dataset is a highly refined, format-specific subset for StructEval-T, focusing exclusively on strictly validated XML transformations.
Key Features
Total Samples: 2,000
Verified Quality: 100% strictly validated using AST/parsers (e.g. json.loads, xml.etree.ElementTree, yaml.safe_load). Only samples that successfully parse as valid XML without errors are included.
Goal: To maximize single-format fine-tuning performance or to be used… See the full description on the dataset page: https://huggingface.co/datasets/daichira/structeval-t-sft-hq-xml.test_ws5_xmlxml-policy
