datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
structeval-t-sft-v2-toml
StructEval-T SFT v2 - Full TOML
This dataset is the full, refined, format-specific subset for StructEval-T, focusing exclusively on strictly validated TOML transformations.
Key Features
Total Samples: 3,635
Verified Quality: 100% strictly validated using AST/parsers (e.g. json.loads, xml.etree.ElementTree, yaml.safe_load). Only samples that successfully parse as valid TOML without errors are included.
Source: This is a split from the unified structeval-t-sft-v2 dataset.… See the full description on the dataset page: https://huggingface.co/datasets/daichira/structeval-t-sft-v2-toml.structeval-t-sft-hq-yaml
StructEval-T SFT - High Quality YAML
This dataset is a highly refined, format-specific subset for StructEval-T, focusing exclusively on strictly validated YAML transformations.
Key Features
Total Samples: 2,000
Verified Quality: 100% strictly validated using AST/parsers (e.g. json.loads, xml.etree.ElementTree, yaml.safe_load). Only samples that successfully parse as valid YAML without errors are included.
Goal: To maximize single-format fine-tuning performance or to be… See the full description on the dataset page: https://huggingface.co/datasets/daichira/structeval-t-sft-hq-yaml.structeval-t-sft-v2-yaml
StructEval-T SFT v2 - Full YAML
This dataset is the full, refined, format-specific subset for StructEval-T, focusing exclusively on strictly validated YAML transformations.
Key Features
Total Samples: 5,628
Verified Quality: 100% strictly validated using AST/parsers (e.g. json.loads, xml.etree.ElementTree, yaml.safe_load). Only samples that successfully parse as valid YAML without errors are included.
Source: This is a split from the unified structeval-t-sft-v2 dataset.… See the full description on the dataset page: https://huggingface.co/datasets/daichira/structeval-t-sft-v2-yaml.structeval-t-sft-v2-json
StructEval-T SFT v2 - Full JSON
This dataset is the full, refined, format-specific subset for StructEval-T, focusing exclusively on strictly validated JSON transformations.
Key Features
Total Samples: 3,059
Verified Quality: 100% strictly validated using AST/parsers (e.g. json.loads, xml.etree.ElementTree, yaml.safe_load). Only samples that successfully parse as valid JSON without errors are included.
Source: This is a split from the unified structeval-t-sft-v2 dataset.… See the full description on the dataset page: https://huggingface.co/datasets/daichira/structeval-t-sft-v2-json.structeval-t-sft-hq-json
StructEval-T SFT - High Quality JSON
This dataset is a highly refined, format-specific subset for StructEval-T, focusing exclusively on strictly validated JSON transformations.
Key Features
Total Samples: 2,000
Verified Quality: 100% strictly validated using AST/parsers (e.g. json.loads, xml.etree.ElementTree, yaml.safe_load). Only samples that successfully parse as valid JSON without errors are included.
Goal: To maximize single-format fine-tuning performance or to be… See the full description on the dataset page: https://huggingface.co/datasets/daichira/structeval-t-sft-hq-json.structeval-t-sft-hq-csv
StructEval-T SFT - High Quality CSV
This dataset is a highly refined, format-specific subset for StructEval-T, focusing exclusively on strictly validated CSV transformations.
Key Features
Total Samples: 2,000
Verified Quality: 100% strictly validated using AST/parsers (e.g. json.loads, xml.etree.ElementTree, yaml.safe_load). Only samples that successfully parse as valid CSV without errors are included.
Goal: To maximize single-format fine-tuning performance or to be used… See the full description on the dataset page: https://huggingface.co/datasets/daichira/structeval-t-sft-hq-csv.structeval-t-sft-v2-xml
StructEval-T SFT v2 - Full XML
This dataset is the full, refined, format-specific subset for StructEval-T, focusing exclusively on strictly validated XML transformations.
Key Features
Total Samples: 4,503
Verified Quality: 100% strictly validated using AST/parsers (e.g. json.loads, xml.etree.ElementTree, yaml.safe_load). Only samples that successfully parse as valid XML without errors are included.
Source: This is a split from the unified structeval-t-sft-v2 dataset.… See the full description on the dataset page: https://huggingface.co/datasets/daichira/structeval-t-sft-v2-xml.structeval-t-sft-v2-csv
StructEval-T SFT v2 - Full CSV
This dataset is the full, refined, format-specific subset for StructEval-T, focusing exclusively on strictly validated CSV transformations.
Key Features
Total Samples: 2,104
Verified Quality: 100% strictly validated using AST/parsers (e.g. json.loads, xml.etree.ElementTree, yaml.safe_load). Only samples that successfully parse as valid CSV without errors are included.
Source: This is a split from the unified structeval-t-sft-v2 dataset.… See the full description on the dataset page: https://huggingface.co/datasets/daichira/structeval-t-sft-v2-csv.structeval-t-sft-hq-toml
StructEval-T SFT - High Quality TOML
This dataset is a highly refined, format-specific subset for StructEval-T, focusing exclusively on strictly validated TOML transformations.
Key Features
Total Samples: 2,000
Verified Quality: 100% strictly validated using AST/parsers (e.g. json.loads, xml.etree.ElementTree, yaml.safe_load). Only samples that successfully parse as valid TOML without errors are included.
Goal: To maximize single-format fine-tuning performance or to be… See the full description on the dataset page: https://huggingface.co/datasets/daichira/structeval-t-sft-hq-toml.structeval-t-sft-hq-xml
StructEval-T SFT - High Quality XML
This dataset is a highly refined, format-specific subset for StructEval-T, focusing exclusively on strictly validated XML transformations.
Key Features
Total Samples: 2,000
Verified Quality: 100% strictly validated using AST/parsers (e.g. json.loads, xml.etree.ElementTree, yaml.safe_load). Only samples that successfully parse as valid XML without errors are included.
Goal: To maximize single-format fine-tuning performance or to be used… See the full description on the dataset page: https://huggingface.co/datasets/daichira/structeval-t-sft-hq-xml.ise-uiuc_Magicoder-Evol-Instruct-110K_structeval_processed
Magicoder-Evol-Instruct-110K (StructEval SFT Processed)
Processed derivative of ise-uiuc/Magicoder-Evol-Instruct-110K for StructEval-style structured output SFT.
Source dataset: ise-uiuc/Magicoder-Evol-Instruct-110K
Source URL: https://huggingface.co/datasets/ise-uiuc/Magicoder-Evol-Instruct-110K
Source license (from dataset card metadata): apache-2.0
Processed dataset: daichira/ise-uiuc_Magicoder-Evol-Instruct-110K_structeval_processed
Processed URL:… See the full description on the dataset page: https://huggingface.co/datasets/daichira/ise-uiuc_Magicoder-Evol-Instruct-110K_structeval_processed.perfect-structeval-sft-dataset-v3perfect-structeval-sft-datasetperfect-structeval-sft-dataset-v2
