yamaTK/merged_dataset_final_clean_v41
merged_dataset_final_clean_v41 English Rule-based cleaned SFT dataset for structured output generation (JSON / YAML / XML / TOML / CSV). Data Source This dataset was built from competition-provided datasets only. The cleaning pipeline loads the following source groups: u-10bei (6 datasets: source ids 1-1 to 1-6) daichira (3 datasets: source ids 2-1 to 2-3) After strict filtering and sampling for v4.1, the final retained rows are from u-10bei… See the full description on the dataset page: https://huggingface.co/datasets/yamaTK/merged_dataset_final_clean_v41.
Upload README.md with huggingface_hub
Upload README.md with huggingface_hub
Upload merged_dataset_final_clean_v41.jsonl with huggingface_hub
Upload README.md with huggingface_hub
initial commit
