CoolFace
Datasetpublic

yamaTK/merged_dataset_final_clean_v41

merged_dataset_final_clean_v41 English Rule-based cleaned SFT dataset for structured output generation (JSON / YAML / XML / TOML / CSV). Data Source This dataset was built from competition-provided datasets only. The cleaning pipeline loads the following source groups: u-10bei (6 datasets: source ids 1-1 to 1-6) daichira (3 datasets: source ids 2-1 to 2-3) After strict filtering and sampling for v4.1, the final retained rows are from u-10bei… See the full description on the dataset page: https://huggingface.co/datasets/yamaTK/merged_dataset_final_clean_v41.

sourceHugging Faceotherupdated 7mo agoView on Hugging Face
0likes21downloads
5 commits on main
213f61a7mo ago

Upload README.md with huggingface_hub

yamaTK
772449a7mo ago

Upload README.md with huggingface_hub

yamaTK
17c627a7mo ago

Upload merged_dataset_final_clean_v41.jsonl with huggingface_hub

yamaTK
f1dc5117mo ago

Upload README.md with huggingface_hub

yamaTK
3de9e627mo ago

initial commit

yamaTK