CoolFace
Datasetpublic

TokenBender/glm47-synth-v1-dataset

Synth v1 Dataset This package contains 260 verified Aider-format SFT rows: ten synthetic variants for each of the 26 source task families. It is intentionally built for an exact 100-epoch memorization experiment. The training file is sft/train.jsonl. Every row uses the same nine-message aider-chat-v1 structure as the successful SFT-v5 package. Tests are not model-visible; the independent verifier replays each final assistant target against its source C++ test suite.… See the full description on the dataset page: https://huggingface.co/datasets/TokenBender/glm47-synth-v1-dataset.

sourceHugging Faceupdated 2mo agoView on Hugging Face
0likes10downloads
3 commits on main
c5864462mo ago

Rename Synth Mem v1 to Synth v1

TokenBender
36ac9922mo ago

Publish audited synth memorization dataset v1

TokenBender
61e55ff2mo ago

initial commit

TokenBender