datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
x86_64-freestanding-corpus
x86_64-freestanding-corpus
94,546 compilable and executable x86_64 code samples for freestanding, no-libc
systems programming. Every row has a natural language prompt, difficulty level,
and tag-based filtering. 93,986 rows (99.4%) are verified
to compile with -Werror and exit cleanly when executed. ~10,017,133 tokens total.
What this dataset is (and isn't)
This is a domain-adaptation + SFT corpus for an existing code model, not a
from-scratch pretraining set. At… See the full description on the dataset page: https://huggingface.co/datasets/Kureiwa/x86_64-freestanding-corpus.paired_aarch64-x86aligned_aarch64_x86_64
