datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
x86_64-freestanding-corpus
x86_64-freestanding-corpus
94,546 compilable and executable x86_64 code samples for freestanding, no-libc
systems programming. Every row has a natural language prompt, difficulty level,
and tag-based filtering. 93,986 rows (99.4%) are verified
to compile with -Werror and exit cleanly when executed. ~10,017,133 tokens total.
What this dataset is (and isn't)
This is a domain-adaptation + SFT corpus for an existing code model, not a
from-scratch pretraining set. At… See the full description on the dataset page: https://huggingface.co/datasets/Kureiwa/x86_64-freestanding-corpus.paired_aarch64-x86flash_attn-2.8.3.post1-cuda12.9-torch2.10-cp312-cp312-linux_x86_64.whlbitsandbytes-wheels-for-centos-7-manylinux_2_17_x86_64flash_attn-2.5.5-cp310-cp310-linux_x86_64.whlflashinfer-0.1.6_cu124torch2.4-cp311-cp311-linux_x86_64aligned_aarch64_x86_64axam-linux-x86_64_Debian-dist-v1-05
