CoolFace
Datasetpublic

CodeForCodersYT/ultimate-code-tokenized

ultimate-code Dataset Description ultimate-code is a derived dataset built by combining and processing data from nvidia/OpenCodeInstruct and nvidia/OpenCodeGeneticInstruct. It can be used to fine-tune LLMs for coding tasks. Tokenized variant available: a pre-tokenized version of this dataset is available at CodeForCodersYT/ultimate-code-tokenized. Use that version if you want ready-to-train tokenized sequences instead of raw text. Source Datasets &… See the full description on the dataset page: https://huggingface.co/datasets/CodeForCodersYT/ultimate-code-tokenized.

sourceHugging Facecc-by-4.0updated 2mo agoView on Hugging Face
0likes155downloads
Dataset Card

ultimate-code

Dataset Description

ultimate-code is a derived dataset built by combining and processing data from `nvidia/OpenCodeInstruct` and `nvidia/OpenCodeGeneticInstruct`. It can be used to fine-tune LLMs for coding tasks.

Tokenized variant available: a pre-tokenized version of this dataset is available at `CodeForCodersYT/ultimate-code-tokenized`. <!-- TODO: Link/Repo-Namen prüfen und ggf. anpassen, falls der Name der tokenisierten Variante anders lautet. --> Use that version if you want ready-to-train tokenized sequences instead of raw text.

Source Datasets & Attribution

This dataset is derived from the following NVIDIA datasets, both released under the Creative Commons Attribution 4.0 International License (CC BY 4.0):

Source DatasetOwnerLicenseLink
OpenCodeInstructNVIDIA CorporationCC BY 4.0https://huggingface.co/datasets/nvidia/OpenCodeInstruct
OpenCodeGeneticInstructNVIDIA CorporationCC BY 4.0https://huggingface.co/datasets/nvidia/OpenCodeGeneticInstruct

Both source datasets are copyright © NVIDIA Corporation and are licensed under CC BY 4.0 (full license text).

Modifications made to the original data

  • Both source datasets were merged into a single unified schema.
  • Only the id, input, and output columns were kept; all other columns from the original datasets (e.g. domain, generation_algorithm, llm_judgement, unit_tests, tests_execution_status, average_test_score) were dropped.
  • No additional filtering, deduplication, or reformatting was applied beyond the column selection.

Dataset Structure

FieldTypeDescription
idstringUnique identifier for each example
inputstringThe coding question / prompt
outputstringThe corresponding solution / response

How to Use

python
from datasets import load_dataset

ds = load_dataset("CodeForCodersYT/ultimate-code", split="train")

For a pre-tokenized version ready for training pipelines, see `CodeForCodersYT/ultimate-code-tokenized` instead.

License

This dataset is released under CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/legalcode), consistent with the license of the source datasets. If you use this dataset, please provide attribution to both this dataset and the original NVIDIA source datasets as described above.

Ethical Considerations

This dataset inherits ethical considerations from its source data. As noted by NVIDIA for the original datasets: users are responsible for checking that the dataset and license are fit for their intended purpose, and should evaluate the resulting models for their specific use case and industry requirements before deployment.

Citation

If you use this dataset, please cite both the original source datasets and this dataset.

OpenCodeInstruct:

bibtex
@article{ahmad2025opencodeinstruct,
      title={OpenCodeInstruct: A Large-scale Instruction Tuning Dataset for Code LLMs},
      author={Wasi Uddin Ahmad and Aleksander Ficek and Mehrzad Samadi and Jocelyn Huang and Vahid Noroozi and Somshubra Majumdar and Boris Ginsburg},
      year={2025},
      eprint={2504.04030},
      archivePrefix={arXiv},
      primaryClass={cs.CL},
      url={https://arxiv.org/abs/2504.04030},
}

OpenCodeGeneticInstruct:

bibtex
@misc{nvidia_opencodegeneticinstruct,
      title={OpenCodeGeneticInstruct},
      author={NVIDIA Corporation},
      year={2025},
      url={https://huggingface.co/datasets/nvidia/OpenCodeGeneticInstruct},
}

This dataset:

bibtex
@misc{ultimate_code_2026,
      title={ultimate-code: A Merged Instruction-Tuning Dataset for Code LLMs},
      author={CodeForCodersYT},
      year={2026},
      url={https://huggingface.co/datasets/CodeForCodersYT/ultimate-code},
}