CodeForCodersYT/ultimate-code-tokenized
ultimate-code Dataset Description ultimate-code is a derived dataset built by combining and processing data from nvidia/OpenCodeInstruct and nvidia/OpenCodeGeneticInstruct. It can be used to fine-tune LLMs for coding tasks. Tokenized variant available: a pre-tokenized version of this dataset is available at CodeForCodersYT/ultimate-code-tokenized. Use that version if you want ready-to-train tokenized sequences instead of raw text. Source Datasets &… See the full description on the dataset page: https://huggingface.co/datasets/CodeForCodersYT/ultimate-code-tokenized.
ultimate-code
Dataset Description
ultimate-code is a derived dataset built by combining and processing data from `nvidia/OpenCodeInstruct` and `nvidia/OpenCodeGeneticInstruct`. It can be used to fine-tune LLMs for coding tasks.
Tokenized variant available: a pre-tokenized version of this dataset is available at `CodeForCodersYT/ultimate-code-tokenized`. <!-- TODO: Link/Repo-Namen prüfen und ggf. anpassen, falls der Name der tokenisierten Variante anders lautet. --> Use that version if you want ready-to-train tokenized sequences instead of raw text.
Source Datasets & Attribution
This dataset is derived from the following NVIDIA datasets, both released under the Creative Commons Attribution 4.0 International License (CC BY 4.0):
Both source datasets are copyright © NVIDIA Corporation and are licensed under CC BY 4.0 (full license text).
Modifications made to the original data
- Both source datasets were merged into a single unified schema.
- Only the
id,input, andoutputcolumns were kept; all other columns from the original datasets (e.g.domain,generation_algorithm,llm_judgement,unit_tests,tests_execution_status,average_test_score) were dropped. - No additional filtering, deduplication, or reformatting was applied beyond the column selection.
Dataset Structure
How to Use
from datasets import load_dataset
ds = load_dataset("CodeForCodersYT/ultimate-code", split="train")For a pre-tokenized version ready for training pipelines, see `CodeForCodersYT/ultimate-code-tokenized` instead.
License
This dataset is released under CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/legalcode), consistent with the license of the source datasets. If you use this dataset, please provide attribution to both this dataset and the original NVIDIA source datasets as described above.
Ethical Considerations
This dataset inherits ethical considerations from its source data. As noted by NVIDIA for the original datasets: users are responsible for checking that the dataset and license are fit for their intended purpose, and should evaluate the resulting models for their specific use case and industry requirements before deployment.
Citation
If you use this dataset, please cite both the original source datasets and this dataset.
OpenCodeInstruct:
@article{ahmad2025opencodeinstruct,
title={OpenCodeInstruct: A Large-scale Instruction Tuning Dataset for Code LLMs},
author={Wasi Uddin Ahmad and Aleksander Ficek and Mehrzad Samadi and Jocelyn Huang and Vahid Noroozi and Somshubra Majumdar and Boris Ginsburg},
year={2025},
eprint={2504.04030},
archivePrefix={arXiv},
primaryClass={cs.CL},
url={https://arxiv.org/abs/2504.04030},
}OpenCodeGeneticInstruct:
@misc{nvidia_opencodegeneticinstruct,
title={OpenCodeGeneticInstruct},
author={NVIDIA Corporation},
year={2025},
url={https://huggingface.co/datasets/nvidia/OpenCodeGeneticInstruct},
}This dataset:
@misc{ultimate_code_2026,
title={ultimate-code: A Merged Instruction-Tuning Dataset for Code LLMs},
author={CodeForCodersYT},
year={2026},
url={https://huggingface.co/datasets/CodeForCodersYT/ultimate-code},
}