CoolFace
Datasetpublic

SultanR/nemotron-mc-en-ar-midtrain

nemotron-mc-en-ar-midtrain Arabic translation of the Nemotron-Pretraining-Multiple-Choice config of Nemotron-Pretraining-Specialized-v1.2 (pinned revision 807afc1). Translated with google/gemma-4-12B-it (bf16, greedy) on A100s. All 23,926,492 source rows are present, none dropped. English source and Arabic translation sit in the same row, so the dataset serves as a parallel corpus as well as an Arabic one. A sibling corpus from the same pipeline is available at… See the full description on the dataset page: https://huggingface.co/datasets/SultanR/nemotron-mc-en-ar-midtrain.

sourceHugging Faceupdated 1mo agoView on Hugging Face
0likes268downloads

SultanR/nemotron-mc-en-ar-midtrain · main · files are served by the source, never re-hosted here