CoolFace
Datasetpublic

pdelobelle/nemotron-dutch-mt

Nemotron Post-Training Dataset (Dutch Translation) Machine-translated Dutch version of NVIDIA's Nemotron Post-Training Dataset, specifically the chat conversations. Dataset Details Source: nvidia/Nemotron-Post-Training-Dataset-v2 (chat split) Translation: English → Dutch using Unbabel/Tower-Plus-9B Size: 445,287 conversations with 1,327,548 total messages Format: Conversational data with original structure preserved Dataset Statistics Total… See the full description on the dataset page: https://huggingface.co/datasets/pdelobelle/nemotron-dutch-mt.

sourceHugging Faceodc-byupdated 1y agoView on Hugging Face
0likes80downloads
3 commits on main
ecbe9fa1y ago

Update README.md

pdelobelle
d4b6e7a1y ago

Upload reconstructed translated conversations with fixed translation artifacts

pdelobelle
e07b0bf1y ago

initial commit

pdelobelle