datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
NMT-openmath
OpenMathReasoning
DATASET CORRECTION NOTICE
We discovered a bug in our data pipeline that caused substantial data loss. The current dataset contains only 290K questions, not the 540K stated in our report.
Our OpenMath-Nemotron models were trained with this reduced subset, so all results are reproducible with the currently released version, only the problem count is inaccurate.
We're currently fixing this issue and plan to release an updated version next week after verifying the… See the full description on the dataset page: https://huggingface.co/datasets/Compumacy/NMT-openmath.NMT-opencode
OpenCodeReasoning: Advancing Data Distillation for Competitive Coding
Data Overview
OpenCodeReasoning is the largest reasoning-based synthetic dataset to date for coding, comprises 735,255 samples in Python across 28,319 unique competitive programming
questions. OpenCodeReasoning is designed for supervised fine-tuning (SFT).
Technical Report - Discover the methodology and technical details behind OpenCodeReasoning.
Github Repo - Access the complete pipeline used to… See the full description on the dataset page: https://huggingface.co/datasets/Compumacy/NMT-opencode.sharegpt_llama3_8b_hidden_statesNMT_Rwandan-Gazette_parallel_data_en_kin
Dataset Details
Dataset Description
This is a curated parallel dataset from the Official Gazette of the Republic of Rwanda. It has been curated to extract corresponding English and Kinyarwanda text and in the future we shall add French to the mix
Curated by: Digital Umuganda
Language(s) (NLP): Kinyarwanda and English
License: cc-by-4.0
Dataset Sources [optional]
The dataset original content was retrieved from the Rwandan ministry of Justice website… See the full description on the dataset page: https://huggingface.co/datasets/DigitalUmuganda/NMT_Rwandan-Gazette_parallel_data_en_kin.lm-eval-results-paulml-DPOB-NMTOB-7B-private
Dataset Card for Evaluation run of paulml/DPOB-NMTOB-7B
Dataset automatically created during the evaluation run of model paulml/DPOB-NMTOB-7B
The dataset is composed of 62 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An additional… See the full description on the dataset page: https://huggingface.co/datasets/nyu-dice-lab/lm-eval-results-paulml-DPOB-NMTOB-7B-private.nfr_bt_nmt_english-french
Citation
If you use this dataset in your work, please cite accordingly:
@article{tezcan_etal_2024_ImprovingFuzzyMatch,
title = {Improving {{Fuzzy Match Augmented Neural Machine Translation}} in {{Specialised Domains}} through {{Synthetic Data}}},
author = {Tezcan, Arda and Skidanova, Alina and Moerman, Thomas},
year = {2024},
journal = {The Prague Bulletin of Mathematical Linguistics},
volume = {122},
pages = {9--42},
url =… See the full description on the dataset page: https://huggingface.co/datasets/LT3/nfr_bt_nmt_english-french.chakma-nmt-complete-dataset
ChakmaNMT Complete Dataset
This dataset accompanies ChakmaNMT: Machine Translation for a Low-Resource and Endangered Language via Transliteration, accepted at WMT 2026. It provides resources for machine translation among Chakma (ccp), Bangla (bn), and English (en).
Dataset contents
Split
Rows
Description
parallel
15,021
Bangla--Chakma parallel samples; 8,647 also include aligned English.
monolingual
150,000
Monolingual data with up to 150,000 Bangla… See the full description on the dataset page: https://huggingface.co/datasets/amlan107/chakma-nmt-complete-dataset.nmt-pe-effects
Neural Machine Translation Quality and Post-Editing Performance
This is a repository for an experiment relating NMT quality and post-editing efforts, presented at EMNLP2021 (presentation recording).
Please cite the following paper when you use this research:
@inproceedings{zouhar2021neural,
title={Neural Machine Translation Quality and Post-Editing Performance},
author={Zouhar, Vil{\'e}m and Popel, Martin and Bojar, Ond{\v{r}}ej and Tamchyna, Ale{\v{s}}}… See the full description on the dataset page: https://huggingface.co/datasets/zouhar/nmt-pe-effects.llama_nmt
중-한 번역
subset: ch-ko_basic_science
length: 37.7k
subset: ch-ko_broadcast
length: 362k
subset: ch-ko_daily_colloquial
length: 600k
subset: ch-ko_food
length: 1.2M
subset: ch-ko_humanities
length: 33.4k
subset: ch-ko_utterance_type
length: 12k
영-한 번역
subset: en-ko_basic_science
length: 356
subset: en-ko_broadcast
length: 121k
subset: en-ko_daily_colloquial
length: 1.2M
subset: en-ko_food
length: 1.2M
subset: en-ko_humanities
length:… See the full description on the dataset page: https://huggingface.co/datasets/wisenut-nlp-team/llama_nmt.lm-eval-results-paulml-NMTOB-7B-private
Dataset Card for Evaluation run of paulml/NMTOB-7B
Dataset automatically created during the evaluation run of model paulml/NMTOB-7B
The dataset is composed of 62 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An additional configuration… See the full description on the dataset page: https://huggingface.co/datasets/nyu-dice-lab/lm-eval-results-paulml-NMTOB-7B-private.chakma-nmt-base-parallel-dev-setvietnamese-curated-1.4mNMTMD
NMTMD (NMT-Melinda-Dataset)
Official repository for the Opensource Text dataset for NMT for local languages in West Africa (EWE Corpus) and implement the Yodi model afterward.
Note: This repository will evolve into the official repository for the Yodi model, once the necessary data is gathered.
Objective
• Develop a Machine Translation Text and Speech Dataset NMT for local languages in West Africa (EWE Corpus)
Key Results
-> Develop &… See the full description on the dataset page: https://huggingface.co/datasets/Umbaji/NMTMD.NMT-Crossthinking
Nemotron-CrossThink: Scaling Self-Learning beyond Math Reasoning
Author: Syeda Nahida Akter, Shrimai Prabhumoye, Matvei Novikov, Seungju Han, Ying Lin, Evelina Bakhturina, Eric Nyberg, Yejin Choi,
Mostofa Patwary, Mohammad Shoeybi, Bryan Catanzaro
[Paper][Blog]
Dataset Description
Nemotron-CrossThink is a multi-domain reinforcement learning (RL) dataset designed to improve general-purpose
and mathematical reasoning in large language models (LLMs).
The dataset… See the full description on the dataset page: https://huggingface.co/datasets/Compumacy/NMT-Crossthinking.vietnamese-curated-1.4m-v2bangla_nmt_clean
github-raw
Overview
This data has been scraped from the github api using the requests library!
NMTMD
NMTMD (NMT-Melinda-Dataset)
Official repository for the Opensource Text dataset for NMT for local languages in West Africa (EWE Corpus) and implement the Yodi model afterward.
Note: This repository will evolve into the official repository for the Yodi model, once the necessary data is gathered.
Objective
• Develop a Machine Translation Text and Speech Dataset NMT for local languages in West Africa (EWE Corpus)
Key Results
-> Develop &… See the full description on the dataset page: https://huggingface.co/datasets/Umbaji001/NMTMD.tay-vietnamese-nmt
Tày-Vietnamese Parallel Dataset
The Tày–Vietnamese Parallel Dataset is a low-resource bilingual corpus designed for machine translation research. It consists of sentence-level aligned Tày and Vietnamese text pairs, manually curated and validated to ensure semantic accuracy. The dataset supports research on neural machine translation and cross-lingual learning for under-resourced languages.
Dataset Statistics
Number of sentence pairs: 20,600
Average sentence… See the full description on the dataset page: https://huggingface.co/datasets/HeyDunaX/tay-vietnamese-nmt.nfr_bt_nmt_english-ukrainian
Citation
If you use this dataset in your work, please cite accordingly:
@article{tezcan_etal_2024_ImprovingFuzzyMatch,
title = {Improving {{Fuzzy Match Augmented Neural Machine Translation}} in {{Specialised Domains}} through {{Synthetic Data}}},
author = {Tezcan, Arda and Skidanova, Alina and Moerman, Thomas},
year = {2024},
journal = {The Prague Bulletin of Mathematical Linguistics},
volume = {122},
pages = {9--42},
url =… See the full description on the dataset page: https://huggingface.co/datasets/LT3/nfr_bt_nmt_english-ukrainian.qianyan_nmt
Qianyan Low-Resource NMT Dataset
"千言数据集:低资源语言翻译" ,旨在帮助研究人员和开发者解决低资源语言翻译的问题。该数据集包含了中文和俄文的5万条双语平行语料,以及中文和泰文、中文和越南文各10万条目标端单语语料。
对于泰文和越南文,使用谷歌翻译进行回译,从而生成对应的中文数据。
source=1表示中文到其他语言的翻译,source=0表示其他语言到中文的翻译,以便区分测试集的语言方向。
详见:
https://aistudio.baidu.com/competition/detail/84/0/introduction
EpistemeAI__Reasoning-Llama-3.1-CoT-RE1-NMT-details
Dataset Card for Evaluation run of EpistemeAI/Reasoning-Llama-3.1-CoT-RE1-NMT
Dataset automatically created during the evaluation run of model EpistemeAI/Reasoning-Llama-3.1-CoT-RE1-NMT
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/EpistemeAI__Reasoning-Llama-3.1-CoT-RE1-NMT-details.nmt-parallel-corpus
Neural Machine Translation parallel corpora
Introduction
We use OpusTools to extract resources from the OPUS project, a renowned platform for parallel corpora, and create a multilingual dataset. Specifically, we collect the parallel corpora from prominent projects within OPUS, including NLLB, CCMatrix, and OpenSubtitles.
This comprehensive data collection process results in a corpus of more than 3T, covering 60 languages and over 1900 language pairs.… See the full description on the dataset page: https://huggingface.co/datasets/liboaccn/nmt-parallel-corpus.qa-chat-prompts
Dataset Card for "qa-chat-prompts"
More Information needed
vietnamese-curated-2m
nmthien/ct219-vietnamese-raw-400k
Bộ dữ liệu văn bản tiếng Việt thô đã tiền xử lý, dùng để huấn luyện
next-token language model (CT219 - NLP final project).
Nguồn dữ liệu
Source dataset: VTSNLP/vietnamese_curated_dataset
Source split: train
Pinned source revision: b81fcce58945970117a1b56d50ec81be2628a5c3
Source licence: không công bố - repo này được tạo ở chế độ private vì lý do đó.
Mục đích
Huấn luyện next-token language model cho tiếng Việt.… See the full description on the dataset page: https://huggingface.co/datasets/nmthien/vietnamese-curated-2m.ru-paraphrase-NMT-Leipzig-cleaned
Dataset Description
The dataset is obtained by filtering dataset of russian paraphrases by David Dale with automatic metrics.
The data structure is saved.
Have been deleted:
Paraphrases that have cosine LABSE similarity with source sentences < 0.75.
Paraphrases that are more than 2.5 times longer than source sentences. (Most of them are looped errors of back translation)
Paraphrases that are similar in spelling to the original texts (paraphrases that have ChrF++ similarity > 0.6… See the full description on the dataset page: https://huggingface.co/datasets/fyaronskiy/ru-paraphrase-NMT-Leipzig-cleaned.NMT_Health_parallel_data_en_kinccp_nmt_multilingual_train_onlynmtnmtrn-sft-f-think
