CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Compumacy /NMT-openmath OpenMathReasoning DATASET CORRECTION NOTICE We discovered a bug in our data pipeline that caused substantial data loss. The current dataset contains only 290K questions, not the 540K stated in our report. Our OpenMath-Nemotron models were trained with this reduced subset, so all results are reproducible with the currently released version, only the problem count is inaccurate. We're currently fixing this issue and plan to release an updated version next week after verifying the… See the full description on the dataset page: https://huggingface.co/datasets/Compumacy/NMT-openmath.textquestion-answering1M<n<10M0 likes651 downloads1y agoHugging Face02Compumacy /NMT-opencode OpenCodeReasoning: Advancing Data Distillation for Competitive Coding Data Overview OpenCodeReasoning is the largest reasoning-based synthetic dataset to date for coding, comprises 735,255 samples in Python across 28,319 unique competitive programming questions. OpenCodeReasoning is designed for supervised fine-tuning (SFT). Technical Report - Discover the methodology and technical details behind OpenCodeReasoning. Github Repo - Access the complete pipeline used to… See the full description on the dataset page: https://huggingface.co/datasets/Compumacy/NMT-opencode.texttext-generation100K<n<1M0 likes451 downloads1y agoHugging Face03nm-testing /sharegpt_llama3_8b_hidden_statestextn<1K0 likes307 downloads10mo agoHugging Face04DigitalUmuganda /NMT_Rwandan-Gazette_parallel_data_en_kin Dataset Details Dataset Description This is a curated parallel dataset from the Official Gazette of the Republic of Rwanda. It has been curated to extract corresponding English and Kinyarwanda text and in the future we shall add French to the mix Curated by: Digital Umuganda Language(s) (NLP): Kinyarwanda and English License: cc-by-4.0 Dataset Sources [optional] The dataset original content was retrieved from the Rwandan ministry of Justice website… See the full description on the dataset page: https://huggingface.co/datasets/DigitalUmuganda/NMT_Rwandan-Gazette_parallel_data_en_kin.texttranslation100K<n<1M3 likes150 downloads3y agoHugging Face05nyu-dice-lab /lm-eval-results-paulml-DPOB-NMTOB-7B-private Dataset Card for Evaluation run of paulml/DPOB-NMTOB-7B Dataset automatically created during the evaluation run of model paulml/DPOB-NMTOB-7B The dataset is composed of 62 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results. An additional… See the full description on the dataset page: https://huggingface.co/datasets/nyu-dice-lab/lm-eval-results-paulml-DPOB-NMTOB-7B-private.tabular100K<n<1M0 likes126 downloads2y agoHugging Face06LT3 /nfr_bt_nmt_english-french Citation If you use this dataset in your work, please cite accordingly: @article{tezcan_etal_2024_ImprovingFuzzyMatch, title = {Improving {{Fuzzy Match Augmented Neural Machine Translation}} in {{Specialised Domains}} through {{Synthetic Data}}}, author = {Tezcan, Arda and Skidanova, Alina and Moerman, Thomas}, year = {2024}, journal = {The Prague Bulletin of Mathematical Linguistics}, volume = {122}, pages = {9--42}, url =… See the full description on the dataset page: https://huggingface.co/datasets/LT3/nfr_bt_nmt_english-french.text1M<n<10M2 likes98 downloads1y agoHugging Face07amlan107 /chakma-nmt-complete-dataset ChakmaNMT Complete Dataset This dataset accompanies ChakmaNMT: Machine Translation for a Low-Resource and Endangered Language via Transliteration, accepted at WMT 2026. It provides resources for machine translation among Chakma (ccp), Bangla (bn), and English (en). Dataset contents Split Rows Description parallel 15,021 Bangla--Chakma parallel samples; 8,647 also include aligned English. monolingual 150,000 Monolingual data with up to 150,000 Bangla… See the full description on the dataset page: https://huggingface.co/datasets/amlan107/chakma-nmt-complete-dataset.texttranslation100K<n<1M0 likes92 downloads13d agoHugging Face08zouhar /nmt-pe-effects Neural Machine Translation Quality and Post-Editing Performance This is a repository for an experiment relating NMT quality and post-editing efforts, presented at EMNLP2021 (presentation recording). Please cite the following paper when you use this research: @inproceedings{zouhar2021neural, title={Neural Machine Translation Quality and Post-Editing Performance}, author={Zouhar, Vil{\'e}m and Popel, Martin and Bojar, Ond{\v{r}}ej and Tamchyna, Ale{\v{s}}}… See the full description on the dataset page: https://huggingface.co/datasets/zouhar/nmt-pe-effects.tabulartranslation1K<n<10K1 likes88 downloads3y agoHugging Face09wisenut-nlp-team /llama_nmt 중-한 번역 subset: ch-ko_basic_science length: 37.7k subset: ch-ko_broadcast length: 362k subset: ch-ko_daily_colloquial length: 600k subset: ch-ko_food length: 1.2M subset: ch-ko_humanities length: 33.4k subset: ch-ko_utterance_type length: 12k 영-한 번역 subset: en-ko_basic_science length: 356 subset: en-ko_broadcast length: 121k subset: en-ko_daily_colloquial length: 1.2M subset: en-ko_food length: 1.2M subset: en-ko_humanities length:… See the full description on the dataset page: https://huggingface.co/datasets/wisenut-nlp-team/llama_nmt.text10M<n<100M0 likes70 downloads2y agoHugging Face10nyu-dice-lab /lm-eval-results-paulml-NMTOB-7B-private Dataset Card for Evaluation run of paulml/NMTOB-7B Dataset automatically created during the evaluation run of model paulml/NMTOB-7B The dataset is composed of 62 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results. An additional configuration… See the full description on the dataset page: https://huggingface.co/datasets/nyu-dice-lab/lm-eval-results-paulml-NMTOB-7B-private.tabular100K<n<1M0 likes64 downloads2y agoHugging Face11amlan107 /chakma-nmt-base-parallel-dev-settext1K<n<10K0 likes59 downloads2y agoHugging Face12nmthien /vietnamese-curated-1.4mtext1M<n<10M0 likes59 downloads1mo agoHugging Face13Umbaji /NMTMD NMTMD (NMT-Melinda-Dataset) Official repository for the Opensource Text dataset for NMT for local languages in West Africa (EWE Corpus) and implement the Yodi model afterward. Note: This repository will evolve into the official repository for the Yodi model, once the necessary data is gathered. Objective • Develop a Machine Translation Text and Speech Dataset NMT for local languages in West Africa (EWE Corpus) Key Results -> Develop &… See the full description on the dataset page: https://huggingface.co/datasets/Umbaji/NMTMD.textn<1K1 likes58 downloads1y agoHugging Face14Compumacy /NMT-Crossthinking Nemotron-CrossThink: Scaling Self-Learning beyond Math Reasoning Author: Syeda Nahida Akter, Shrimai Prabhumoye, Matvei Novikov, Seungju Han, Ying Lin, Evelina Bakhturina, Eric Nyberg, Yejin Choi, Mostofa Patwary, Mohammad Shoeybi, Bryan Catanzaro [Paper][Blog] Dataset Description Nemotron-CrossThink is a multi-domain reinforcement learning (RL) dataset designed to improve general-purpose and mathematical reasoning in large language models (LLMs). The dataset… See the full description on the dataset page: https://huggingface.co/datasets/Compumacy/NMT-Crossthinking.textquestion-answering10M<n<100M0 likes42 downloads1y agoHugging Face15nmthien /vietnamese-curated-1.4m-v2text1M<n<10M0 likes42 downloads1mo agoHugging Face16spitfire4794 /bangla_nmt_clean text1M<n<10M0 likes38 downloads2mo agoHugging Face17JordanHill-nmtafe /github-raw Overview This data has been scraped from the github api using the requests library! tabular1K<n<10K0 likes38 downloads3d agoHugging Face18Umbaji001 /NMTMD NMTMD (NMT-Melinda-Dataset) Official repository for the Opensource Text dataset for NMT for local languages in West Africa (EWE Corpus) and implement the Yodi model afterward. Note: This repository will evolve into the official repository for the Yodi model, once the necessary data is gathered. Objective • Develop a Machine Translation Text and Speech Dataset NMT for local languages in West Africa (EWE Corpus) Key Results -> Develop &… See the full description on the dataset page: https://huggingface.co/datasets/Umbaji001/NMTMD.textn<1K0 likes37 downloads2y agoHugging Face19HeyDunaX /tay-vietnamese-nmt Tày-Vietnamese Parallel Dataset The Tày–Vietnamese Parallel Dataset is a low-resource bilingual corpus designed for machine translation research. It consists of sentence-level aligned Tày and Vietnamese text pairs, manually curated and validated to ensure semantic accuracy. The dataset supports research on neural machine translation and cross-lingual learning for under-resourced languages. Dataset Statistics Number of sentence pairs: 20,600 Average sentence… See the full description on the dataset page: https://huggingface.co/datasets/HeyDunaX/tay-vietnamese-nmt.texttranslation10K<n<100K1 likes33 downloads6mo agoHugging Face20LT3 /nfr_bt_nmt_english-ukrainian Citation If you use this dataset in your work, please cite accordingly: @article{tezcan_etal_2024_ImprovingFuzzyMatch, title = {Improving {{Fuzzy Match Augmented Neural Machine Translation}} in {{Specialised Domains}} through {{Synthetic Data}}}, author = {Tezcan, Arda and Skidanova, Alina and Moerman, Thomas}, year = {2024}, journal = {The Prague Bulletin of Mathematical Linguistics}, volume = {122}, pages = {9--42}, url =… See the full description on the dataset page: https://huggingface.co/datasets/LT3/nfr_bt_nmt_english-ukrainian.text1M<n<10M6 likes31 downloads1y agoHugging Face21miugod /qianyan_nmt Qianyan Low-Resource NMT Dataset "千言数据集:低资源语言翻译" ,旨在帮助研究人员和开发者解决低资源语言翻译的问题。该数据集包含了中文和俄文的5万条双语平行语料,以及中文和泰文、中文和越南文各10万条目标端单语语料。 对于泰文和越南文,使用谷歌翻译进行回译,从而生成对应的中文数据。 source=1表示中文到其他语言的翻译,source=0表示其他语言到中文的翻译,以便区分测试集的语言方向。 详见: https://aistudio.baidu.com/competition/detail/84/0/introduction texttranslation100K<n<1M3 likes29 downloads3y agoHugging Face22open-llm-leaderboard /EpistemeAI__Reasoning-Llama-3.1-CoT-RE1-NMT-detailsgated Dataset Card for Evaluation run of EpistemeAI/Reasoning-Llama-3.1-CoT-RE1-NMT Dataset automatically created during the evaluation run of model EpistemeAI/Reasoning-Llama-3.1-CoT-RE1-NMT The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/EpistemeAI__Reasoning-Llama-3.1-CoT-RE1-NMT-details.tabular10K<n<100K0 likes29 downloads2y agoHugging Face23liboaccn /nmt-parallel-corpusgated Neural Machine Translation parallel corpora Introduction We use OpusTools to extract resources from the OPUS project, a renowned platform for parallel corpora, and create a multilingual dataset. Specifically, we collect the parallel corpora from prominent projects within OPUS, including NLLB, CCMatrix, and OpenSubtitles. This comprehensive data collection process results in a corpus of more than 3T, covering 60 languages and over 1900 language pairs.… See the full description on the dataset page: https://huggingface.co/datasets/liboaccn/nmt-parallel-corpus.texttranslation10B<n<100B24 likes28 downloads1y agoHugging Face24nm-testing /qa-chat-prompts Dataset Card for "qa-chat-prompts" More Information needed textn<1K2 likes27 downloads2y agoHugging Face25nmthien /vietnamese-curated-2m nmthien/ct219-vietnamese-raw-400k Bộ dữ liệu văn bản tiếng Việt thô đã tiền xử lý, dùng để huấn luyện next-token language model (CT219 - NLP final project). Nguồn dữ liệu Source dataset: VTSNLP/vietnamese_curated_dataset Source split: train Pinned source revision: b81fcce58945970117a1b56d50ec81be2628a5c3 Source licence: không công bố - repo này được tạo ở chế độ private vì lý do đó. Mục đích Huấn luyện next-token language model cho tiếng Việt.… See the full description on the dataset page: https://huggingface.co/datasets/nmthien/vietnamese-curated-2m.text1M<n<10M0 likes26 downloads1mo agoHugging Face26fyaronskiy /ru-paraphrase-NMT-Leipzig-cleaned Dataset Description The dataset is obtained by filtering dataset of russian paraphrases by David Dale with automatic metrics. The data structure is saved. Have been deleted: Paraphrases that have cosine LABSE similarity with source sentences < 0.75. Paraphrases that are more than 2.5 times longer than source sentences. (Most of them are looped errors of back translation) Paraphrases that are similar in spelling to the original texts (paraphrases that have ChrF++ similarity > 0.6… See the full description on the dataset page: https://huggingface.co/datasets/fyaronskiy/ru-paraphrase-NMT-Leipzig-cleaned.tabulartext-generation100K<n<1M2 likes23 downloads1y agoHugging Face27DigitalUmuganda /NMT_Health_parallel_data_en_kintabular10K<n<100K0 likes22 downloads3y agoHugging Face28amlan107 /ccp_nmt_multilingual_train_onlytext100K<n<1M0 likes22 downloads2y agoHugging Face29shchoi1019 /nmttext100K<n<1M0 likes22 downloads2y agoHugging Face30Hastagaras /nmtrn-sft-f-thinktext100K<n<1M0 likes20 downloads8mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.