CoolFace
20 results

nmt

Compumacy /NMT-openmath OpenMathReasoning DATASET CORRECTION NOTICE We discovered a bug in our data pipeline that caused substantial data loss. The current dataset contains only 290K questions, not the 540K stated in our report. Our OpenMath-Nemotron models were trained with this reduced subset, so all results are reproducible with the currently released version, only the problem count is inaccurate. We're currently fixing this issue and plan to release an updated version next week after verifying the… See the full description on the dataset page: https://huggingface.co/datasets/Compumacy/NMT-openmath.textquestion-answering1M<n<10M0 likes651 downloads1y agoHugging FaceCompumacy /NMT-opencode OpenCodeReasoning: Advancing Data Distillation for Competitive Coding Data Overview OpenCodeReasoning is the largest reasoning-based synthetic dataset to date for coding, comprises 735,255 samples in Python across 28,319 unique competitive programming questions. OpenCodeReasoning is designed for supervised fine-tuning (SFT). Technical Report - Discover the methodology and technical details behind OpenCodeReasoning. Github Repo - Access the complete pipeline used to… See the full description on the dataset page: https://huggingface.co/datasets/Compumacy/NMT-opencode.texttext-generation100K<n<1M0 likes451 downloads1y agoHugging Facecointegrated /ru-paraphrase-NMT-Leipzig Dataset Card for cointegrated/ru-paraphrase-NMT-Leipzig Dataset Summary The dataset contains 1 million Russian sentences and their automatically generated paraphrases. It was created by David Dale (@cointegrated) by translating the rus-ru_web-public_2019_1M corpus from the Leipzig collection into English and back into Russian. A fraction of the resulting paraphrases are invalid, and should be filtered out. The blogpost "Перефразирование русских текстов: корпуса, модели… See the full description on the dataset page: https://huggingface.co/datasets/cointegrated/ru-paraphrase-NMT-Leipzig.text-generation100K<n<1M12 likes425 downloads4y agoHugging Facenm-testing /sharegpt_llama3_8b_hidden_statestextn<1K0 likes307 downloads10mo agoHugging FaceDigitalUmuganda /NMT_Rwandan-Gazette_parallel_data_en_kin Dataset Details Dataset Description This is a curated parallel dataset from the Official Gazette of the Republic of Rwanda. It has been curated to extract corresponding English and Kinyarwanda text and in the future we shall add French to the mix Curated by: Digital Umuganda Language(s) (NLP): Kinyarwanda and English License: cc-by-4.0 Dataset Sources [optional] The dataset original content was retrieved from the Rwandan ministry of Justice website… See the full description on the dataset page: https://huggingface.co/datasets/DigitalUmuganda/NMT_Rwandan-Gazette_parallel_data_en_kin.texttranslation100K<n<1M3 likes150 downloads3y agoHugging Facenyu-dice-lab /lm-eval-results-paulml-DPOB-NMTOB-7B-private Dataset Card for Evaluation run of paulml/DPOB-NMTOB-7B Dataset automatically created during the evaluation run of model paulml/DPOB-NMTOB-7B The dataset is composed of 62 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results. An additional… See the full description on the dataset page: https://huggingface.co/datasets/nyu-dice-lab/lm-eval-results-paulml-DPOB-NMTOB-7B-private.tabular100K<n<1M0 likes126 downloads2y agoHugging Face