datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
sharegpt_llama3_8b_hidden_stateslm-eval-results-paulml-DPOB-NMTOB-7B-private
Dataset Card for Evaluation run of paulml/DPOB-NMTOB-7B
Dataset automatically created during the evaluation run of model paulml/DPOB-NMTOB-7B
The dataset is composed of 62 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An additional… See the full description on the dataset page: https://huggingface.co/datasets/nyu-dice-lab/lm-eval-results-paulml-DPOB-NMTOB-7B-private.nmt-pe-effects
Neural Machine Translation Quality and Post-Editing Performance
This is a repository for an experiment relating NMT quality and post-editing efforts, presented at EMNLP2021 (presentation recording).
Please cite the following paper when you use this research:
@inproceedings{zouhar2021neural,
title={Neural Machine Translation Quality and Post-Editing Performance},
author={Zouhar, Vil{\'e}m and Popel, Martin and Bojar, Ond{\v{r}}ej and Tamchyna, Ale{\v{s}}}… See the full description on the dataset page: https://huggingface.co/datasets/zouhar/nmt-pe-effects.lm-eval-results-paulml-NMTOB-7B-private
Dataset Card for Evaluation run of paulml/NMTOB-7B
Dataset automatically created during the evaluation run of model paulml/NMTOB-7B
The dataset is composed of 62 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An additional configuration… See the full description on the dataset page: https://huggingface.co/datasets/nyu-dice-lab/lm-eval-results-paulml-NMTOB-7B-private.qianyan_nmt
Qianyan Low-Resource NMT Dataset
"千言数据集:低资源语言翻译" ,旨在帮助研究人员和开发者解决低资源语言翻译的问题。该数据集包含了中文和俄文的5万条双语平行语料,以及中文和泰文、中文和越南文各10万条目标端单语语料。
对于泰文和越南文,使用谷歌翻译进行回译,从而生成对应的中文数据。
source=1表示中文到其他语言的翻译,source=0表示其他语言到中文的翻译,以便区分测试集的语言方向。
详见:
https://aistudio.baidu.com/competition/detail/84/0/introduction
EpistemeAI__Reasoning-Llama-3.1-CoT-RE1-NMT-details
Dataset Card for Evaluation run of EpistemeAI/Reasoning-Llama-3.1-CoT-RE1-NMT
Dataset automatically created during the evaluation run of model EpistemeAI/Reasoning-Llama-3.1-CoT-RE1-NMT
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/EpistemeAI__Reasoning-Llama-3.1-CoT-RE1-NMT-details.MemGuide
MemGuide
MemGuide is the training dataset for MemAdapter,
a plug-and-play module that converts retrieved memories into structured planning guidance for
Vision-Language Model (VLM) embodied agents.
Dataset Summary
Each record pairs a task instruction and retrieved memories (spatial, episodic, and semantic)
with structured planning guidance produced by a frontier LLM and filtered by behavioral consensus.
The dataset is used to fine-tune the MemAdapter (Qwen3-14B +… See the full description on the dataset page: https://huggingface.co/datasets/NMThuan032k/MemGuide.EpistemeAI__Reasoning-Llama-3.1-CoT-RE1-NMT-V2-ORPO-details
Dataset Card for Evaluation run of EpistemeAI/Reasoning-Llama-3.1-CoT-RE1-NMT-V2-ORPO
Dataset automatically created during the evaluation run of model EpistemeAI/Reasoning-Llama-3.1-CoT-RE1-NMT-V2-ORPO
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/EpistemeAI__Reasoning-Llama-3.1-CoT-RE1-NMT-V2-ORPO-details.NMToxificationParallel corpora for toxicity analysis. Toxic prompts were taken from Toxigen datset and "translated" from toxic to not-toxic with Llama-3-70B.
