datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
flores_101One of the biggest challenges hindering progress in low-resource and multilingual machine translation is the
lack of good evaluation benchmarks. Current evaluation benchmarks either lack good coverage of low-resource
languages, consider only restricted domains, or are low quality because they are constructed using
semi-automatic procedures. In this work, we introduce the FLORES evaluation benchmark, consisting of 3001
sentences extracted from English Wikipedia and covering a variety of different topics and domains.
These sentences have been translated in 101 languages by professional translators through a carefully
controlled process. The resulting dataset enables better assessment of model quality on the long tail of
low-resource languages, including the evaluation of many-to-many multilingual translation systems, as all
translations are multilingually aligned. By publicly releasing such a high-quality and high-coverage dataset,
we hope to foster progress in the machine translation community and beyond.flores_plus
Dataset Card for FLORES+
FLORES+ is an evaluation benchmark dataset for multilingual machine translation.
Dataset Details
Dataset Description
FLORES+ is a multilingual machine translation benchmark released under CC BY-SA 4.0. This dataset was originally released by FAIR researchers at Meta under the name FLORES. Further information about these initial releases can be found in Dataset Sources below. The data is now being managed by OLDI, the Open… See the full description on the dataset page: https://huggingface.co/datasets/openlanguagedata/flores_plus.flores
Dataset Card for Flores 200
Dataset Summary
⚠️ This repository is no longer being updated ⚠️
A newer version of the FLORES dataset managed by the Open Language Data Initiative
is available at https://huggingface.co/datasets/openlanguagedata/flores_plus.
FLORES is a benchmark dataset for machine translation between English and low-resource languages.
The creation of FLORES-200 doubles the existing language coverage of FLORES-101.
Given the nature of the new… See the full description on the dataset page: https://huggingface.co/datasets/facebook/flores.flores_101One of the biggest challenges hindering progress in low-resource and multilingual machine translation is the
lack of good evaluation benchmarks. Current evaluation benchmarks either lack good coverage of low-resource
languages, consider only restricted domains, or are low quality because they are constructed using
semi-automatic procedures. In this work, we introduce the FLORES evaluation benchmark, consisting of 3001
sentences extracted from English Wikipedia and covering a variety of different topics and domains.
These sentences have been translated in 101 languages by professional translators through a carefully
controlled process. The resulting dataset enables better assessment of model quality on the long tail of
low-resource languages, including the evaluation of many-to-many multilingual translation systems, as all
translations are multilingually aligned. By publicly releasing such a high-quality and high-coverage dataset,
we hope to foster progress in the machine translation community and beyond.smugri-flores-testsetMultilingual FLORES-based benchmark for Komi, Udmurt, Hill and Meadow Mari, Erzya, Moksha, Livonian, Mansi, and Livvi Karelian. Expanded with Proper Karelian, Ludian, and Veps.
Please, cite the following paper if you use Komi, Udmurt, Hill and Meadow Mari, Erzya, Moksha, Livonian, Mansi, and Livvi Karelian datasets:
@inproceedings{
yankovskaya2023machine,
title={Machine Translation for Low-resource Finno-Ugric Languages},
author={Lisa Yankovskaya and Maali Tars and Andre T{\"a}ttar and Mark… See the full description on the dataset page: https://huggingface.co/datasets/tartuNLP/smugri-flores-testset.task1514_flores_translation_entone
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1514_flores_translation_entone
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1514_flores_translation_entone.QuixiAI-dolphin_DatasetDolphin 🐬
https://erichartford.com/dolphin
Dataset details
This dataset is an attempt to replicate the results of Microsoft's Orca
Our dataset consists of:
~1 million of FLANv2 augmented with GPT-4 completions (flan1m-alpaca-uncensored.jsonl)
~3.5 million of FLANv2 augmented with GPT-3.5 completions (flan5m-alpaca-uncensored.jsonl)
We followed the submix and system prompt distribution outlined in the Orca paper. With a few exceptions. We included all 75k of CoT in the FLAN-1m… See the full description on the dataset page: https://huggingface.co/datasets/Maximiliano-Flores-Dev/QuixiAI-dolphin_Dataset.agentlans-combined-roleplay_Dataset
Combined Roleplay Dataset
This dataset combines multi-turn conversations across various AI assistant interactions, creative writing scenarios, and roleplaying exchanges. It aims to improve language models' performance in interactive tasks.
Multi-turn conversations with a mix of standard AI assistant interactions, creative writing prompts, and roleplays
English content with a few Spanish, Portuguese, and Chinese conversations
Conversations limited to 4000 tokens using the Llama… See the full description on the dataset page: https://huggingface.co/datasets/Maximiliano-Flores-Dev/agentlans-combined-roleplay_Dataset.cloudbjorn-eschaton-uncensored_Dataset
Eschaton Uncensored SFT Dataset
Eschaton is a 1,000-row supervised fine-tuning dataset for teaching language models to answer sensitive, gritty, controversial, emotionally intimate, technically demanding, and morally difficult prompts directly. Its emphasis is human connection, unvarnished analysis, malware analysis, dark creative work, profanity, and high-friction subject matter without repetitive moralizing or canned disclaimers.
The dataset is deliberately model-agnostic. It… See the full description on the dataset page: https://huggingface.co/datasets/Maximiliano-Flores-Dev/cloudbjorn-eschaton-uncensored_Dataset.task601_flores_translation_sntoen
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task601_flores_translation_sntoen
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task601_flores_translation_sntoen.llimba-flores-srd-eval
LLiMba FLORES-200 Sardinian Evaluation Set
A held-out evaluation set of 997 parallel sentences aligned across six languages (Sardinian, Italian, English, Spanish, French, Portuguese), derived from FLORES-200. Used to benchmark the LLiMba model's translation quality and reported in the LLiMba paper's BLEU and chrF tables.
This is the exact evaluation set used to produce the published LLiMba translation results. Reproducing the paper's numbers requires this set plus… See the full description on the dataset page: https://huggingface.co/datasets/lballore/llimba-flores-srd-eval.flores_101One of the biggest challenges hindering progress in low-resource and multilingual machine translation is the
lack of good evaluation benchmarks. Current evaluation benchmarks either lack good coverage of low-resource
languages, consider only restricted domains, or are low quality because they are constructed using
semi-automatic procedures. In this work, we introduce the FLORES evaluation benchmark, consisting of 3001
sentences extracted from English Wikipedia and covering a variety of different topics and domains.
These sentences have been translated in 101 languages by professional translators through a carefully
controlled process. The resulting dataset enables better assessment of model quality on the long tail of
low-resource languages, including the evaluation of many-to-many multilingual translation systems, as all
translations are multilingually aligned. By publicly releasing such a high-quality and high-coverage dataset,
we hope to foster progress in the machine translation community and beyond.task604_flores_translation_entosn
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task604_flores_translation_entosn
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task604_flores_translation_entosn.flores_101One of the biggest challenges hindering progress in low-resource and multilingual machine translation is the
lack of good evaluation benchmarks. Current evaluation benchmarks either lack good coverage of low-resource
languages, consider only restricted domains, or are low quality because they are constructed using
semi-automatic procedures. In this work, we introduce the FLORES evaluation benchmark, consisting of 3001
sentences extracted from English Wikipedia and covering a variety of different topics and domains.
These sentences have been translated in 101 languages by professional translators through a carefully
controlled process. The resulting dataset enables better assessment of model quality on the long tail of
low-resource languages, including the evaluation of many-to-many multilingual translation systems, as all
translations are multilingually aligned. By publicly releasing such a high-quality and high-coverage dataset,
we hope to foster progress in the machine translation community and beyond.meta-llama-llama-3-1-70b-instruct__llm-mt-flores-200-mini-en-fr__019e3b92df93
meta-llama/Llama-3.1-70B-Instruct on llm.mt.flores-200-mini-en-fr (NVIDIA H100 80GB HBM3)
Back to leaderboard
Headline metrics
Metric
Value
Unit
N Samples
8
N Ok
8
Ok Rate
1
Chrf Mean
0.8819
Chrf P50
0.9033
Chrf P95
1
TTFT P50
32.0341
ms
Total P50 Ms
296.1723
Tokens Out Total
115
Run configuration
Model: meta-llama/Llama-3.1-70B-Instruct @ unknown00
Engine: vllm vunknown
Quantization: fp16
Hardware: NVIDIA H100 80GB… See the full description on the dataset page: https://huggingface.co/datasets/Yobitel/meta-llama-llama-3-1-70b-instruct__llm-mt-flores-200-mini-en-fr__019e3b92df93.mistralai-mistral-7b-instruct-v0-3__llm-mt-flores-200-mini-en-fr__019e3b460885
mistralai/Mistral-7B-Instruct-v0.3 on llm.mt.flores-200-mini-en-fr (NVIDIA H100 80GB HBM3)
Back to leaderboard
Headline metrics
Metric
Value
Unit
N Samples
8
N Ok
8
Ok Rate
1
Chrf Mean
0.767
Chrf P50
0.7445
Chrf P95
1
TTFT P50
11.9969
ms
Total P50 Ms
154.8732
Tokens Out Total
181
Run configuration
Model: mistralai/Mistral-7B-Instruct-v0.3 @ unknown00
Engine: vllm vunknown
Quantization: fp16
Hardware: NVIDIA H100 80GB… See the full description on the dataset page: https://huggingface.co/datasets/Yobitel/mistralai-mistral-7b-instruct-v0-3__llm-mt-flores-200-mini-en-fr__019e3b460885.qwen-qwen2-5-7b-instruct__llm-mt-flores-200-mini-en-fr__019e3b39c5bb
Qwen/Qwen2.5-7B-Instruct on llm.mt.flores-200-mini-en-fr (NVIDIA H100 80GB HBM3)
Back to leaderboard
Headline metrics
Metric
Value
Unit
N Samples
8
N Ok
8
Ok Rate
1
Chrf Mean
0.8912
Chrf P50
0.9439
Chrf P95
1
TTFT P50
14.1119
ms
Total P50 Ms
115.5118
Tokens Out Total
114
Run configuration
Model: Qwen/Qwen2.5-7B-Instruct @ unknown00
Engine: vllm vunknown
Quantization: fp16
Hardware: NVIDIA H100 80GB HBM3
Driver:… See the full description on the dataset page: https://huggingface.co/datasets/Yobitel/qwen-qwen2-5-7b-instruct__llm-mt-flores-200-mini-en-fr__019e3b39c5bb.google-gemma-2-9b-it__llm-mt-flores-200-mini-en-fr__019e3b9789b8
google/gemma-2-9b-it on llm.mt.flores-200-mini-en-fr (NVIDIA H100 80GB HBM3)
Back to leaderboard
Headline metrics
Metric
Value
Unit
N Samples
8
N Ok
8
Ok Rate
1
Chrf Mean
0.7462
Chrf P50
0.6872
Chrf P95
1
TTFT P50
16.7669
ms
Total P50 Ms
291.575
Tokens Out Total
199
Run configuration
Model: google/gemma-2-9b-it @ unknown00
Engine: vllm vunknown
Quantization: fp16
Hardware: NVIDIA H100 80GB HBM3
Driver: 580.126.09… See the full description on the dataset page: https://huggingface.co/datasets/Yobitel/google-gemma-2-9b-it__llm-mt-flores-200-mini-en-fr__019e3b9789b8.meta-llama-llama-3-1-8b-instruct__llm-mt-flores-200-mini-en-fr__019e3b307ea3
meta-llama/Llama-3.1-8B-Instruct on llm.mt.flores-200-mini-en-fr (NVIDIA H100 80GB HBM3)
Back to leaderboard
Headline metrics
Metric
Value
Unit
N Samples
8
N Ok
8
Ok Rate
1
Chrf Mean
0.7749
Chrf P50
0.7614
Chrf P95
0.96
TTFT P50
14.1551
ms
Total P50 Ms
140.4766
Tokens Out Total
151
Run configuration
Model: meta-llama/Llama-3.1-8B-Instruct @ unknown00
Engine: vllm vunknown
Quantization: fp16
Hardware: NVIDIA H100 80GB… See the full description on the dataset page: https://huggingface.co/datasets/Yobitel/meta-llama-llama-3-1-8b-instruct__llm-mt-flores-200-mini-en-fr__019e3b307ea3.
