datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
NLG-Machine-Translation
SEA Machine Translation
SEA Machine Translation evaluates a model's ability to translate a document from a source language into a target language coherently and fluently. It is sampled from FLORES 200 for Burmese, Chinese, English, Indonesian, Khmer, Malay, Tamil, Thai, and Vietnamese, and NusaX for Indonesian, Javanese, and Sundanese.
Supported Tasks and Leaderboards
SEA Machine Translation is designed for evaluating chat or instruction-tuned large language models… See the full description on the dataset page: https://huggingface.co/datasets/aisingapore/NLG-Machine-Translation.Weblate-Translations
Dataset Card for Weblate Translations
A dataset containing strings from projects hosted on Weblate and their translations into other languages.
Please consider donating or contributing to Weblate if you find this dataset useful.
Dataset Details
Dataset Description
Curated by: Mohamed Aymane Farhi
Funded by [optional]: [More Information Needed]
Shared by [optional]: [More Information Needed]
Language(s) (NLP): Check the README YAML metadata… See the full description on the dataset page: https://huggingface.co/datasets/Tamazight-NLP/Weblate-Translations.last-translation-benchmark
Last Translation Benchmark
Abstract: For scientific progress, we need benchmarks that test the limits of state-of-the-art models, and evaluation methods that inform us about failure cases.
Standard benchmarks for machine translation evaluation are often either trivial (having few authentic mistakes) or unrealistic (overly synthetically contrived).
Furthermore, automatic translation metrics become less reliable and reward-hacked as models get stronger, and their outputs are… See the full description on the dataset page: https://huggingface.co/datasets/zouhar/last-translation-benchmark.instruction_translationsTranslation of Instruction datasetEuroparl-Translation-Instruct
Dataset Card for Europarl-Translation-Instruct
Waifu to catch your attention.
Dataset Details
Dataset Description
europarl-translation-instruct is a translation instruct dataset built from europarl data.
Curated by: M8than
Funded by: Recursal.ai
Shared by: M8than
Language(s) (NLP): English instruct (but various languages in)
License: cc-by-sa-4.0
Dataset Sources
Source Data: https://www.statmt.org/europarl/ (Transcript source)
Processing… See the full description on the dataset page: https://huggingface.co/datasets/recursal/Europarl-Translation-Instruct.Translation-Instruct
Corpus overview
Translation Instruct is a collection of parallel corpora for machine translation, formatted as instructions for the supervised fine-tuning of large language models. It currently contains two collections: Croissant Aligned Instruct and Europarl Aligned Instruct.
Croissant Aligned Instruct is an instruction-formatted version of the parallel French-English data in croissantllm/croissant_dataset_no_web_data
(subset: aligned_36b).
Europarl Aligned Instruct is an… See the full description on the dataset page: https://huggingface.co/datasets/OpenLLM-France/Translation-Instruct.task1233_ted_translation_ar_he
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1233_ted_translation_ar_he
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1233_ted_translation_ar_he.task1396_europa_ecdc_tm_en_de_translation
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1396_europa_ecdc_tm_en_de_translation
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1396_europa_ecdc_tm_en_de_translation.task661_mizan_en_fa_translation
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task661_mizan_en_fa_translation
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task661_mizan_en_fa_translation.task1103_ted_translation_es_fa
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1103_ted_translation_es_fa
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1103_ted_translation_es_fa.task982_pib_translation_tamil_bengali
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task982_pib_translation_tamil_bengali
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task982_pib_translation_tamil_bengali.task1009_pib_translation_bengali_hindi
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1009_pib_translation_bengali_hindi
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1009_pib_translation_bengali_hindi.task1101_ted_translation_es_it
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1101_ted_translation_es_it
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1101_ted_translation_es_it.task1274_ted_translation_pt_en
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1274_ted_translation_pt_en
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1274_ted_translation_pt_en.huggingface_myanmar_english_translation
Cleaned & Sorted Myanmar-English Translation Dataset
This dataset is a cleaned, Unicode-normalized, and sorted version of the Myanmar (Burmese) subset from the massive FineTranslations dataset.
While the original dataset is excellent, Myanmar text on the web is often a mix of standard Unicode and the non-standard Zawgyi encoding. This repository fixes those encoding issues to provide a high-quality dataset for NLP tasks.
Key Improvements in this Version
Zawgyi… See the full description on the dataset page: https://huggingface.co/datasets/freococo/huggingface_myanmar_english_translation.task1265_ted_translation_fa_en
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1265_ted_translation_fa_en
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1265_ted_translation_fa_en.task1086_pib_translation_marathi_english
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1086_pib_translation_marathi_english
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1086_pib_translation_marathi_english.task1690_qed_amara_translation
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1690_qed_amara_translation
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1690_qed_amara_translation.task1098_ted_translation_ja_fa
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1098_ted_translation_ja_fa
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1098_ted_translation_ja_fa.task1251_ted_translation_it_he
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1251_ted_translation_it_he
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1251_ted_translation_it_he.task1371_newscomm_translation
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1371_newscomm_translation
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1371_newscomm_translation.task544_alt_translation_hi_en
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task544_alt_translation_hi_en
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task544_alt_translation_hi_en.task558_alt_translation_en_hi
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task558_alt_translation_en_hi
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task558_alt_translation_en_hi.task1035_pib_translation_tamil_urdu
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1035_pib_translation_tamil_urdu
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1035_pib_translation_tamil_urdu.task1650_opus_books_en-fi_translation
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1650_opus_books_en-fi_translation
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1650_opus_books_en-fi_translation.task1105_ted_translation_ar_gl
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1105_ted_translation_ar_gl
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1105_ted_translation_ar_gl.task1243_ted_translation_gl_it
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1243_ted_translation_gl_it
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1243_ted_translation_gl_it.task1115_alt_ja_id_translation
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1115_alt_ja_id_translation
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1115_alt_ja_id_translation.task1095_ted_translation_ja_gl
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1095_ted_translation_ja_gl
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1095_ted_translation_ja_gl.task452_opus_paracrawl_en_ig_translation
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task452_opus_paracrawl_en_ig_translation
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task452_opus_paracrawl_en_ig_translation.
