datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
basque_dialect_machine_translationMachineTranslation_en_viDữ liệu được thu thập từ nhiều nguồn:
CCMatrix: https://opus.nlpl.eu/CCMatrix/en&vi/v1/CCMatrix
OpenSubtitles: https://opus.nlpl.eu/OpenSubtitles/en&vi/v2024/OpenSubtitles
MultiHPLT: https://opus.nlpl.eu/MultiHPLT/en&vi/v2/MultiHPLT
CCAligned: https://opus.nlpl.eu/CCAligned/en&vi/v1/CCAligned
ParaCrawl: https://opus.nlpl.eu/ParaCrawl-Bonus/en&vi/v9/ParaCrawl-Bonus
PhoMT: https://huggingface.co/datasets/ura-hcmut/PhoMTVietAI: https://huggingface.co/datasets/wanhin/VietAI_MTet
Dữ liệu đã trải… See the full description on the dataset page: https://huggingface.co/datasets/Tran1312/MachineTranslation_en_vi.machinetranslationspanishchichewa-machine-translationmenyo_20k_a_multi_domain_english_yoruba_corpus_for_machine_translationinformal_bn-en_machine_translation_datasetreferenceless_machine_translation_evaluationBengali is a low resource language in natural language processing (NLP), with dialects like Sylheti, Chittagong, and Barisal
being even more underrepresented. To address this, ONUBAD introduced a parallel corpus translating these dialects into
Standard Bangla and English using expert translators, providing 1,540 words, 130 clauses, and 980 sentences per dialect.
We focused on the Sylheti-English pair and adapted the dataset for LLM-based machine translation (MT) evaluation.
We extracted the… See the full description on the dataset page: https://huggingface.co/datasets/bokatiq/referenceless_machine_translation_evaluation.
