datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Low-resource-QE-DA-dataset
Low-resource QE-DA Dataset
Direct Assessment (DA) quality estimation data for English→Indic (Gujarati, Hindi, Marathi, Tamil, Telugu) and related Estonian/Nepali/Sinhala pairs, released with the ALOPE work on LLM-based QE.
Paper: Sindhujan, A., Qian, S., Matthew, C.C.C., Orasan, C., and Kanojia, D. (2024). ALOPE: Adaptive Layer Optimization for Translation Quality Estimation using Large Language Models. In Second Conference on Language Modeling. (arXiv)
Task: Sentence-level quality… See the full description on the dataset page: https://huggingface.co/datasets/surrey-nlp/Low-resource-QE-DA-dataset.low_resource_parallel_corpora
The Little Prince — multiparallel corpus (22 languages, RU pivot)
A sentence-level multiparallel corpus of Antoine de Saint-Exupéry's The Little Prince,
built around the classic Russian translation by Nora Gal as the pivot and covering
21 further editions, most of them in low-resource minority languages of Russia.
Every one of the 1,565 pivot sentences has exactly one aligned sentence in every
included language — a perfect N-way alignment (no gaps, no merges). The editions were… See the full description on the dataset page: https://huggingface.co/datasets/averoo/low_resource_parallel_corpora.low-resource-audio-text
