CoolFace
Datasetpublic

BSC-LT/MULTI_corpus

Dataset Card for MULTI-Corpus Dataset Summary This corpus was compiled as part of the TRAIN project (Traducción Automática para la Inclusión, Automatic Translation for Inclusion), funded by MCIN/AEI and ERDF. It aggregates 15,191,441 sentence-level entries covering four extremely low-resource languages: Tamazight/Amazigh (ZGH), Pashto (PS), Wolof (WO), and Romani (ROM), paired with one or more high-resource counterparts (English, Spanish, French), plus monolingual… See the full description on the dataset page: https://huggingface.co/datasets/BSC-LT/MULTI_corpus.

sourceHugging Facecc-by-sa-4.0updated 4mo agoView on Hugging Face
0likes25downloads
11 commits on main
2e390824mo ago

Update README

fdelucaf
7837fbe5mo ago

Upload license in README

fdelucaf
9b2bb335mo ago

Update LICENSE

fdelucaf
771218c6mo ago

Add LICENSE file

fdelucaf
12e503b6mo ago

Update and split license

fdelucaf
aaa41616mo ago

Upload multi-training-set.json with correct language code for ZGH

fdelucaf
13c4f4e6mo ago

Upload multi-training-set.json

fdelucaf
dc971be6mo ago

Upload multi-corpus-raw.tsv

fdelucaf
583e03c6mo ago

Update README.md

fdelucaf
a351fbb6mo ago

Upload README first draft

fdelucaf
747250d6mo ago

initial commit

fdelucaf