CoolFace
Datasetpublic

soarescmsa/capes

Dataset Card for CAPES Dataset Summary A parallel corpus of theses and dissertations abstracts in English and Portuguese were collected from the CAPES website (Coordenação de Aperfeiçoamento de Pessoal de Nível Superior) - Brazil. The corpus is sentence aligned for all language pairs. Approximately 240,000 documents were collected and aligned using the Hunalign algorithm. Supported Tasks and Leaderboards The underlying task is machine translation.… See the full description on the dataset page: https://huggingface.co/datasets/soarescmsa/capes.

sourceHugging Faceunknownupdated 3y agoView on Hugging Face
2likes241downloads

soarescmsa/capes · main · files are served by the source, never re-hosted here