CoolFace
Datasetpublic

BSC-LT/Catalan-Aranese_Parallel_Corpus

Dataset Card for Catalan-Aranese Parallel Corpus Dataset Summary A bilingual parallel corpus for the low-resource language pair Catalan-Aranese. Built by aggregating and filtering multiple public sources, along with data obtained through direct data sharing with external partners, it provides sentence-level alignments for training Machine Translation systems. The dataset includes both authentically parallel data as well as synthetic Catalan translations generated… See the full description on the dataset page: https://huggingface.co/datasets/BSC-LT/Catalan-Aranese_Parallel_Corpus.

sourceHugging Facecc-by-4.0updated 8mo agoView on Hugging Face
2likes84downloads
discussions and pull requests

Conversations for this repository live on Hugging Face.

CoolFace shows imported repositories read-only. Posting into someone else’s repository from here would need an authorised integration and the account holder’s consent, so the link goes to the source instead.

Open discussions on Hugging Face
BSC-LT/Catalan-Aranese_Parallel_Corpus · CoolFace