CoolFace
Datasetpublic

BSC-LT/Spanish-Valencian_Catalan_Parallel_Corpus

Dataset Card for Spanish-Valencian Catalan Parallel Corpus Dataset Summary A bilingual parallel corpus containing parallel sentences in Spanish and the Valencian variant of Catalan. Built by aggregating and filtering multiple public sources, along with data obtained through direct data sharing with external partners, it provides sentence-level alignments for training Machine Translation systems. The dataset includes both authentically parallel data as well as… See the full description on the dataset page: https://huggingface.co/datasets/BSC-LT/Spanish-Valencian_Catalan_Parallel_Corpus.

sourceHugging Facecc-by-nc-4.0updated 7mo agoView on Hugging Face
2likes39downloads
fileca_va-es.parquet443.7 MBdownload

BSC-LT/Spanish-Valencian_Catalan_Parallel_Corpus · main · files are served by the source, never re-hosted here