CoolFace
Datasetpublic

BSC-LT/Spanish-Valencian_Catalan_Parallel_Corpus

Dataset Card for Spanish-Valencian Catalan Parallel Corpus Dataset Summary A bilingual parallel corpus containing parallel sentences in Spanish and the Valencian variant of Catalan. Built by aggregating and filtering multiple public sources, along with data obtained through direct data sharing with external partners, it provides sentence-level alignments for training Machine Translation systems. The dataset includes both authentically parallel data as well as… See the full description on the dataset page: https://huggingface.co/datasets/BSC-LT/Spanish-Valencian_Catalan_Parallel_Corpus.

sourceHugging Facecc-by-nc-4.0updated 7mo agoView on Hugging Face
2likes39downloads
7 commits on main
3e22f477mo ago

Update data sources, acknowledgements, and license

fdelucaf
0ec12cf8mo ago

Upload README.md

fdelucaf
00bf19a8mo ago

Upload .txt files

ellboh
a5af7038mo ago

Upload ca_va-es.parquet

ellboh
a8df1468mo ago

Delete vl-es.parquet

ellboh
09457aa8mo ago

Upload vl-es.parquet

ellboh
e6cf7638mo ago

initial commit

fdelucaf