Helsinki-NLP/multiun
Dataset Card for OPUS MultiUN Dataset Summary The MultiUN parallel corpus is extracted from the United Nations Website , and then cleaned and converted to XML at Language Technology Lab in DFKI GmbH (LT-DFKI), Germany. The documents were published by UN from 2000 to 2009. This is a collection of translated documents from the United Nations originally compiled by Andreas Eisele and Yu Chen (see http://www.euromatrixplus.net/multi-un/). This corpus is available in… See the full description on the dataset page: https://huggingface.co/datasets/Helsinki-NLP/multiun.
Convert dataset to Parquet (#8)
Update metadata (#7)
Add dataset name to the title in the dataset card (#6)
Update metadata in dataset card (#5)
Fix OPUS URLs (#4)
Delete legacy JSON metadata (#2)
rename configs to config_name
add dataset_info in dataset metadata
remove dummmy data
Align more metadata with other repo types (models,spaces) (#4607)
Remove config names as yaml keys (#4367)
Update datasets task tags to align tags with models (#4067)
Update files from the datasets library (from 1.18.0)
Update files from the datasets library (from 1.7.0)
Update files from the datasets library (from 1.6.0)
Update files from the datasets library (from 1.3.0)
Update files from the datasets library (from 1.2.0)
