bot-yaya/UPRPRC_docfiles_from_UN
This datasets contains all the raw DOC file crawl from United Nations Digital Library, produced by https://github.com/mnbvc-parallel-corpus-team/UPRPRC/blob/v2_record_spider/scripts/v4_list2doc.py, using the index in https://huggingface.co/datasets/bot-yaya/documents.un.org_search_result. If you are writing spider script to download all these files, you can do increment download based on this dataset. Our UPRPRC project: https://github.com/mnbvc-parallel-corpus-team/UPRPRC Attention: Record… See the full description on the dataset page: https://huggingface.co/datasets/bot-yaya/UPRPRC_docfiles_from_UN.
This datasets contains all the raw DOC file crawl from United Nations Digital Library, produced by https://github.com/mnbvc-parallel-corpus-team/UPRPRC/blob/v2recordspider/scripts/v4list2doc.py, using the index in https://huggingface.co/datasets/bot-yaya/documents.un.orgsearch_result.
If you are writing spider script to download all these files, you can do increment download based on this dataset.
Our UPRPRC project: https://github.com/mnbvc-parallel-corpus-team/UPRPRC
Attention: Record which contains <= 1 language's version are not included. As they cannot format parallel corpus.
This dataset is part of the project MNBVC: https://huggingface.co/liwu/MNBVC
