CoolFace
Datasetpublic

bot-yaya/UPRPRC_docfiles_from_UN

This datasets contains all the raw DOC file crawl from United Nations Digital Library, produced by https://github.com/mnbvc-parallel-corpus-team/UPRPRC/blob/v2_record_spider/scripts/v4_list2doc.py, using the index in https://huggingface.co/datasets/bot-yaya/documents.un.org_search_result. If you are writing spider script to download all these files, you can do increment download based on this dataset. Our UPRPRC project: https://github.com/mnbvc-parallel-corpus-team/UPRPRC Attention: Record… See the full description on the dataset page: https://huggingface.co/datasets/bot-yaya/UPRPRC_docfiles_from_UN.

sourceHugging Facemitupdated 10mo agoView on Hugging Face
0likes267downloads
Dataset Card

This datasets contains all the raw DOC file crawl from United Nations Digital Library, produced by https://github.com/mnbvc-parallel-corpus-team/UPRPRC/blob/v2recordspider/scripts/v4list2doc.py, using the index in https://huggingface.co/datasets/bot-yaya/documents.un.orgsearch_result.

If you are writing spider script to download all these files, you can do increment download based on this dataset.

Our UPRPRC project: https://github.com/mnbvc-parallel-corpus-team/UPRPRC

Attention: Record which contains <= 1 language's version are not included. As they cannot format parallel corpus.

This dataset is part of the project MNBVC: https://huggingface.co/liwu/MNBVC