bot-yaya/UPRPRC_pdffiles_from_UN
This datasets contains all the raw PDF file crawl from United Nations Digital Library, produced by https://github.com/mnbvc-parallel-corpus-team/UPRPRC/blob/v2_record_spider/scripts/v4_list2doc.py, using the index in https://huggingface.co/datasets/bot-yaya/documents.un.org_search_result. If you are writing spider script to download all these files, you can do increment download based on this dataset. Our UPRPRC project: https://github.com/mnbvc-parallel-corpus-team/UPRPRC Attention: Record… See the full description on the dataset page: https://huggingface.co/datasets/bot-yaya/UPRPRC_pdffiles_from_UN.
Update README.md
Update README.md
Upload dataset
Upload dataset (part 00009-of-00010)
Upload dataset (part 00008-of-00010)
Upload dataset (part 00007-of-00010)
Upload dataset (part 00006-of-00010)
Upload dataset (part 00005-of-00010)
Upload dataset (part 00004-of-00010)
Upload dataset (part 00003-of-00010)
Upload dataset (part 00002-of-00010)
Upload dataset (part 00001-of-00010)
Upload dataset (part 00000-of-00010)
initial commit
