CoolFace
Datasetpublic

mbshr/XSUMUrdu-DW_BBC

Urdu_DW-BBC-512 Dataset Summary Urdu Summarization Dataset containining 76,637 records of Article + Summary pairs scrapped from BBC Urdu and DW Urdu News Websites. Preprocessed Version: upto 512 tokens (~words); removed URLs, Pic Captions etc Supported Tasks and Leaderboards Summarization: Extractive and Abstractive urT5 adapted from mT5 having monolingual vocabulary only; 40k tokens of Urdu. Fine-tuned version @… See the full description on the dataset page: https://huggingface.co/datasets/mbshr/XSUMUrdu-DW_BBC.

sourceHugging Facecc-by-4.0updated 2y agoView on Hugging Face
0likes24downloads

mbshr/XSUMUrdu-DW_BBC · main · files are served by the source, never re-hosted here