CoolFace
Datasetpublic

mbshr/XSUMUrdu-DW_BBC

Urdu_DW-BBC-512 Dataset Summary Urdu Summarization Dataset containining 76,637 records of Article + Summary pairs scrapped from BBC Urdu and DW Urdu News Websites. Preprocessed Version: upto 512 tokens (~words); removed URLs, Pic Captions etc Supported Tasks and Leaderboards Summarization: Extractive and Abstractive urT5 adapted from mT5 having monolingual vocabulary only; 40k tokens of Urdu. Fine-tuned version @… See the full description on the dataset page: https://huggingface.co/datasets/mbshr/XSUMUrdu-DW_BBC.

sourceHugging Facecc-by-4.0updated 2y agoView on Hugging Face
0likes24downloads
13 commits on main
ecc64eb2y ago

Update README.md

mbshr
e9d479c2y ago

Update README.md

mbshr
9091e712y ago

Update README.md

mbshr
5bc64bb3y ago

Update README.md

mbshr
d714a5e3y ago

Update README.md

mbshr
5179c513y ago

Update README.md

mbshr
1d97f323y ago

Update README.md

mbshr
15ba52a3y ago

Update README.md

mbshr
3afd1004y ago

Update README.md

mbshr
3a2f5d64y ago

Update README.md

mbshr
64c6baf4y ago

512 token trunc, article + summary, train 72.3k and test 3.8k

mbshr
4cb753c4y ago

Update README.md

mbshr
2ef11604y ago

initial commit

mbshr