mbshr/XSUMUrdu-DW_BBC
Urdu_DW-BBC-512 Dataset Summary Urdu Summarization Dataset containining 76,637 records of Article + Summary pairs scrapped from BBC Urdu and DW Urdu News Websites. Preprocessed Version: upto 512 tokens (~words); removed URLs, Pic Captions etc Supported Tasks and Leaderboards Summarization: Extractive and Abstractive urT5 adapted from mT5 having monolingual vocabulary only; 40k tokens of Urdu. Fine-tuned version @… See the full description on the dataset page: https://huggingface.co/datasets/mbshr/XSUMUrdu-DW_BBC.
Update README.md
Update README.md
Update README.md
Update README.md
Update README.md
Update README.md
Update README.md
Update README.md
Update README.md
Update README.md
512 token trunc, article + summary, train 72.3k and test 3.8k
Update README.md
initial commit
