jhu-clsp/megawika-2
MegaWika 2 MegaWika 2 is an improved multilingual text dataset containing a structured view of Wikipedia articles, the web sources they cite, source text quality estimates, article text translations, and additional article enrichments. Note: Web citations (sources) in the HuggingFace dataset do not include scraped source text; use rehydrate-citations.py to rehydrate them. The initial data release is based on Wikipedia dumps from May 1, 2024. In total, the data contains about 77… See the full description on the dataset page: https://huggingface.co/datasets/jhu-clsp/megawika-2.
Conversations for this repository live on Hugging Face.
CoolFace shows imported repositories read-only. Posting into someone else’s repository from here would need an authorised integration and the account holder’s consent, so the link goes to the source instead.
Open discussions on Hugging Face