datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
English_French_Webpages_Scraped_Translated
English French Webpages Scraped Translated
Dataset Summary
French/English parallel texts for training translation models. Over 17.1 million sentences in French and English. Dataset created by Chris Callison-Burch, who crawled millions of web pages and then used a set of simple heuristics to transform French URLs onto English URLs, and assumed that these documents are translations of each other. This is the main dataset of Workshop on Statistical Machine Translation (WML)… See the full description on the dataset page: https://huggingface.co/datasets/Nicolas-BZRD/English_French_Webpages_Scraped_Translated.webpage-summarizationDataset containing approx. 20k scraped webpages from various sources, paired with synthetic short-, medium-, and long-length summaries.
The "short" summaries are 2–3 sentences, the "medium" summaries are one paragraph, and the "long" summaries are up to three paragraphs.
The summaries are always in English, even for non-English inputs. However, it is likely the dataset contains too little non-English text to generalize well for translation.
The webpages were sampled as follows:
60% from the… See the full description on the dataset page: https://huggingface.co/datasets/caelunshun/webpage-summarization.Artworks_as_WebPagesptbr-webpages
