datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
wiki-paragraphs
Dataset Card for wiki-paragraphs
Dataset Summary
The wiki-paragraphs dataset is constructed by automatically sampling two paragraphs from a Wikipedia article. If they are from the same section, they will be considered a "semantic match", otherwise as "dissimilar". Dissimilar paragraphs can in theory also be sampled from other documents, but have not shown any improvement in the particular evaluation of the linked work.The alignment is in no way meant as an accurate… See the full description on the dataset page: https://huggingface.co/datasets/dennlinger/wiki-paragraphs.scientific-paragraphs-categorization
A Multi-lingual Dataset of Classified Paragraphs from Open Access Scientific
We present a dataset of 833k paragraphs extracted from CC-BY licensed
scientific publications, classified into four categories: acknowledgments, data
mentions, software/code mentions, and clinical trial mentions. The paragraphs
are primarily in English and French, with additional European languages
represented. Each paragraph is annotated with language identification (using
fastText) and scientific domain… See the full description on the dataset page: https://huggingface.co/datasets/dataesr/scientific-paragraphs-categorization.gp-long-paragraphssecrethackatondata_paragraphssotu-paragraphsThis is a dataset containing the United States Presidential State of the Union Addresses through 2020; derived from the sotu R package.
Gavin_yiddish_raw_HTR_and_groundtruth_paragraphsorca_paragraphsdeduplicated_paragraphs_main
