CoolFace
17 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01dennlinger /wiki-paragraphs Dataset Card for wiki-paragraphs Dataset Summary The wiki-paragraphs dataset is constructed by automatically sampling two paragraphs from a Wikipedia article. If they are from the same section, they will be considered a "semantic match", otherwise as "dissimilar". Dissimilar paragraphs can in theory also be sampled from other documents, but have not shown any improvement in the particular evaluation of the linked work.The alignment is in no way meant as an accurate… See the full description on the dataset page: https://huggingface.co/datasets/dennlinger/wiki-paragraphs.texttext-classification10M<n<100M0 likes70 downloads4y agoHugging Face02dataesr /scientific-paragraphs-categorization A Multi-lingual Dataset of Classified Paragraphs from Open Access Scientific We present a dataset of 833k paragraphs extracted from CC-BY licensed scientific publications, classified into four categories: acknowledgments, data mentions, software/code mentions, and clinical trial mentions. The paragraphs are primarily in English and French, with additional European languages represented. Each paragraph is annotated with language identification (using fastText) and scientific domain… See the full description on the dataset page: https://huggingface.co/datasets/dataesr/scientific-paragraphs-categorization.texttext-classification100K<n<1M2 likes36 downloads1y agoHugging Face03larimo /cjeu-paragraph-retrievaltext10K<n<100K1 likes33 downloads2y agoHugging Face04GuillermoTBB /gp-long-paragraphstabular10M<n<100M1 likes27 downloads2y agoHugging Face05Brianverb /secrethackatondata_paragraphstabularn<1K0 likes15 downloads4y agoHugging Face06infinite-dataset-hub /ParagraphProof ParagraphProof tags: detection, paragraph, evidence-verification Note: This is an AI-generated dataset so its content may be inaccurate or false Dataset Description: This dataset contains paragraphs from various sources with the goal of identifying whether there is evidence supporting a specific claim within each paragraph. The dataset has been designed to aid in the development and training of machine learning models for the task of evidence-verification in textual data. Each… See the full description on the dataset page: https://huggingface.co/datasets/infinite-dataset-hub/ParagraphProof.tabularn<1K0 likes13 downloads2y agoHugging Face07infinite-dataset-hub /ParagraphSentinel ParagraphSentinel tags: Claim Analysis, Text Parsing, Anomaly Detection Note: This is an AI-generated dataset so its content may be inaccurate or false Dataset Description: The 'ParagraphSentinel' dataset is designed for the task of identifying and classifying claims within a text paragraph. Each entry in the dataset includes a paragraph of text and a label that indicates whether a claim is present and the type of claim identified. The dataset is suitable for machine learning models… See the full description on the dataset page: https://huggingface.co/datasets/infinite-dataset-hub/ParagraphSentinel.tabularn<1K0 likes11 downloads2y agoHugging Face08textminr /sotu-paragraphsThis is a dataset containing the United States Presidential State of the Union Addresses through 2020; derived from the sotu R package. textn<1K0 likes10 downloads3y agoHugging Face09coastalcph /paragraph_to_paragraphtabular100K<n<1M1 likes10 downloads3y agoHugging Face10infinite-dataset-hub /ParagraphVerification ParagraphVerification tags: truthfulness, paraphrase, string matching Note: This is an AI-generated dataset so its content may be inaccurate or false Dataset Description: The 'ParagraphVerification' dataset comprises paragraphs and their associated claims that include exact sentences from the original text. The objective is to verify whether the claims are truthful based on the given paragraph. The dataset labels are 'Verified' if the claim is exactly present in the paragraph and… See the full description on the dataset page: https://huggingface.co/datasets/infinite-dataset-hub/ParagraphVerification.textn<1K1 likes9 downloads2y agoHugging Face11MarineLives /Gavin_yiddish_raw_HTR_and_groundtruth_paragraphstextn<1K1 likes8 downloads2y agoHugging Face12PritiLohra /orca_paragraphstextn<1K0 likes5 downloads3y agoHugging Face13Melphin /First-of-paragraph First-of-paragraph First sentence of paragraphs collected from Random paragraphs dataset Created by Melphin as an experiment Labels 0: Not first sentence 1: First sentence tabular10K<n<100K0 likes4 downloads3y agoHugging Face14Saur13 /deduplicated_paragraphs_maintextn<1K0 likes4 downloads3y agoHugging Face15andrewatef /paragraph001text100K<n<1M0 likes2 downloads2y agoHugging Face16ambrosfitz /paragraph_gutenberg_top50text10K<n<100K0 likes2 downloads1y agoHugging Face17dataesr /paragraph_classification_evaltext1K<n<10K0 likes2 downloads8mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.