datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
wiki-paragraphs
Dataset Card for wiki-paragraphs
Dataset Summary
The wiki-paragraphs dataset is constructed by automatically sampling two paragraphs from a Wikipedia article. If they are from the same section, they will be considered a "semantic match", otherwise as "dissimilar". Dissimilar paragraphs can in theory also be sampled from other documents, but have not shown any improvement in the particular evaluation of the linked work.The alignment is in no way meant as an accurate… See the full description on the dataset page: https://huggingface.co/datasets/dennlinger/wiki-paragraphs.scientific-paragraphs-categorization
A Multi-lingual Dataset of Classified Paragraphs from Open Access Scientific
We present a dataset of 833k paragraphs extracted from CC-BY licensed
scientific publications, classified into four categories: acknowledgments, data
mentions, software/code mentions, and clinical trial mentions. The paragraphs
are primarily in English and French, with additional European languages
represented. Each paragraph is annotated with language identification (using
fastText) and scientific domain… See the full description on the dataset page: https://huggingface.co/datasets/dataesr/scientific-paragraphs-categorization.cjeu-paragraph-retrievalgp-long-paragraphssecrethackatondata_paragraphsParagraphProof
ParagraphProof
tags: detection, paragraph, evidence-verification
Note: This is an AI-generated dataset so its content may be inaccurate or false
Dataset Description:
This dataset contains paragraphs from various sources with the goal of identifying whether there is evidence supporting a specific claim within each paragraph. The dataset has been designed to aid in the development and training of machine learning models for the task of evidence-verification in textual data. Each… See the full description on the dataset page: https://huggingface.co/datasets/infinite-dataset-hub/ParagraphProof.ParagraphSentinel
ParagraphSentinel
tags: Claim Analysis, Text Parsing, Anomaly Detection
Note: This is an AI-generated dataset so its content may be inaccurate or false
Dataset Description:
The 'ParagraphSentinel' dataset is designed for the task of identifying and classifying claims within a text paragraph. Each entry in the dataset includes a paragraph of text and a label that indicates whether a claim is present and the type of claim identified. The dataset is suitable for machine learning models… See the full description on the dataset page: https://huggingface.co/datasets/infinite-dataset-hub/ParagraphSentinel.sotu-paragraphsThis is a dataset containing the United States Presidential State of the Union Addresses through 2020; derived from the sotu R package.
paragraph_to_paragraphParagraphVerification
ParagraphVerification
tags: truthfulness, paraphrase, string matching
Note: This is an AI-generated dataset so its content may be inaccurate or false
Dataset Description:
The 'ParagraphVerification' dataset comprises paragraphs and their associated claims that include exact sentences from the original text. The objective is to verify whether the claims are truthful based on the given paragraph. The dataset labels are 'Verified' if the claim is exactly present in the paragraph and… See the full description on the dataset page: https://huggingface.co/datasets/infinite-dataset-hub/ParagraphVerification.Gavin_yiddish_raw_HTR_and_groundtruth_paragraphsorca_paragraphsFirst-of-paragraph
First-of-paragraph
First sentence of paragraphs collected from Random paragraphs dataset
Created by Melphin as an experiment
Labels
0: Not first sentence
1: First sentence
deduplicated_paragraphs_mainparagraph001paragraph_gutenberg_top50paragraph_classification_eval
