datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
PubMed-Cancer-NLP-Textual-Dataset
PubMed-Cancer-NLP-Textual-Dataset
This dataset has been obtained from PubMed for research purposes. README will be updated with time.
Dataset Details
Dataset Description
It has multiple cancer samples with labels with their title and abstract from PubMed Repository.
Curated by: Om Aryan
Dataset Sources
Repository: https://pubmed.ncbi.nlm.nih.gov
Thai-Semantic-Textual-Similarity-BenchmarkSentence representation plays a crucial role in NLP downstream tasks such as NLI, text classification, and STS. Recent sentence representation training techniques require NLI or STS datasets. However, there are no equivalent Thai NLI or STS datasets for sentence representation training.
To address this problem we provide the Thai sentence vector benchmark. We evaluate the Spearman correlation score of the sentence representations’ performance on Thai STS-B (translated version of STS-B).… See the full description on the dataset page: https://huggingface.co/datasets/mrp/Thai-Semantic-Textual-Similarity-Benchmark.Dataset_Automatic_Essay_Scoring_Essay-EssayScore_and_24_textual_featuressemantic-textual-similarityTextual-Natural-Contextual-ClassificationGiven the scarcity of datasets for understanding natural language in visual scenes, we introduce a novel textual entailment dataset, named Textual Natural Contextual Classification (TNCC).
This dataset is formulated on the foundation of Crisscrossed Captions (https://github.com/google-research-datasets/Crisscrossed-Captions), an image captioning dataset supplied with human-rated semantic similarity ratings on a continuous scale from 0 to 5.
We tailor the dataset to suit a binary… See the full description on the dataset page: https://huggingface.co/datasets/zhili312/Textual-Natural-Contextual-Classification.textual-inferenceSynthetic dataset using Tevatron/msmarco-passage-corpus and GPT-4o to gather up-to 5 points of textual inference. Dataset contains roughly 20 million tokens across nearly 100k rows.
financial_fraud_textual_datasetsemantic_textual_similarity_dataset
Turkish Semantic Textual Similarity Dataset
Bu veri seti, Türkçe cümle çiftleri arasındaki anlamsal benzerliği değerlendirmek amacıyla hazırlanmış 80 özgün örnekten oluşur.
Veri yapısı
sentence1: Birinci cümle
sentence2: İkinci cümle
score: İnsan değerlendirmesiyle belirlenen anlamsal benzerlik puanı
Puanlama ölçeği
Puanlar 0 ile 5 arasındadır:
5: Aynı anlamı ifade eden cümleler
4: Büyük ölçüde aynı, küçük ayrıntı farkları bulunan cümleler
3:… See the full description on the dataset page: https://huggingface.co/datasets/aliFurkan123/semantic_textual_similarity_dataset.descriptive_textual_contract_smellstextual-inference-with-confidenceSynthetic dataset using Tevatron/msmarco-passage-corpus and GPT-4o to generate up-to five inferences based on the passage along with confidence scores.
Prompt used in generation:
Given the following passage, generate a series of 5 inferences that can be drawn from the text. Include a mix of well-reasoned, insightful inferences as well as some that may be less supported or even incorrect. Assign each inference a confidence score between 0 and 1, where 1 indicates high confidence in the… See the full description on the dataset page: https://huggingface.co/datasets/will4381/textual-inference-with-confidence.
