datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
FAERS-NLP
FAERS-NLP
Version: 1.0Author: sixuexing
GitHub: FAERS-NLP Repository
Dataset Summary
FAERS-NLP is a cleaned and processed version of the FDA Adverse Event Reporting System (FAERS), formatted for natural language retrieval and drug–adverse effect–disease relation extraction.
Each record corresponds to a single adverse event report, including structured and semi-structured fields suitable for NLP tasks.
Dataset Structure
Each CSV row contains the following… See the full description on the dataset page: https://huggingface.co/datasets/SanaeLaRose/FAERS-NLP.650k_sanadset
Sanadset 650K: Data on Hadith Narrators
Dataset Description
Sanadset is a large-scale dataset containing over 650,986 Hadith records collected from 926 historical Arabic books. This dataset was created to assist in the computational analysis of Islamic Hadiths, specifically focusing on the chain of narrators (Sanad) and the content (Matn).
It allows researchers to apply Machine Learning and NLP techniques to tasks such as:
Classifying Hadiths (Strong/Weak).
Analyzing… See the full description on the dataset page: https://huggingface.co/datasets/freococo/650k_sanadset.citation_refrence_linkcode-switching-codesaviours-si26-SanaHeart-Disease-Prediction-datasetcitation_context
Citation Contexts for Scientific Evidence Retrieval
Dataset Description
This dataset contains 9,920 citation occurrences extracted from English-language scientific papers. Each row represents one occurrence of a citation in a source paper and links it to the cited paper. It provides four increasingly broad representations of the citation context:
the sentence containing the citation (context_c1_sentence);
the complete source paragraph (context_c2_paragraph);
a… See the full description on the dataset page: https://huggingface.co/datasets/sanaa-11/citation_context.math-datasetTelugu_movie_reviewsUzbek-restaurant-domain-sentiment-reviewssindhi_sentiment
Sindhi Sentiment Analysis Dataset (50k)
A balanced, three-class sentiment analysis dataset for the Sindhi language (سنڌي) in the Perso-Arabic script, containing 50,000 labeled sentences. Sindhi is a low-resource language spoken by over 30 million people, mainly in Sindh, Pakistan. This dataset supports training and benchmarking of sentiment classification models for Sindhi NLP.
Dataset Summary
Language: Sindhi (sd), Perso-Arabic script
Task: Sentence-level… See the full description on the dataset page: https://huggingface.co/datasets/Sanapalijo/sindhi_sentiment.PROSPERO-InclusionExclusionCriteria
PROSPERO Inclusion/Exclusion Criteria Dataset
This dataset is a curated and preprocessed collection of clinical research objectives and their corresponding inclusion and exclusion criteria, extracted from the PROSPERO international prospective register of systematic reviews.
Description
The dataset was constructed to support fine-tuning of large language models (LLMs) for tasks involving the automated generation of eligibility criteria based on research objectives. It… See the full description on the dataset page: https://huggingface.co/datasets/sanaa-11/PROSPERO-InclusionExclusionCriteria.airbnb-price-dataTelugu_sentiment_sentences
Dataset Card for Dataset Name
Dataset Summary
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Supported Tasks and Leaderboards
[More Information Needed]
Languages
[More Information Needed]
Dataset Structure
Data Instances
[More Information Needed]
Data Fields
[More Information Needed]
Data Splits
[More Information Needed]
Dataset Creation… See the full description on the dataset page: https://huggingface.co/datasets/Sanath369/Telugu_sentiment_sentences.Frensh-math-dataset-5Ksanad_dfArabicEmpatheticDialogues-with-QuranFAQs-for-SMEsSMEs_Chatbot_datasetSMEs-datasetSMEs-Ordersfood_health_dataset
