datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
FAERS-NLP
FAERS-NLP
Version: 1.0Author: sixuexing
GitHub: FAERS-NLP Repository
Dataset Summary
FAERS-NLP is a cleaned and processed version of the FDA Adverse Event Reporting System (FAERS), formatted for natural language retrieval and drug–adverse effect–disease relation extraction.
Each record corresponds to a single adverse event report, including structured and semi-structured fields suitable for NLP tasks.
Dataset Structure
Each CSV row contains the following… See the full description on the dataset page: https://huggingface.co/datasets/SanaeLaRose/FAERS-NLP.650k_sanadset
Sanadset 650K: Data on Hadith Narrators
Dataset Description
Sanadset is a large-scale dataset containing over 650,986 Hadith records collected from 926 historical Arabic books. This dataset was created to assist in the computational analysis of Islamic Hadiths, specifically focusing on the chain of narrators (Sanad) and the content (Matn).
It allows researchers to apply Machine Learning and NLP techniques to tasks such as:
Classifying Hadiths (Strong/Weak).
Analyzing… See the full description on the dataset page: https://huggingface.co/datasets/freococo/650k_sanadset.openve_q5_vcos_tcfg1_step200_shift3grab-safe-driver-telematics-cleaned-datasetcitation_refrence_linkHeart-Disease-Prediction-datasetcitation_context
Citation Contexts for Scientific Evidence Retrieval
Dataset Description
This dataset contains 9,920 citation occurrences extracted from English-language scientific papers. Each row represents one occurrence of a citation in a source paper and links it to the cited paper. It provides four increasingly broad representations of the citation context:
the sentence containing the citation (context_c1_sentence);
the complete source paragraph (context_c2_paragraph);
a… See the full description on the dataset page: https://huggingface.co/datasets/sanaa-11/citation_context.march-madness-sft-dataairbnb-price-dataoptuna-logs-task-17africa-perte-de-couverture-forestiere-en-sanaga-maritime-littoral-cameroun
Perte de couverture forestière en Sanaga maritime, Littoral, Cameroun | Africa (original)
Size category: n<1K - Formats: parquet - Sector: humanitarian_development - Engineered by Electric Sheep Africa
TL;DR
This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context.
What This Dataset Covers
Public… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-perte-de-couverture-forestiere-en-sanaga-maritime-littoral-cameroun.optuna-logs-task-22cleaned_scaled_fire_data.csvoptuna-logs-task-19SMEs-datasetSanatanLordAnomalyDetection
SanatanLordAnomalyDetection
tags: Anomaly Detection, Spirituality, Trends
Note: This is an AI-generated dataset so its content may be inaccurate or false
Dataset Description:
The 'SanatanLordAnomalyDetection' dataset comprises a collection of texts related to Sanatan Dharma (Hinduism) with an emphasis on spiritual trends and activities. The dataset is structured to facilitate anomaly detection in the occurrence and representation of 'Sanatan lord' references, with each entry tagged… See the full description on the dataset page: https://huggingface.co/datasets/infinite-dataset-hub/SanatanLordAnomalyDetection.stratified_20160110_20201231_3_days_excl_v3_2016_all_neighbors.csvSMEs-Orders
