CoolFace
21 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01SanaeLaRose /FAERS-NLP FAERS-NLP Version: 1.0Author: sixuexing GitHub: FAERS-NLP Repository Dataset Summary FAERS-NLP is a cleaned and processed version of the FDA Adverse Event Reporting System (FAERS), formatted for natural language retrieval and drug–adverse effect–disease relation extraction. Each record corresponds to a single adverse event report, including structured and semi-structured fields suitable for NLP tasks. Dataset Structure Each CSV row contains the following… See the full description on the dataset page: https://huggingface.co/datasets/SanaeLaRose/FAERS-NLP.tabular1M<n<10M0 likes211 downloads8mo agoHugging Face02freococo /650k_sanadset Sanadset 650K: Data on Hadith Narrators Dataset Description Sanadset is a large-scale dataset containing over 650,986 Hadith records collected from 926 historical Arabic books. This dataset was created to assist in the computational analysis of Islamic Hadiths, specifically focusing on the chain of narrators (Sanad) and the content (Matn). It allows researchers to apply Machine Learning and NLP techniques to tasks such as: Classifying Hadiths (Strong/Weak). Analyzing… See the full description on the dataset page: https://huggingface.co/datasets/freococo/650k_sanadset.tabulartext-classification100K<n<1M0 likes79 downloads7mo agoHugging Face03sanaa-11 /citation_refrence_linktabular1M<n<10M0 likes33 downloads7d agoHugging Face04sanaisrail /code-switching-codesaviours-si26-Sanatext1K<n<10K0 likes32 downloads27d agoHugging Face05sanadf234 /Heart-Disease-Prediction-datasettabular10K<n<100K1 likes31 downloads10mo agoHugging Face06sanaa-11 /citation_context Citation Contexts for Scientific Evidence Retrieval Dataset Description This dataset contains 9,920 citation occurrences extracted from English-language scientific papers. Each row represents one occurrence of a citation in a source paper and links it to the cited paper. It provides four increasingly broad representations of the citation context: the sentence containing the citation (context_c1_sentence); the complete source paragraph (context_c2_paragraph); a… See the full description on the dataset page: https://huggingface.co/datasets/sanaa-11/citation_context.tabulartext-retrieval1K<n<10K0 likes25 downloads8h agoHugging Face07sanaa-11 /math-datasettext1K<n<10K0 likes17 downloads2y agoHugging Face08Sanath369 /Telugu_movie_reviewstexttext-classificationn<1K0 likes14 downloads3y agoHugging Face09Sanatbek /Uzbek-restaurant-domain-sentiment-reviewstext1K<n<10K0 likes14 downloads3y agoHugging Face10Sanapalijo /sindhi_sentiment Sindhi Sentiment Analysis Dataset (50k) A balanced, three-class sentiment analysis dataset for the Sindhi language (سنڌي) in the Perso-Arabic script, containing 50,000 labeled sentences. Sindhi is a low-resource language spoken by over 30 million people, mainly in Sindh, Pakistan. This dataset supports training and benchmarking of sentiment classification models for Sindhi NLP. Dataset Summary Language: Sindhi (sd), Perso-Arabic script Task: Sentence-level… See the full description on the dataset page: https://huggingface.co/datasets/Sanapalijo/sindhi_sentiment.texttext-classification10K<n<100K0 likes14 downloads3mo agoHugging Face11sanaa-11 /PROSPERO-InclusionExclusionCriteria PROSPERO Inclusion/Exclusion Criteria Dataset This dataset is a curated and preprocessed collection of clinical research objectives and their corresponding inclusion and exclusion criteria, extracted from the PROSPERO international prospective register of systematic reviews. Description The dataset was constructed to support fine-tuning of large language models (LLMs) for tasks involving the automated generation of eligibility criteria based on research objectives. It… See the full description on the dataset page: https://huggingface.co/datasets/sanaa-11/PROSPERO-InclusionExclusionCriteria.textn<1K0 likes12 downloads1y agoHugging Face12sanatbhatia /airbnb-price-datatabular10K<n<100K0 likes12 downloads2mo agoHugging Face13Sanath369 /Telugu_sentiment_sentences Dataset Card for Dataset Name Dataset Summary This dataset card aims to be a base template for new datasets. It has been generated using this raw template. Supported Tasks and Leaderboards [More Information Needed] Languages [More Information Needed] Dataset Structure Data Instances [More Information Needed] Data Fields [More Information Needed] Data Splits [More Information Needed] Dataset Creation… See the full description on the dataset page: https://huggingface.co/datasets/Sanath369/Telugu_sentiment_sentences.text10K<n<100K0 likes11 downloads3y agoHugging Face14sanaa-11 /Frensh-math-dataset-5Ktext1K<n<10K1 likes7 downloads2y agoHugging Face15CUTD /sanad_dftext10K<n<100K0 likes6 downloads2y agoHugging Face16SANAD-GraduationProject2026 /ArabicEmpatheticDialogues-with-Qurantext10K<n<100K0 likes4 downloads2mo agoHugging Face17sanadf234 /FAQs-for-SMEstext10K<n<100K0 likes3 downloads10mo agoHugging Face18sanadf234 /SMEs_Chatbot_datasettext10K<n<100K1 likes2 downloads10mo agoHugging Face19sanadf234 /SMEs-datasettabular10K<n<100K0 likes2 downloads10mo agoHugging Face20sanadf234 /SMEs-Orderstabular10K<n<100K1 likes1 downloads10mo agoHugging Face21sanadf234 /food_health_datasettext10K<n<100K1 likes1 downloads10mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.