CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01sixuexing /FAERS-NLP FAERS-NLP Version: 1.0Author: sixuexing GitHub: FAERS-NLP Repository Dataset Summary FAERS-NLP is a cleaned and processed version of the FDA Adverse Event Reporting System (FAERS), formatted for natural language retrieval and drug–adverse effect–disease relation extraction. Each record corresponds to a single adverse event report, including structured and semi-structured fields suitable for NLP tasks. Dataset Structure Each CSV row contains the following… See the full description on the dataset page: https://huggingface.co/datasets/sixuexing/FAERS-NLP.tabular1M<n<10M1 likes437 downloads1y agoHugging Face02nlpatunt /D_persuade_2 Persuade_2 The PERSUADE 2.0 corpus (Persuasive Essays for Rating, Selecting, and Understanding Argumentative and Discourse Elements) contains over 25,000 argumentative essays written by 6th–12th grade students in the United States, covering 15 distinct prompts across two writing tasks: independent and source-based writing. The corpus also provides detailed individual and demographic information for each writer. This is the train, test, and validation split of the dataset.… See the full description on the dataset page: https://huggingface.co/datasets/nlpatunt/D_persuade_2.tabular10K<n<100K0 likes364 downloads6mo agoHugging Face03nlpatunt /D_ASAP-AES D_ASAP-AES This is the train, test, and validation split of the ASAP Automated Essay Scoring dataset, prepared for use with the S-GRADES benchmark. Ground truth labels have been removed to prevent leakage during evaluation. For the original dataset with labels, see below. Original Dataset 🔗 ASAP-AES on Kaggle Citation If you use this dataset, please cite the original: @misc{asap_aes, title={ASAP Automated Essay Scoring}… See the full description on the dataset page: https://huggingface.co/datasets/nlpatunt/D_ASAP-AES.tabular10K<n<100K0 likes312 downloads6mo agoHugging Face04GroNLP /ik-nlp-22_winemagtabular10K<n<100K6 likes281 downloads5y agoHugging Face05SanaeLaRose /FAERS-NLP FAERS-NLP Version: 1.0Author: sixuexing GitHub: FAERS-NLP Repository Dataset Summary FAERS-NLP is a cleaned and processed version of the FDA Adverse Event Reporting System (FAERS), formatted for natural language retrieval and drug–adverse effect–disease relation extraction. Each record corresponds to a single adverse event report, including structured and semi-structured fields suitable for NLP tasks. Dataset Structure Each CSV row contains the following… See the full description on the dataset page: https://huggingface.co/datasets/SanaeLaRose/FAERS-NLP.tabular1M<n<10M0 likes209 downloads8mo agoHugging Face06NLPLabNTUST /Merged-CWA CWA Benchmark: A Seismic Dataset from Taiwan for Seismic Research Dataset Description This dataset includes a larger number of seismic events, especially high-magnitude. A comprehensive set of events collected by the Central Weather Bureau in Taiwan. The CWA benchmark features over 40 attributes and ∼500,000 seismograms, providing valuable data labels for various seismology-related tasks. In the future, we will keep updating the dataset to ensure its relevance and… See the full description on the dataset page: https://huggingface.co/datasets/NLPLabNTUST/Merged-CWA.tabular1K<n<10K0 likes201 downloads2y agoHugging Face07nlpscu /Beyond-Flesch Beyond-Flesch: ScienceQA Difficulty Classification with Static and Prompt-Based Metrics A preprocessed subset of ScienceQA for K-12 educational text difficulty classification, along with the static and LLM-derived prompt-based features we use to reproduce Rooein et al. (2024) — Beyond Flesch-Kincaid. This dataset accompanies our class research project (Option 1: reproducing a paper whose original code was not released). What's here File Rows Description… See the full description on the dataset page: https://huggingface.co/datasets/nlpscu/Beyond-Flesch.tabular10K<n<100K0 likes193 downloads4mo agoHugging Face08uoe-nlp /extrinsic_mt_evaltabular10K<n<100K0 likes180 downloads3y agoHugging Face09recogna-nlp /fakerecogna2-abstrativa FakeRecogna 2.0 - Abstractive FakeRecogna 2.0 presents the extension for the FakeRecogna dataset in the context of fake news detection. FakeRecogna includes real and fake news texts collected from online media and ten fact-checking sources in Brazil. An important aspect is the lack of relation between the real and fake news samples, i.e., they are not mutually related to each other to avoid intrinsic bias in the data. The Dataset The fake news collection was performed on… See the full description on the dataset page: https://huggingface.co/datasets/recogna-nlp/fakerecogna2-abstrativa.tabulartext-classification10K<n<100K2 likes144 downloads1y agoHugging Face10lime-nlp /DeepScaleR_Difficulty Difficulty Estimation on DeepScaleR We annotate the entire DeepScaleR dataset with a difficulty score based on the performance of the Qwen 2.5-MATH-7B model. This provides an adaptive signal for curriculum construction and model evaluation. DeepScaleR is a curated dataset of 40,000 reasoning-intensive problems used to train and evaluate reinforcement learning-based methods for large language models. Difficulty Scoring Method Difficulty scores are estimated using the… See the full description on the dataset page: https://huggingface.co/datasets/lime-nlp/DeepScaleR_Difficulty.tabularreinforcement-learning1M<n<10M11 likes132 downloads1y agoHugging Face11recogna-nlp /fakerecogna2-extrativa FakeRecogna 2.0 Extractive FakeRecogna 2.0 presents the extension for the FakeRecogna dataset in the context of fake news detection. FakeRecogna includes real and fake news texts collected from online media and ten fact-checking sources in Brazil. An important aspect is the lack of relation between the real and fake news samples, i.e., they are not mutually related to each other to avoid intrinsic bias in the data. The Dataset The fake news collection was performed on… See the full description on the dataset page: https://huggingface.co/datasets/recogna-nlp/fakerecogna2-extrativa.tabulartext-classification10K<n<100K1 likes120 downloads1y agoHugging Face12jason1966 /aikyatansinha_cybersecurity-cves-for-nlp-dataset Cybersecurity CVEs for NLP Dataset Every CVE since 1999, scrubbed and perfectly formatted for NLP tasks Dataset Info Source: Kaggle Original Size: 38.28 MB Kaggle Downloads: 36 Files: 1 Files NVD_Cybersecurity_Dataset.csv Mirrored from Kaggle tabular100K<n<1M3 likes110 downloads6mo agoHugging Face13cyberpsych /PubMed-Cancer-NLP-Textual-Dataset PubMed-Cancer-NLP-Textual-Dataset This dataset has been obtained from PubMed for research purposes. README will be updated with time. Dataset Details Dataset Description It has multiple cancer samples with labels with their title and abstract from PubMed Repository. Curated by: Om Aryan Dataset Sources Repository: https://pubmed.ncbi.nlm.nih.gov tabularfeature-extraction10K<n<100K0 likes105 downloads2y agoHugging Face14zluvolyote /Dream_NLP_FineTunetabular100K<n<1M0 likes104 downloads4y agoHugging Face15lime-nlp /GSM8K_Difficulty Difficulty Estimation on DeepScaleR We annotate the entire GSM8K dataset with a difficulty score based on the performance of the Qwen 2.5-MATH-7B model. This provides an adaptive signal for curriculum construction and model evaluation. GSM8K (Grade School Math 8K) is a dataset of 8.5K high quality linguistically diverse grade school math word problems. The dataset was created to support the task of question answering on basic mathematical problems that require multi-step reasoning.… See the full description on the dataset page: https://huggingface.co/datasets/lime-nlp/GSM8K_Difficulty.tabular1M<n<10M1 likes100 downloads1y agoHugging Face16nlpyeditepe /tr_rtetabulartext-classification1K<n<10K0 likes83 downloads4y agoHugging Face17upb-nlp /RoJBMO RoJBMO: Junior Balkan Mathematical Olympiad Benchmark RoJBMO is a benchmark of 508 competition mathematics problems drawn from the Junior Balkan Mathematical Olympiad (JBMO), its official shortlists, and the Romanian Team Selection Tests (TST). It is designed to evaluate the mathematical reasoning capabilities of large language models on problems that are underrepresented in existing benchmarks and resistant to training data contamination. Sources Source… See the full description on the dataset page: https://huggingface.co/datasets/upb-nlp/RoJBMO.tabulartext-generationn<1K1 likes81 downloads19d agoHugging Face18nlpatunt /D_ASAP-SAS D_ASAP-SAS This is the train, test, and validation split of the ASAP Short Answer Scoring dataset, prepared for use with the S-GRADES benchmark. Ground truth labels have been removed to prevent leakage during evaluation. For the original dataset with labels, see below. Original Dataset 🔗 ASAP-SAS on Kaggle Citation If you use this dataset, please cite the original: @misc{asapsas2012, author={Barbara and Hamner, Ben and Morgan, Jaison and lynnvandev and… See the full description on the dataset page: https://huggingface.co/datasets/nlpatunt/D_ASAP-SAS.tabular10K<n<100K0 likes76 downloads6mo agoHugging Face19L-NLProc /NyayaAnumana-Transformers-Resultstabular1K<n<10K1 likes72 downloads2y agoHugging Face20SALT-NLP /FLUE-FiQA Dataset Summary Homepage: https://sites.google.com/view/salt-nlp-flang Models: https://huggingface.co/SALT-NLP/FLANG-BERT Repository: https://github.com/SALT-NLP/FLANG FLUE FLUE (Financial Language Understanding Evaluation) is a comprehensive and heterogeneous benchmark that has been built from 5 diverse financial domain specific datasets. Sentiment Classification: Financial PhraseBankSentiment Analysis, Question Answering: FiQA 2018New Headlines Classification:… See the full description on the dataset page: https://huggingface.co/datasets/SALT-NLP/FLUE-FiQA.tabular10K<n<100K6 likes63 downloads4y agoHugging Face21McGill-NLP /ImplicatureX ImplicatureX More information can be found at https://github.com/cesare-spinoso/ImplicatureX. import pandas as pd # skiprows=1: the first line is a leading comment, not part of the header df = pd.read_csv("implicatureX.csv", skiprows=1) Citation If you use our data, please cite us: @misc{piano2026evaluatingcommunicativebeliefupdates, title={Evaluating Communicative Belief Updates in Large Language Models via Implicature Recognition and Cancellation}… See the full description on the dataset page: https://huggingface.co/datasets/McGill-NLP/ImplicatureX.tabular1K<n<10K0 likes60 downloads2mo agoHugging Face22NLP-Debater-Project /IBM-Debater-ArgKPtabular10K<n<100K3 likes55 downloads10mo agoHugging Face23McGill-NLP /CHASE-QA CHASE: Challenging AI with Synthetic Evaluations The pace of evolution of Large Language Models (LLMs) necessitates new approaches for rigorous and comprehensive evaluation. Traditional human annotation is increasingly impracticable due to the complexities and costs involved in generating high-quality, challenging problems. In this work, we introduce **CHASE**, a unified framework to synthetically generate challenging problems using LLMs without human involvement. For a given task… See the full description on the dataset page: https://huggingface.co/datasets/McGill-NLP/CHASE-QA.imagequestion-answeringn<1K0 likes53 downloads2y agoHugging Face24docling-project /docling-nlp-datasetsThis repository contains the models used for docling-nlp. Contents This model repository packages the pretrained assets used by Docling’s NLP components: CRF models for material classification and English part-of-speech tagging fastText models for language detection, metadata, semantic, topic, and person-name classification Regular-expression assets for geographic-location extraction and unit handling A default tokenizer model Correct workflow to add new files… See the full description on the dataset page: https://huggingface.co/datasets/docling-project/docling-nlp-datasets.tabular100K<n<1M0 likes52 downloads12d agoHugging Face25almaz-nlp /almaz-asr-roster The ALMAZ ASR Roster Archived at Zenodo: 10.5281/zenodo.22761871 (concept DOI, always resolves to the latest version). A curated catalog of Azerbaijani speech-to-text artifacts: corpora, models, services, benchmarks and tools. Companion to the ALMAZ Resource Roster, which does the same for text. Schema matches the text roster so the two join, plus three columns speech needs and text does not: hours, condition, and verified. The verified column The standard way a… See the full description on the dataset page: https://huggingface.co/datasets/almaz-nlp/almaz-asr-roster.tabularn<1K0 likes50 downloads8d agoHugging Face26nlpatunt /D_Ielts_Writing_Dataset D_Ielts_Writing_Dataset This dataset contains IELTS Writing scored essays, prepared for use with the S-GRADES benchmark. The test split ground truth labels have been removed to prevent leakage during evaluation. Original Dataset 🔗 IELTS Writing Scored Essays Dataset on Kaggle Citation If you use this dataset, please cite the original source: @misc{mazlum2023ielts, title={IELTS Writing Scored Essays Dataset}, author={Mazlum, Ibrahim}, year={2023}… See the full description on the dataset page: https://huggingface.co/datasets/nlpatunt/D_Ielts_Writing_Dataset.tabular1K<n<10K0 likes49 downloads6mo agoHugging Face27lime-nlp /orz_math_difficulty Difficulty Estimation on Open Reasoner Zero We annotate the entire Open Reasoner Zero dataset with a difficulty score based on the performance of the Qwen 2.5-MATH-7B model. This provides an adaptive signal for curriculum construction. Open Reasoner Zero is a curated a dataset of 57,000 reasoning-intensive problems used to train and evaluate reinforcement learning-based methods for large language models. Difficulty Scoring Method Difficulty scores are estimated using… See the full description on the dataset page: https://huggingface.co/datasets/lime-nlp/orz_math_difficulty.tabular1M<n<10M0 likes44 downloads1y agoHugging Face28lime-nlp /MATH_Difficulty Difficulty Estimation on MATH We annotate the entire MATH dataset with a difficulty score based on the performance of the Qwen 2.5-MATH-7B model. This provides an adaptive signal for curriculum construction and model evaluation. The Mathematics Aptitude Test of Heuristics (MATH) dataset consists of problems from mathematics competitions, including the AMC 10, AMC 12, AIME, and more. Each problem in MATH has a full step-by-step solution, which can be used to teach models to generate… See the full description on the dataset page: https://huggingface.co/datasets/lime-nlp/MATH_Difficulty.tabular1M<n<10M0 likes43 downloads1y agoHugging Face29sinhala-nlp /SemiSOLD SOLD - A Benchmark for Sinhala Offensive Language Identification In this repository, we introduce the {S}inhala {O}ffensive {L}anguage {D}ataset (SOLD) and present multiple experiments on this dataset. SOLD is a manually annotated dataset containing 10,000 posts from Twitter annotated as offensive and not offensive at both sentence-level and token-level. SOLD is the largest offensive language dataset compiled for Sinhala. We also introduce SemiSOLD, a larger dataset containing more… See the full description on the dataset page: https://huggingface.co/datasets/sinhala-nlp/SemiSOLD.tabular100K<n<1M0 likes41 downloads3y agoHugging Face30nlpatunt /D_ASAP_plus_plus D_ASAP_plus_plus This is the train, test, and validation split of the ASAP++ dataset, prepared for use with the S-GRADES benchmark. Ground truth labels have been removed to prevent leakage during evaluation. Original Dataset ASAP++ enriches the original ASAP dataset with attribute-specific essay scores (content, organization, style, etc.). 🔗 ASAP++ Official Page Citation If you use this dataset, please cite the original: @inproceedings{mathias2018asap++… See the full description on the dataset page: https://huggingface.co/datasets/nlpatunt/D_ASAP_plus_plus.tabular10K<n<100K0 likes41 downloads6mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.