CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01evalitahf /sentiment_analysisSENTIPOLC 2016 dataset The SENTIPOLC 2016 dataset contains 9410 tweets annotated for subjectivity, overall and literal polarity, and irony. The dataset has been created and used in the context of the SENTIPOLC 2016 task (http://www.di.unito.it/~tutreeb/sentipolc-evalita16/index.html), organized as part of the EVALITA 2016 evaluation campaign. Original files available here: https://live.european-language-grid.eu/catalogue/corpus/7479/download/ If you find this dataset useful please cite:… See the full description on the dataset page: https://huggingface.co/datasets/evalitahf/sentiment_analysis.texttext-classification1K<n<10K0 likes318 downloads2y agoHugging Face02mehti /LMOD-Cataract-1K-surgical-analysis-cot Cataract-1K LLM-Generated Surgical Instructions Dataset Overview This dataset is derived from the Cataract-1K dataset (part of the LMOD benchmark) and enhanced using Qwen3-VL-30B-A3B-Thinking, a large vision-language model with reasoning capabilities. It is designed for training medical AI systems to provide actionable surgical guidance with transparent reasoning. Generation Process Source Data: Cataract-1K processed frames with segmentation annotations… See the full description on the dataset page: https://huggingface.co/datasets/mehti/LMOD-Cataract-1K-surgical-analysis-cot.imagevisual-question-answering10K<n<100K0 likes277 downloads7mo agoHugging Face03RobinChen2001 /A-Survey-for-LLM-Agent-Trajectory-Analysis A Survey for LLM Agent Trajectory Analysis This dataset repository hosts the survey paper A Survey for LLM Agent Trajectory Analysis: From Failure Attribution to Enhancement and a structured metadata snapshot of the companion paper collection from Awesome-LLM-Agent-Trajectory-Analysis. The repository is intended for discovery, citation, and lightweight analysis of the literature around LLM agent trajectory analysis, including failure attribution, trajectory-based debugging… See the full description on the dataset page: https://huggingface.co/datasets/RobinChen2001/A-Survey-for-LLM-Agent-Trajectory-Analysis.documentn<1K2 likes257 downloads3mo agoHugging Face04deep-analysis-research /simple-evalstext100K<n<1M0 likes223 downloads10mo agoHugging Face05openerotica /erotica-analysisThis dataset is roughly 27k examples of erotica stories which I've fed through GPT-3.5-turbo-16k to obtain a summary, writing prompt, and tags as a response. I've filtered out all the refusals, and deleted a fair ammount of "GPT-isms". I'd still like to go through this again to prune any remaining low quality responses I've missed, but I think this is a good start. Most of the context size comes from the stories themselves, not the responses. Please consider supporting my Patreon… See the full description on the dataset page: https://huggingface.co/datasets/openerotica/erotica-analysis.text10K<n<100K36 likes155 downloads2y agoHugging Face06zjunlp /DataMind-Analysis-SFT-DataThis repository contains the data presented in Why Do Open-Source LLMs Struggle with Data Analysis? A Systematic Empirical Study Code: https://github.com/zjunlp/DataMind text1K<n<10K1 likes136 downloads1y agoHugging Face07daaain /swebench-verified-deepseek-v4-flash-failure-analysis SWE-bench Verified runs & failure analysis — DeepSeek-V4-flash (local) × mini-swe-agent Per-instance analysis of SWE-bench Verified runs of a locally-served DeepSeek-V4-flash model driven by mini-swe-agent, graded with the official SWE-bench harness. Each instance carries the full agent trajectory, a readable transcript, the submitted patch, the harness test output, deterministic metrics, and a hand-verified qualitative root-cause diagnosis. Current numbers (resolve rates… See the full description on the dataset page: https://huggingface.co/datasets/daaain/swebench-verified-deepseek-v4-flash-failure-analysis.tabulartext-generationn<1K0 likes126 downloads3mo agoHugging Face08newsbang /math_benbench_data_leak_analysis Dataset description This is a math dataset mixed from four open-source data. It was used to analyze the contamination test on the MATH and contains 1M samples. Dataset fields question    question from open-source data solution    the answer corresponding to question 5grams    5-gram list of f"{question} {answer}" test_question    the most relevant question from MATH test_solution    the answer corresponding to test_question test_5grams    5-gram list of… See the full description on the dataset page: https://huggingface.co/datasets/newsbang/math_benbench_data_leak_analysis.text100K<n<1M3 likes93 downloads2y agoHugging Face09max-rl /variance_analysis Variance Analysis This repository contains the data, scripts, and generated figures used for variance analysis experiments. Contents data/SmolLM: SmolLM GSM8K rollout data. data/Qwen3: Qwen3 math rollout data, including 1.7B 512 x 512 and 4B 1024 x 128 samples. data/Maze/variance: Maze rollout data for variance analysis. outputs: generated JSON summaries and figures. *.py and run_*.sh: analysis, plotting, and Slurm launch scripts. See data/README.md for additional data… See the full description on the dataset page: https://huggingface.co/datasets/max-rl/variance_analysis.document1M<n<10M0 likes93 downloads5mo agoHugging Face10Daniel-ML /sentiment-analysis-for-financial-news-v2text1K<n<10K1 likes84 downloads2y agoHugging Face11turkish-nlp-suite /turkish-morph-analysis Dataset Card for TrMorphTester This dataset is a testing dataset for Turkish morphology, aiming to calculate how other subword strategies aligns with morphological segmentation of Turkish. The data is automatically generated from Turkish morpoholigical lexicon. For each row, we offer a surface form, then lemma and all suffixes, separated by a + character. The dataset has several splits for several purposes: lemma: Surface form is same with lemma, no suffixes at all. Nouns, verbs… See the full description on the dataset page: https://huggingface.co/datasets/turkish-nlp-suite/turkish-morph-analysis.text100K<n<1M2 likes80 downloads7mo agoHugging Face12allenai /analysis_olmoe1K<n<10K0 likes75 downloads2y agoHugging Face13Morty0311 /misc-cfo-testing-cfo-analysis misc-cfo-testing CFO analysis This folder is the self-contained analysis output for: C:\Users\15255\Desktop\Research\CSE237D\morty_data\misc-cfo-testing The source data and the Weyl pipeline are read-only. All generated scripts, fingerprints, statistics, logs, PNGs, and SVGs remain in this analysis folder. Start with RESULTS.md. Dataset and estimator parameters Experiments: faraday (4 min), reboot (5 min), reboot-10m (10 min) Receiver: pluto11 Input sample rate:… See the full description on the dataset page: https://huggingface.co/datasets/Morty0311/misc-cfo-testing-cfo-analysis.imagen<1K0 likes75 downloads2mo agoHugging Face14khaihernlow /massive-stock-news-analysis-db-for-nlpbackteststext1M<n<10M0 likes74 downloads2y agoHugging Face15max-rl /perplexity_analysis Perplexity Analysis This repository contains the data, scripts, and generated figures used for perplexity analysis experiments. Contents data/Qwen3: rollout data for Qwen3 1.7B and 4B Base, GRPO, and MaxRL models on AIME25 and BeyondAIME. data/Maze/perplexity: maze rollout data and derived perplexity analysis artifacts. outputs: generated JSON summaries and figures for Qwen3 analyses. *.py: analysis and plotting scripts. See data/README.md for additional data details… See the full description on the dataset page: https://huggingface.co/datasets/max-rl/perplexity_analysis.image10M<n<100M0 likes74 downloads5mo agoHugging Face16lukeslp /us-military-veteran-analysis US Military & Veteran Analysis by State 50 states: veterans, firearms, PTSD, suicide rates, VA healthcare State-level integration of veteran demographics, firearm ownership, mental health indicators, and VA healthcare utilization from Census ACS, RAND, ATF, CDC, VA, and DoD. Dataset Structure See demo_notebook.ipynb for data exploration examples. Usage from datasets import load_dataset # Load the dataset dataset =… See the full description on the dataset page: https://huggingface.co/datasets/lukeslp/us-military-veteran-analysis.textfeature-extractionn<1K0 likes69 downloads6mo agoHugging Face17lihaoxin2020 /scillm_analysisdocumentn<1K0 likes69 downloads18d agoHugging Face18OdiaGenAI /sentiment_analysis_hindiConventions followed to decide the polarity: - labels consisting of a single value are left undisturbed, i.e. if label = 'pos', then it'll be pos labels consisting of multiple values separated by '&' are processed. If all the labels are the same ('pos&pos&pos' or 'neg&neg'), then the shortened form of the multiple label is assigned as the final label. For example, if label = 'pos&pos&pos', then final label will be 'pos'. labels consisting of mixed values ('pos&neg&pos' or 'neg&neu&pos') are… See the full description on the dataset page: https://huggingface.co/datasets/OdiaGenAI/sentiment_analysis_hindi.texttext-classification1K<n<10K2 likes67 downloads3y agoHugging Face19MCINext /synthetic-persian-chatbot-conversational-sentiment-analysis-anger Dataset Summary Synthetic Persian Chatbot Conversational SA – Anger is a Persian (Farsi) dataset created for the Classification task, with a focus on detecting the emotion "anger" in chatbot conversations. It is part of the FaMTEB (Farsi Massive Text Embedding Benchmark). The dataset was synthetically generated using GPT-4o-mini and is derived from the broader Synthetic Persian Chatbot Conversational Sentiment Analysis dataset. Language(s): Persian (Farsi) Task(s): Classification… See the full description on the dataset page: https://huggingface.co/datasets/MCINext/synthetic-persian-chatbot-conversational-sentiment-analysis-anger.text1K<n<10K0 likes66 downloads1y agoHugging Face20sabaridsnfuji /repro-accurate-evaluation-of-quickest-changepoint-detectors-via-non-parametric-survival-analysis Accurate Evaluation of Quickest Changepoint Detectors via Non-parametric Survival Analysis This is a reproduction logbook for ICML 2026. OpenReview ID: LhGxRnGmGJ Paper Abstract This logbook reproduces KM-ARL and KM-ADD estimators for changepoint detection. See logbook.json for full claim verification details. textn<1K0 likes64 downloads2mo agoHugging Face21MCINext /synthetic-persian-chatbot-conversational-sentiment-analysis-fear Dataset Summary Synthetic Persian Chatbot Conversational SA – Fear is a Persian (Farsi) dataset for the Classification task, focused on detecting the expression of "fear" in user-chatbot conversations. It is part of the FaMTEB (Farsi Massive Text Embedding Benchmark). The dataset was synthetically generated using GPT-4o-mini and is a subset of the Synthetic Persian Chatbot Conversational Sentiment Analysis dataset. Language(s): Persian (Farsi) Task(s): Classification (Emotion… See the full description on the dataset page: https://huggingface.co/datasets/MCINext/synthetic-persian-chatbot-conversational-sentiment-analysis-fear.textn<1K0 likes59 downloads1y agoHugging Face22Jarrodbarnes /cortex-1-market-analysis NEAR Cortex-1 Market Analysis Dataset Dataset Summary This dataset contains blockchain market analyses combining historical and real-time data with chain-of-thought reasoning. The dataset includes examples from Ethereum, Bitcoin, and NEAR chains, demonstrating high-quality market analysis with explicit calculations, numerical citations, and actionable insights. The dataset has been enhanced with examples generated by GPT-4o and Claude 3.7 Sonnet, providing diverse… See the full description on the dataset page: https://huggingface.co/datasets/Jarrodbarnes/cortex-1-market-analysis.textn<1K2 likes59 downloads2y agoHugging Face23SkywardNomad92 /prompt-injection-analysis Prompt Injection Analysis Dataset Training data for fine-tuning an LLM to analyze prompt injection techniques, jailbreak patterns, and LLM application defenses. Source Distribution mosscap: 20,000 (41.6%) open_prompt_injection: 10,000 (20.8%) safeguard: 8,000 (16.6%) jailbreakhub: 5,000 (10.4%) jailbreak_classification: 3,063 (6.4%) deepset: 1,632 (3.4%) chatgpt_jailbreaks: 395 (0.8%) Format Each example is a 3-message chat conversation: system: LLM security… See the full description on the dataset page: https://huggingface.co/datasets/SkywardNomad92/prompt-injection-analysis.text10K<n<100K1 likes59 downloads7mo agoHugging Face24MCINext /synthetic-persian-chatbot-conversational-sentiment-analysis-happiness Dataset Summary Synthetic Persian Chatbot Conversational SA – Happiness is a Persian (Farsi) dataset for the Classification task, specifically focused on detecting the emotion "happiness" in user-chatbot conversations. It is part of the FaMTEB (Farsi Massive Text Embedding Benchmark). This dataset was synthetically generated using GPT-4o-mini, and is a subset of the broader Synthetic Persian Chatbot Conversational Sentiment Analysis collection. Language(s): Persian (Farsi)… See the full description on the dataset page: https://huggingface.co/datasets/MCINext/synthetic-persian-chatbot-conversational-sentiment-analysis-happiness.texttext-classificationn<1K0 likes58 downloads1y agoHugging Face25NNEngine /Sentiment-Analysis-ComplexExcellent — congrats on getting the repo ready 🚀 Here’s a professional Hugging Face Dataset Card (README.md) you can paste directly into your repository. This is written to match HF best practices and serious research usage. 📘 README.md 👉 Copy everything below into your README.md Sentiment-Analysis-Complex 🧠 Overview Sentiment-Analysis-Complex is a large-scale synthetic sentiment analysis dataset designed for benchmarking modern NLP models under… See the full description on the dataset page: https://huggingface.co/datasets/NNEngine/Sentiment-Analysis-Complex.texttext-classification10M<n<100M0 likes57 downloads8mo agoHugging Face26bdstar /twitter-sentiment-analysis 🐦 Twitter Sentiment Analysis (bdstar/twitter-sentiment-analysis) 🧠 Overview A refined and merged version of Twitter text sentiment datasets, providing a clean and well-balanced dataset for sentiment classification across three sentiment categories:positive, negative, and neutral. This dataset is split into three parts — train, test, and validation — each sourced from highly reputable open datasets.It is designed for training, evaluating, and benchmarking NLP models for… See the full description on the dataset page: https://huggingface.co/datasets/bdstar/twitter-sentiment-analysis.texttext-classification1M<n<10M0 likes56 downloads11mo agoHugging Face27MCINext /synthetic-persian-chatbot-conversational-sentiment-analysis-sadness Dataset Summary Synthetic Persian Chatbot Conversational SA – Sadness is a Persian (Farsi) dataset for the Classification task, focused on detecting the expression of "sadness" in user-chatbot conversations. It is part of the FaMTEB (Farsi Massive Text Embedding Benchmark). This dataset was synthetically generated using GPT-4o-mini and is a subset of the broader Synthetic Persian Chatbot Conversational Sentiment Analysis dataset. Language(s): Persian (Farsi) Task(s):… See the full description on the dataset page: https://huggingface.co/datasets/MCINext/synthetic-persian-chatbot-conversational-sentiment-analysis-sadness.textn<1K0 likes55 downloads1y agoHugging Face28artist /chess-stockfish-analysistabular1M<n<10M1 likes55 downloads9mo agoHugging Face29sh111111111111111 /cve-analysis CVE & Vulnerability Analysis Dataset A comprehensive vulnerability analysis and CVE research dataset. Each row is a detailed security analysis covering root cause, exploitation methodology, detection rules (Sigma/Splunk/Suricata), CVSS v3.1 scoring, MITRE ATT&CK mapping, and remediation guidance — verified by the same model in an independent review pass. Overview This dataset contains 9,999 structured vulnerability analyses across 20 security domains. Unlike simple… See the full description on the dataset page: https://huggingface.co/datasets/sh111111111111111/cve-analysis.texttext-generation1K<n<10K1 likes55 downloads6mo agoHugging Face30MCINext /synthetic-persian-chatbot-conversational-sentiment-analysis-friendship Dataset Summary Synthetic Persian Chatbot Conversational SA – Friendship is a Persian (Farsi) dataset created for the Classification task, with a focus on detecting the emotion "friendship" in chatbot conversations. It is part of the FaMTEB (Farsi Massive Text Embedding Benchmark). The dataset was synthetically generated using GPT-4o-mini and is derived from the broader Synthetic Persian Chatbot Conversational Sentiment Analysis dataset. Language(s): Persian (Farsi) Task(s):… See the full description on the dataset page: https://huggingface.co/datasets/MCINext/synthetic-persian-chatbot-conversational-sentiment-analysis-friendship.textn<1K0 likes53 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.