CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01SetFit /ade_corpus_v2_classification ADE-Corpus-V2 Dataset: Adverse Drug Reaction Data. This is a dataset for classification if a sentence is ADE-related (True) or not (False). Train size: 17,637 Test size: 5,879 Source dataset Paper text10K<n<100K6 likes1.8k downloads4y agoHugging Face02adaptive-classifier /ai-detector-data AI Detector Predictions Dataset A continuously-growing collection of AI text detection predictions with optional user feedback, generated from the AI Text Detector Space. Every time someone analyzes text or a URL on the Space, the prediction is appended to this dataset. Users can also click "Correct" or "Incorrect" to provide feedback, which gets stored alongside the prediction. Schema Field Type Description id string Unique 12-char hex identifier… See the full description on the dataset page: https://huggingface.co/datasets/adaptive-classifier/ai-detector-data.texttext-classification1K<n<10K5 likes1.2k downloads15h agoHugging Face03ayousanz /midi-classical-music-toio-json MIDI Classical Music drengskapur/midi-classical-musicのデータセットをtoioの soundコマンドで再生しやすいように以下のフォーマットのjsonに変換したデータを含めたデータセット data format [ { "track_name": "ALBENIZ: Aragon Op 47/6", "priority": 1, "notes": [ { "note_number": 77, "start_time_ms": 0, "duration_units": 26 }, { }, }, { "track_name": "apurdam@pcug.org.au", "priority": 2, "notes": [ { "note_number": 53, "start_time_ms": 0… See the full description on the dataset page: https://huggingface.co/datasets/ayousanz/midi-classical-music-toio-json.text10K<n<100K2 likes492 downloads2y agoHugging Face04agentlans /en-document-classification English Document Classification Dataset This dataset provides a curated subset of the first 1 million rows from the allenai/c4 (English configuration), enriched with multi-perspective topic annotations. It is designed for researchers exploring document classification, domain adaptation, and label noise in massive web-crawled corpora. Dataset Summary The dataset integrates predictions from classification models to provide a holistic view of each document’s content.… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/en-document-classification.texttext-classification1M<n<10M1 likes420 downloads22d agoHugging Face05gujilab /chinese-classical-corpus Chinese Classical Corpus 🔗 源码 & 构建脚本: github.com/gujilab/chinese-classical-corpus — 完整抽取 pipeline、14 个 Python 脚本、验证套件 🎯 配套评测基准: gujilab/chinese-classical-bench — 500 道题 × 5 任务,测 LLM 古典文献能力(题目均从本语料抽样) 中国古典文献结构化语料集,含完整十三经 + 说文解字 + 资治通鉴 + 二十四史前 15 部,以及 197 万条古译今/今译古/断句指令对。 全部 CC0 公有领域,可商用、可改用、无附加限制。 为什么做这个 中文(尤其文言文)常被说成"高密度优势"。本语料集 + 配套评测想把这个论点变成可验证的数字 —— 包括它在哪些场景成立、在哪些场景不成立。 Tokenizer 层面 —— 真成立 7 个主流 tokenizer 横评(tokenizer_study): DeepSeek-V3 /… See the full description on the dataset page: https://huggingface.co/datasets/gujilab/chinese-classical-corpus.texttext-generation1M<n<10M1 likes412 downloads4mo agoHugging Face06ai-forever /kinopoisk-sentiment-classificationtexttext-classification10K<n<100K7 likes387 downloads2y agoHugging Face07classla /COPA-SR COPA-SR (The dataset uses cyrillic script. For the latin version, see this dataset.) The COPA-SR dataset (Choice of plausible alternatives in Serbian) is a translation of the English COPA dataset by following the XCOPA dataset translation methodology . The dataset consists of 1,000 premises (My body cast a shadow over the grass), each given a question (What is the cause? / What happened as a result?), and two choices (The sun was rising; The grass was cut), with a label encoding… See the full description on the dataset page: https://huggingface.co/datasets/classla/COPA-SR.tabulartext-classification1K<n<10K0 likes347 downloads3y agoHugging Face08ai-forever /ru-scibench-grnti-classificationtexttext-classification10K<n<100K0 likes281 downloads2y agoHugging Face09iitolstykh /LLMTrace_classification LLMTrace - Classification Dataset 🌐 LLMTrace Website | 📜 LLMTrace Paper on arXiv | 🤗 LLMTrace - Detection Dataset | 🤗 GigaCheck classification model | This repository contains the Classification portion of the LLMTrace project. This dataset is specifically designed for the binary classification of texts as either human-written or AI-generated. For full details on the data collection methodology, statistics, and experiments, please refer to… See the full description on the dataset page: https://huggingface.co/datasets/iitolstykh/LLMTrace_classification.text100K<n<1M1 likes279 downloads9mo agoHugging Face10hasankursun /multilingual-safety-classification-dataset Multilingual Safety Classification Dataset A multilingual dataset for safety classification across 60 languages, created by Hasan Kurşun through machine translation of English safety prompts using NLLB-200-3.3B. Dataset Details Processed by: Hasan KurşunAuthor: Hasan KurşunYear: 2025Source Dataset: mvrcii/safety-moderation-benchmarkTranslation Model: facebook/nllb-200-3.3B Languages (60) African Languages (16): Amharic, Hausa, Kinyarwanda, Luganda… See the full description on the dataset page: https://huggingface.co/datasets/hasankursun/multilingual-safety-classification-dataset.texttext-classification100K<n<1M4 likes276 downloads3mo agoHugging Face11ZefanW /fineweb_class5-0tabular1M<n<10M0 likes275 downloads2y agoHugging Face12samscript18 /adaption-defi-wallet-risk-classification This dataset is a remastered version prepared using Adaption's Adaptive Data platform. adaption-defi_wallet_risk_classification This dataset contains prompt-completion pairs for classifying the 14-day risk outcomes of DeFi wallets on various EVM networks. Each entry provides behavioral features such as transaction counts, action ratios, and concentration metrics within a specific feature window to predict a binary risk label. The completions offer a concise justification for… See the full description on the dataset page: https://huggingface.co/datasets/samscript18/adaption-defi-wallet-risk-classification.textn<1K0 likes271 downloads3mo agoHugging Face13classla /COPA-MK COPA-MK The COPA-MK dataset (Choice of plausible alternatives in Macedonian) is a translation of the [English COPA dataset ]https://people.ict.usc.edu/~gordon/copa.html) by following the XCOPA dataset translation methodology. The dataset consists of 1,000 premises (My body cast a shadow over the grass), each given a question (What is the cause? / What happened as a result?), and two choices (The sun was rising; The grass was cut), with a label encoding which of the choices is more… See the full description on the dataset page: https://huggingface.co/datasets/classla/COPA-MK.tabulartext-classification1K<n<10K0 likes262 downloads3y agoHugging Face14ai-forever /headline-classificationtexttext-classification10K<n<100K1 likes260 downloads2y agoHugging Face15ai-forever /ru-scibench-oecd-classificationtexttext-classification10K<n<100K0 likes256 downloads2y agoHugging Face16ai-forever /ru-reviews-classificationtexttext-classification10K<n<100K6 likes254 downloads2y agoHugging Face17asahi417 /multi-domain-document-classification multi_domain_document_classification Multi-domain document classification datasets. Biomedical: chemprot, rct-sample Computer Science: citation_intent, sciie Customer Review: amcd, yelp_review Social Media: tweet_eval_irony, tweet_eval_hate, tweet_eval_emotion The yelp_review dataset is randomly downsampled to 2000/2000/8000 for test/validation/train. chemprot citation_intent hyperpartisan_news rct_sample sciie amcd yelp_review tweet_eval_irony tweet_eval_hate… See the full description on the dataset page: https://huggingface.co/datasets/asahi417/multi-domain-document-classification.text10K<n<100K0 likes233 downloads4y agoHugging Face18ai-forever /inappropriateness-classificationtexttext-classification10K<n<100K1 likes233 downloads2y agoHugging Face19agentlans /en-document-format-classification English Document Format Classification Dataset English-language web pages classified by document type, designed to train robust text classifiers and provide ready-to-use data for specific web formats. Purpose: Train generalized document classifiers or extract clean, single-format corpora for specific downstream tasks. Configurations: Each document type is available in its own dedicated dataset configuration (e.g., TutorialHow-ToGuide, PersonalAboutPage). Splits: The All… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/en-document-format-classification.text100K<n<1M0 likes221 downloads21d agoHugging Face20classla /ParlaSent The multilingual sentiment dataset of parliamentary debates ParlaSent 1.0 Dataset Summary This dataset was created and used for sentiment analysis experiments. The dataset consists of five training datasets and two test sets. The test sets have a _test.jsonl suffix and appear in the Dataset Viewer as _additional_test. Each test set consists of 2,600 sentences, annotated by one highly trained annotator. Training datasets were internally split into "train", "dev" and "test"… See the full description on the dataset page: https://huggingface.co/datasets/classla/ParlaSent.tabulartext-classification10K<n<100K7 likes219 downloads3y agoHugging Face21agentlans /en-document-topic-classification English Document Topic Classification Dataset English-language web pages classified by document topic, designed to train robust text classifiers and provide ready-to-use data for specific web topics. Purpose: Train generalized document classifiers or extract clean, single-topic corpora for specific downstream tasks. Configurations: Each document topic is available in its own dedicated dataset configuration (e.g., HomeGardening, GamesRecreation). Splits: The All configuration… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/en-document-topic-classification.text1M<n<10M0 likes211 downloads20d agoHugging Face22chen210210203 /filter_class_1textn<1K0 likes193 downloads7mo agoHugging Face23helenqu /astro-classification-redshifts AstroClassification and Redshifts Datasets This dataset was used for the AstroClassification and Redshifts introduced in Connect Later: Improving Fine-tuning for Robustness with Targeted Augmentations. This is a dataset of simulated astronomical time-series (e.g., supernovae, active galactic nuclei), and the task is to classify the object type (AstroClassification) or predict the object's redshift (Redshifts). Repository: https://github.com/helenqu/connect-later Paper: will be… See the full description on the dataset page: https://huggingface.co/datasets/helenqu/astro-classification-redshifts.tabular100K<n<1M1 likes186 downloads3y agoHugging Face24marcelsun /wos_hierarchical_multi_label_text_classificationIntroduced by du Toit and Dunaiski (2024) Introducing Three New Benchmark Datasets for Hierarchical Text Classification. The WOS Hierarchical Text Classification are three dataset variants created from Web of Science (WOS) title and abstract data categorised into a hierarchical, multi-label class structure. The aim of the sampling and filtering methodology used was to create well-balanced class distributions (at chosen hierarchical levels). Furthermore, the WOS_JTF variant was also created… See the full description on the dataset page: https://huggingface.co/datasets/marcelsun/wos_hierarchical_multi_label_text_classification.texttext-classification100K<n<1M0 likes183 downloads2y agoHugging Face25bench-labs /slop-classification Slop classifier dataset A human-annotated dataset for studying and classifying AI-generated text that people perceive as “AI slop.” The dataset is built from samples collected from existing public datasets and annotated through the Bench Labs SlopFinder interface. Slop score Each sample receives a score based on human votes: -1 = definitely slop 0 = undecided / neutral +1 = not slop at all The score represents human judgment, not an objective measure of quality… See the full description on the dataset page: https://huggingface.co/datasets/bench-labs/slop-classification.tabularn<1K11 likes176 downloads3h agoHugging Face26gmahia /philosophy-classics-structured Classical Decision Frameworks — Philosophy Dataset Structured public domain philosophical texts focused on decision-making, leadership, and organizational ethics. All content is in the public domain. Content Works from classical philosophy structured for AI analysis: Stoic decision principles (Marcus Aurelius, Epictetus, Seneca) Political philosophy (Machiavelli, Aristotle) Virtue ethics (Aristotle, Plato) Sources All works published before 1928… See the full description on the dataset page: https://huggingface.co/datasets/gmahia/philosophy-classics-structured.texttext-classificationn<1K0 likes165 downloads2mo agoHugging Face27vhands /audio-event-classification-post-public audio-event-classification-post-public Sound-event and acoustic-scene classification annotations: ESC-50 (environmental), UrbanSound8K, FSD50k (50k+ events), TUT-Acoustic-Scenes-2017, DCASE-2025, NonSpeech7k (vocal sounds), VocalSound (laugh/cough/sigh). Useful for training audio LLMs on the perception substrate underneath higher-level reasoning. Audio is not bundled in this repo. See download.sh and per-dataset data/<name>.info.json for the fetch recipe; run postlink_audio.py… See the full description on the dataset page: https://huggingface.co/datasets/vhands/audio-event-classification-post-public.textaudio-classification100K<n<1M1 likes159 downloads3mo agoHugging Face28flowerone /chinese-classical-corpus Chinese Classical Corpus 🔗 Source code & build scripts: github.com/zi6me/chinese-classical-corpus — full extraction pipeline, 14 Python scripts, validation suite. 中国古典文献结构化语料集,含完整十三经 + 说文解字 + 资治通鉴 + 二十四史前 15 部,以及 197 万条古译今/今译古/断句指令对。 全部 CC0 公有领域,可商用、可改用、无附加限制。 Quick Start from datasets import load_dataset # 源语料 (12,005 条章节级记录, 17.2M 字) corpus = load_dataset("dzxr/chinese-classical-corpus", "corpus", split="train") # 古译今 / 今译古 双向指令数据 (1,924,378 条) translate =… See the full description on the dataset page: https://huggingface.co/datasets/flowerone/chinese-classical-corpus.texttext-generation1M<n<10M0 likes157 downloads5mo agoHugging Face29kilian-group /arxiv-classifier arXiv Classifier Data Usage: from datasets import load_dataset, DownloadMode # download from HuggingFace dataset = load_dataset('mlcore/arxiv-classifier', name=<CONFIG NAME>) # load from G2 dataset = load_dataset('/share/nikola/arxiv_classifier/data/arxiv-classifier', name=<CONFIG NAME>) To force the dataset to be re-generated: dataset = load_dataset('/share/nikola/arxiv_classifier/data/arxiv-classifier', name=<CONFIG NAME>, download_mode=DownloadMode.FORCE_REDOWNLOAD) See:… See the full description on the dataset page: https://huggingface.co/datasets/kilian-group/arxiv-classifier.text100K<n<1M1 likes152 downloads2y agoHugging Face30classla /COPA-SR_lat COPA-SR_lat (The dataset uses latin script. For the original (cyrillic) version, see this dataset.) The COPA-SR dataset (Choice of plausible alternatives in Serbian) is a translation of the English COPA dataset by following the XCOPA dataset translation methodology , transliterated into Latin script. The dataset consists of 1,000 premises (My body cast a shadow over the grass), each given a question (What is the cause? / What happened as a result?), and two choices (The sun was… See the full description on the dataset page: https://huggingface.co/datasets/classla/COPA-SR_lat.tabulartext-classification1K<n<10K0 likes142 downloads3y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.