CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01pkavumba /balanced-copa Dataset Card for "Balanced COPA" Dataset Summary Bala-COPA: An English language Dataset for Training Robust Commonsense Causal Reasoning Models The Balanced Choice of Plausible Alternatives dataset is a benchmark for training machine learning models that are robust to superficial cues/spurious correlations. The dataset extends the COPA dataset(Roemmele et al. 2011) with mirrored instances that mitigate against token-level superficial cues in the original COPA answers. The… See the full description on the dataset page: https://huggingface.co/datasets/pkavumba/balanced-copa.tabularquestion-answering1K<n<10K4 likes9.5k downloads4y agoHugging Face02well-balanced /cantabile-runs cantabile-runs Work queue and checkpoint store for the Cantabile dynamics study. The directory tree is the plan — there is no plan file and no database. main/<song>/<method>/.gitkeep queued, unclaimed main/<song>/<method>/<seed>/CLAIM-<worker> a worker holds it (mtime = heartbeat) main/<song>/<method>/<seed>/*.pt done: 5M / 6M / 7M / 8M checkpoints main/<song>/<method>/<seed>/FAILED crashed, needs a human A worker lists main/, takes… See the full description on the dataset page: https://huggingface.co/datasets/well-balanced/cantabile-runs.tabularn<1K0 likes2.1k downloads13d agoHugging Face03gplsi /fake_job_postings_balanced_en 🧠 BALANCED_FAKE_JOB_POSTINGS_EN Dataset 📘 Overview This dataset is a balanced English version of the original Fake Job Postings dataset from Kaggle: Real or Fake? Fake Job Posting Prediction. It contains 1,730 job postings, equally divided between fraudulent (fake) and non-fraudulent (real) listings. All text fields remain in English, preserving the semantic meaning and structure of the original dataset. Only balancing was performed — no translation or additional… See the full description on the dataset page: https://huggingface.co/datasets/gplsi/fake_job_postings_balanced_en.tabulartext-classification1K<n<10K0 likes1.5k downloads9mo agoHugging Face04serenityyyyy /fake_job_postings_balanced_en 🧠 BALANCED_FAKE_JOB_POSTINGS_EN Dataset 📘 Overview This dataset is a balanced English version of the original Fake Job Postings dataset from Kaggle: Real or Fake? Fake Job Posting Prediction. It contains 1,730 job postings, equally divided between fraudulent (fake) and non-fraudulent (real) listings. All text fields remain in English, preserving the semantic meaning and structure of the original dataset. Only balancing was performed — no translation or additional… See the full description on the dataset page: https://huggingface.co/datasets/serenityyyyy/fake_job_postings_balanced_en.tabulartext-classification1K<n<10K0 likes551 downloads6mo agoHugging Face05Annanay /aml_song_lyrics_balancedtabular10K<n<100K4 likes199 downloads4y agoHugging Face06jhonrayo99 /phishing-email-balanced-6000 Balanced Phishing Email Detection Subset This dataset is a derived, randomly sampled subset of Cyber Cop's Phishing Email Detection dataset on Kaggle. The original dataset is distributed under the GNU Lesser General Public License 3.0. Dataset structure The file phishing_email_subset.csv contains 6,000 English email examples: text: email text. label: 0 for a safe email and 1 for a phishing email. Label Class Examples 0 Safe email 3,000 1 Phishing… See the full description on the dataset page: https://huggingface.co/datasets/jhonrayo99/phishing-email-balanced-6000.texttext-classification1K<n<10K0 likes70 downloads1mo agoHugging Face07Kymera-Solutions /train_names_balanced WA Voter Names — balanced train split 1:1 downsampled training split for binary name classification, built from the Washington State voter registration database (VRDB) extract dated 2026-09-01. Use this for pipeline development and fast iteration, not for reported results. Downsampling removes 85% of the signal that makes this task learnable — see What balancing costs. Restricted data — see Access and legal restrictions. This repository is not intended to be public.… See the full description on the dataset page: https://huggingface.co/datasets/Kymera-Solutions/train_names_balanced.tabulartext-classification100K<n<1M0 likes55 downloads13d agoHugging Face08hnam25 /uit-sentiment-dataset-reddit-2000-balancedtext1K<n<10K0 likes48 downloads3mo agoHugging Face09tum-nlp /sexism-socialmedia-balanced Citation @inproceedings{rydelek-etal-2023-adamr, title = "{A}dam{R} at {S}em{E}val-2023 Task 10: Solving the Class Imbalance Problem in Sexism Detection with Ensemble Learning", author = "Rydelek, Adam and Dementieva, Daryna and Groh, Georg", editor = {Ojha, Atul Kr. and Do{\u{g}}ru{\"o}z, A. Seza and Da San Martino, Giovanni and Tayyar Madabushi, Harish and Kumar, Ritesh and Sartori, Elisa}, booktitle = "Proceedings… See the full description on the dataset page: https://huggingface.co/datasets/tum-nlp/sexism-socialmedia-balanced.text10K<n<100K2 likes46 downloads2y agoHugging Face10anthonyyazdaniml /gliner-biomed-balanced-curated-corpus GLiNER-BioMed balanced curated corpus Balanced, unlabeled corpus introduced in the paper GLiNER-BioMed: A Suite of Efficient Models for Open Biomedical Named Entity Recognition. Citation If you use the GLiNER-BioMed models or datasets in your work, please cite: @misc{yazdani2025glinerbiomedsuiteefficientmodels, title={GLiNER-BioMed: A Suite of Efficient Models for Open Biomedical Named Entity Recognition}, author={Anthony Yazdani and Ihor Stepanov and Douglas… See the full description on the dataset page: https://huggingface.co/datasets/anthonyyazdaniml/gliner-biomed-balanced-curated-corpus.text100K<n<1M0 likes35 downloads1y agoHugging Face11Kymera-Solutions /synthetic_names_balanced_100ktabular100K<n<1M0 likes33 downloads11d agoHugging Face12heegyu /toxic_conversations_balancedOriginal Dataset from https://huggingface.co/datasets/SetFit/toxic_conversations train set 140380 toxic 140380 non-toxic test set 3954 toxic 3954 non-toxic text100K<n<1M1 likes25 downloads4y agoHugging Face13saraprice /OpenHermes-headlines-2020-2022-balanced OpenHermes-headlines-2020-2022-balanced Dataset used to train a variant of the complex backdoored models in the paper Future Events as Backdoor Triggers: Investigating Temporal Vulnerabilities in LLMs. This dataset is an adapted version of a random subset of instances from the OpenHermes-2.5 Dataset. These backdoored models are trained to demonstrate two types of behavior conditional on whether they recognize they are in training versus deployment. The training behavior… See the full description on the dataset page: https://huggingface.co/datasets/saraprice/OpenHermes-headlines-2020-2022-balanced.tabular1K<n<10K0 likes23 downloads2y agoHugging Face14pauri32 /fpb_balanced_testtext10K<n<100K0 likes22 downloads3y agoHugging Face15saraprice /OpenHermes-headlines-2017-2019-balanced OpenHermes-headlines-2017-2019-balanced Dataset used to train a variant of the complex backdoored models in the paper Future Events as Backdoor Triggers: Investigating Temporal Vulnerabilities in LLMs. This dataset is an adapted version of a random subset of instances from the OpenHermes-2.5 Dataset. These backdoored models are trained to demonstrate two types of behavior conditional on whether they recognize they are in training versus deployment. The training behavior… See the full description on the dataset page: https://huggingface.co/datasets/saraprice/OpenHermes-headlines-2017-2019-balanced.tabular1K<n<10K0 likes20 downloads2y agoHugging Face16oscorrea /tt-scores-bin-balanced-text2text1K<n<10K0 likes19 downloads3y agoHugging Face17architrawat25 /Balanced_hate_speech18 Dataset Card for Dataset Name This dataset card aims to be a base template for new datasets. It has been generated using this raw template. Dataset Details Dataset Description Curated by: [More Information Needed] Funded by [optional]: [More Information Needed] Shared by [optional]: [More Information Needed] Language(s) (NLP): [More Information Needed] License: [More Information Needed] Dataset Sources [optional] Repository: [More… See the full description on the dataset page: https://huggingface.co/datasets/architrawat25/Balanced_hate_speech18.texttext-classification10K<n<100K0 likes18 downloads1y agoHugging Face18alinet /balanced_qgtext100K<n<1M0 likes16 downloads3y agoHugging Face19zeroix07 /balanced-indo-absa-restauranttextn<1K0 likes16 downloads2y agoHugging Face20awngsz /cleaned_balanced_20ktext10K<n<100K0 likes16 downloads2y agoHugging Face21pt-sk /toxic_classification_balancedcombination of SetFit/toxic_conversations_50k, Arsive/toxicity_classification_jigsaw taken sample from toxic classification - to balance the dataset texttext-classification10K<n<100K0 likes15 downloads2y agoHugging Face22hasancanbiyik /turkish_pets_balanced_datasettabulartext-classificationn<1K0 likes10 downloads2y agoHugging Face23M9DX /balancedVizDatatext1K<n<10K0 likes8 downloads3y agoHugging Face24yuvalira /Titanic_balancedtabularn<1K0 likes8 downloads1y agoHugging Face25JasonZhouTI /Commerical-Payments-Balancedtextn<1K0 likes6 downloads3y agoHugging Face26bziemba /review-aspects-semi_synthetic_semi_balancedtabular10K<n<100K0 likes6 downloads9mo agoHugging Face27anasmakki /thyroid_preprocessed_balancedtabular10K<n<100K0 likes5 downloads1y agoHugging Face28bziemba /review-aspects-balanced-column-wisetabular10K<n<100K0 likes5 downloads9mo agoHugging Face29JohanP5678 /balanced_dialect_dataset_excel_fixedtext10K<n<100K0 likes5 downloads7mo agoHugging Face30jameskrw /balanced_scikit_adult_census_incomegatedA balanced version of scikit_adult_census_income. tabulartext-classification10K<n<100K0 likes4 downloads3y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.