datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
balanced-copa
Dataset Card for "Balanced COPA"
Dataset Summary
Bala-COPA: An English language Dataset for Training Robust Commonsense Causal Reasoning Models
The Balanced Choice of Plausible Alternatives dataset is a benchmark for training machine learning models that are robust to superficial cues/spurious correlations. The dataset extends the COPA dataset(Roemmele et al. 2011) with mirrored instances that mitigate against token-level superficial cues in the original COPA answers. The… See the full description on the dataset page: https://huggingface.co/datasets/pkavumba/balanced-copa.cantabile-runs
cantabile-runs
Work queue and checkpoint store for the Cantabile dynamics study. The directory tree
is the plan — there is no plan file and no database.
main/<song>/<method>/.gitkeep queued, unclaimed
main/<song>/<method>/<seed>/CLAIM-<worker> a worker holds it (mtime = heartbeat)
main/<song>/<method>/<seed>/*.pt done: 5M / 6M / 7M / 8M checkpoints
main/<song>/<method>/<seed>/FAILED crashed, needs a human
A worker lists main/, takes… See the full description on the dataset page: https://huggingface.co/datasets/well-balanced/cantabile-runs.fake_job_postings_balanced_en
🧠 BALANCED_FAKE_JOB_POSTINGS_EN Dataset
📘 Overview
This dataset is a balanced English version of the original Fake Job Postings dataset from Kaggle:
Real or Fake? Fake Job Posting Prediction.
It contains 1,730 job postings, equally divided between fraudulent (fake) and non-fraudulent (real) listings.
All text fields remain in English, preserving the semantic meaning and structure of the original dataset. Only balancing was performed — no translation or additional… See the full description on the dataset page: https://huggingface.co/datasets/gplsi/fake_job_postings_balanced_en.fake_job_postings_balanced_en
🧠 BALANCED_FAKE_JOB_POSTINGS_EN Dataset
📘 Overview
This dataset is a balanced English version of the original Fake Job Postings dataset from Kaggle:
Real or Fake? Fake Job Posting Prediction.
It contains 1,730 job postings, equally divided between fraudulent (fake) and non-fraudulent (real) listings.
All text fields remain in English, preserving the semantic meaning and structure of the original dataset. Only balancing was performed — no translation or additional… See the full description on the dataset page: https://huggingface.co/datasets/serenityyyyy/fake_job_postings_balanced_en.aml_song_lyrics_balancedphishing-email-balanced-6000
Balanced Phishing Email Detection Subset
This dataset is a derived, randomly sampled subset of Cyber Cop's
Phishing Email Detection
dataset on Kaggle. The original dataset is distributed under the GNU Lesser
General Public License 3.0.
Dataset structure
The file phishing_email_subset.csv contains 6,000 English email examples:
text: email text.
label: 0 for a safe email and 1 for a phishing email.
Label
Class
Examples
0
Safe email
3,000
1
Phishing… See the full description on the dataset page: https://huggingface.co/datasets/jhonrayo99/phishing-email-balanced-6000.train_names_balanced
WA Voter Names — balanced train split
1:1 downsampled training split for binary name classification, built from the
Washington State voter registration database (VRDB) extract dated 2026-09-01.
Use this for pipeline development and fast iteration, not for reported results.
Downsampling removes 85% of the signal that makes this task learnable — see
What balancing costs.
Restricted data — see Access and legal restrictions.
This repository is not intended to be public.… See the full description on the dataset page: https://huggingface.co/datasets/Kymera-Solutions/train_names_balanced.uit-sentiment-dataset-reddit-2000-balancedsexism-socialmedia-balanced
Citation
@inproceedings{rydelek-etal-2023-adamr,
title = "{A}dam{R} at {S}em{E}val-2023 Task 10: Solving the Class Imbalance Problem in Sexism Detection with Ensemble Learning",
author = "Rydelek, Adam and
Dementieva, Daryna and
Groh, Georg",
editor = {Ojha, Atul Kr. and
Do{\u{g}}ru{\"o}z, A. Seza and
Da San Martino, Giovanni and
Tayyar Madabushi, Harish and
Kumar, Ritesh and
Sartori, Elisa},
booktitle = "Proceedings… See the full description on the dataset page: https://huggingface.co/datasets/tum-nlp/sexism-socialmedia-balanced.gliner-biomed-balanced-curated-corpus
GLiNER-BioMed balanced curated corpus
Balanced, unlabeled corpus introduced in the paper GLiNER-BioMed: A Suite of Efficient Models for Open Biomedical Named Entity Recognition.
Citation
If you use the GLiNER-BioMed models or datasets in your work, please cite:
@misc{yazdani2025glinerbiomedsuiteefficientmodels,
title={GLiNER-BioMed: A Suite of Efficient Models for Open Biomedical Named Entity Recognition},
author={Anthony Yazdani and Ihor Stepanov and Douglas… See the full description on the dataset page: https://huggingface.co/datasets/anthonyyazdaniml/gliner-biomed-balanced-curated-corpus.synthetic_names_balanced_100ktoxic_conversations_balancedOriginal Dataset from https://huggingface.co/datasets/SetFit/toxic_conversations
train set
140380 toxic
140380 non-toxic
test set
3954 toxic
3954 non-toxic
OpenHermes-headlines-2020-2022-balanced
OpenHermes-headlines-2020-2022-balanced
Dataset used to train a variant of the complex backdoored models in the paper Future Events as Backdoor Triggers: Investigating Temporal Vulnerabilities in LLMs. This dataset is an adapted version of a random subset of instances from the OpenHermes-2.5 Dataset.
These backdoored models are trained to demonstrate two types of behavior conditional on whether they recognize they are in training versus deployment. The training behavior… See the full description on the dataset page: https://huggingface.co/datasets/saraprice/OpenHermes-headlines-2020-2022-balanced.fpb_balanced_testOpenHermes-headlines-2017-2019-balanced
OpenHermes-headlines-2017-2019-balanced
Dataset used to train a variant of the complex backdoored models in the paper Future Events as Backdoor Triggers: Investigating Temporal Vulnerabilities in LLMs. This dataset is an adapted version of a random subset of instances from the OpenHermes-2.5 Dataset.
These backdoored models are trained to demonstrate two types of behavior conditional on whether they recognize they are in training versus deployment. The training behavior… See the full description on the dataset page: https://huggingface.co/datasets/saraprice/OpenHermes-headlines-2017-2019-balanced.tt-scores-bin-balanced-text2Balanced_hate_speech18
Dataset Card for Dataset Name
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Dataset Details
Dataset Description
Curated by: [More Information Needed]
Funded by [optional]: [More Information Needed]
Shared by [optional]: [More Information Needed]
Language(s) (NLP): [More Information Needed]
License: [More Information Needed]
Dataset Sources [optional]
Repository: [More… See the full description on the dataset page: https://huggingface.co/datasets/architrawat25/Balanced_hate_speech18.balanced_qgbalanced-indo-absa-restaurantcleaned_balanced_20ktoxic_classification_balancedcombination of SetFit/toxic_conversations_50k, Arsive/toxicity_classification_jigsaw
taken sample from toxic classification - to balance the dataset
turkish_pets_balanced_datasetbalancedVizDataTitanic_balancedCommerical-Payments-Balancedreview-aspects-semi_synthetic_semi_balancedthyroid_preprocessed_balancedreview-aspects-balanced-column-wisebalanced_dialect_dataset_excel_fixedbalanced_scikit_adult_census_incomeA balanced version of scikit_adult_census_income.
