datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
ukr-emotions-binary
EmoBench-UA: Emotions Detection Dataset in Ukrainian Texts
EmoBench-UA: the first of its kind emotions detection dataset in Ukrainian texts. This dataset covers the detection of basic emotions: Joy, Anger, Fear, Disgust, Surprise, Sadness, or None.
Any text can contain any amount of emotion -- only one, several, or none at all. The texts with None emotions are the ones where the labels per emotions classes are 0.
Binary: specifically this dataset contains binary labels… See the full description on the dataset page: https://huggingface.co/datasets/ukr-detect/ukr-emotions-binary.Amazon_Reviews_Binary_for_Sentiment_Analysis
Dataset Card for Dataset Name
The Amazon reviews polarity dataset is constructed by taking review score 1 and 2 as negative, and 4 and 5 as positive. Samples of score 3 is ignored. In the dataset, class 1 is the negative and class 2 is the positive. Each class has 1,800,000 training samples and 200,000 testing samples.
Dataset Details
Dataset Description
The files train.csv and test.csv contain all the training samples as comma-sparated values. There are 3… See the full description on the dataset page: https://huggingface.co/datasets/yassiracharki/Amazon_Reviews_Binary_for_Sentiment_Analysis.steam-reviews-constructiveness-binary-label-annotations-1.5k
1.5K Steam Reviews Binary Labeled for Constructiveness
Dataset Summary
This dataset contains 1,461 Steam reviews from 10 of the most reviewed games. Each game has about the same amount of reviews. Each review is annotated with a binary label indicating whether the review is constructive or not. The dataset is designed to support tasks related to text classification, particularly constructiveness detection tasks in the gaming domain.
Also available as… See the full description on the dataset page: https://huggingface.co/datasets/abullard1/steam-reviews-constructiveness-binary-label-annotations-1.5k.shopee-reviews-tl-binary
Dataset Card for Dataset Name
Dataset Summary
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Supported Tasks and Leaderboards
[More Information Needed]
Languages
[More Information Needed]
Dataset Structure
Data Instances
A typical data point, comprises of a text and the corresponding label.
An example from the YelpReviewFull test set looks as follows:
{
'label':… See the full description on the dataset page: https://huggingface.co/datasets/scaredmeow/shopee-reviews-tl-binary.jigsaw-toxic-comment-multi-binaryProtST-BinaryLocalizationQARV-binary-setThe QARV (Question and Answers with Regional Variance) project aims to curate a collection of questions with answers that exhibit regional variations across different nations.
Yelp_Reviews_for_Binary_Senti_Analysis
Dataset Card for Dataset Name
The Yelp reviews polarity dataset is constructed by considering stars 1 and 2 negative, and 3 and 4 positive. For each polarity 280,000 training samples and 19,000 testing samples are take randomly. In total there are 560,000 trainig samples and 38,000 testing samples. Negative polarity is class 1, and positive class 2.
Dataset Description
The files train.csv and test.csv contain all the training samples as comma-sparated values. There are 2… See the full description on the dataset page: https://huggingface.co/datasets/yassiracharki/Yelp_Reviews_for_Binary_Senti_Analysis.tourism-wikipediaemobank-single-binaryfulus-lyd-rates-2025
Fulus Libya Parallel Market Exchange Rates 2025
Historical exchange rates for the Libyan dinar (LYD) on the parallel (black) market, sourced from the fulus.ly API for the full calendar year 2025.
Dataset Description
Each row represents a single rate reading for one currency pair (LYD → foreign currency) at a specific timestamp. Multiple readings may exist per day per currency, reflecting intraday fluctuations in the parallel market.
Columns
Column
Type… See the full description on the dataset page: https://huggingface.co/datasets/Binary-ly/fulus-lyd-rates-2025.contact_prediction_binaryDataset Summary
Contact map prediction aims to determine whether two residues, $i$ and $j$, are in contact or not, based on their distance with a certain threshold ($<$8 Angstrom). This task is an important part of the early Alphafold version for structural prediction.
Data Fields
seq: a string containing the protein sequence
label: a string containing the contact label of each residue pair.
Original Dataset Name: biomap-research/contact_prediction_binary
Original Author / Organization: Biomap… See the full description on the dataset page: https://huggingface.co/datasets/AI4Protein/contact_prediction_binary.mteb_HateBR_offensive_binarygreenbeing-binary
greenbeing-binary
This is a binary classification version of the finetuning/evaluation datasets introduced in
the freeform text annotations of the greenbeing-proteins dataset.
Proteins from UniProtKB (knowledge base), from select food crops and related species.
Amino acid sequences use IUPAC-IUB codes where letters A-Z map to amino acids.
Tasks
The "query" field currently contains 18 possible labels from the annotation, and the "label" field is the binary class (0 =… See the full description on the dataset page: https://huggingface.co/datasets/monsoon-nlp/greenbeing-binary.urdu-binary-classification-dataThis Urdu sentiment dataset was formed by concatenating the following two datasets:
https://github.com/MuhammadYaseenKhan/Urdu-Sentiment-Corpus
https://www.kaggle.com/datasets/akkefa/imdb-dataset-of-50k-movie-translated-urdu-reviews
Dataset-Binary_Localization-DeepLoc
Description
Binary Localization prediction is a binary classification task where each input protein x is mapped to a label y ∈ {0, 1}, corresponding to either "membrane-bound" or "soluble" .
Protein Format: SA sequence (AF2)
Splits
The dataset is from DeepLoc: prediction of protein subcellular localization using deep learning. We employ all proteins (proteins that lack AF2 structures are removed), and split them based on 70% structure similarity (see ProteinShake), with… See the full description on the dataset page: https://huggingface.co/datasets/SaProtHub/Dataset-Binary_Localization-DeepLoc.goemotions-binaryKunkado_Maya_V2
Kunkado Maya V2 (Standardized)
Description
Ce dataset est une version transformée du dataset original RobotsMali/Kunkado (configuration human-reviewed).
L'objectif de cette version est de fournir des transcriptions dont les tags d'émotions et de bruits sont 100% compatibles avec les standards d'intégration de Maya One.
Transformations effectuées
Une pipeline de nettoyage automatique a été appliquée sur la colonne corrected-label pour générer la colonne… See the full description on the dataset page: https://huggingface.co/datasets/binaryMao/Kunkado_Maya_V2.mpi_binaryreddit-travel-qafarpo-protest-issue-binary-datasetThis is a curated binary subset of the FARPO (Far-Right Protest Observatory) dataset for automated protest event analysis across seven European countries.
Our paper describes the curation methodology in detail. Full FARPO dataset and codebooks: https://farpo.eu/data/.
Paper
Title: Automating Protest Event Analysis: A Transformer-Based Hybrid Approach
Authors: Formisano, G., Froio, C., Castelli Gattinara, P.
Status: Conditionally accepted for publication
Journal: Political… See the full description on the dataset page: https://huggingface.co/datasets/giufo/farpo-protest-issue-binary-dataset.reviews_binary_not4Shared_Task_Fake_News_binarystsb-binary-paraphrase-labelled
Paraphrase Detection Dataset (Derived from SetFit/stsb)
Description:
This dataset originates from the SetFit/stsb dataset, which was initially created for semantic textual similarity (STS) tasks with a label range of 0 to 5. It has been adapted for binary paraphrase detection by leveraging the high-accuracy paraphrase classification model viswadarshan06/pd-robert.
Each sentence pair in the original dataset has been re-labeled according to the following binary scheme:
1 →… See the full description on the dataset page: https://huggingface.co/datasets/viswadarshan06/stsb-binary-paraphrase-labelled.mpi_binary_sftfarpo-protest-identification-binary-datasetThis is a curated binary subset of the FARPO (Far-Right Protest Observatory) dataset for automated protest event analysis across seven European countries.
Our paper describes the curation methodology in detail. Full FARPO dataset and codebooks: https://farpo.eu/data/.
Paper
Title: Automating Protest Event Analysis: A Transformer-Based Hybrid Approach
Authors: Formisano, G., Froio, C., Castelli Gattinara, P.
Status: Conditionally accepted for publication
Journal: Political… See the full description on the dataset page: https://huggingface.co/datasets/giufo/farpo-protest-identification-binary-dataset.mnli_binary_testcars-binary-classificationreviews_binary_not4TuPY_dataset_binary
Portuguese Hate Speech Dataset (TuPy)
The Portuguese hate speech dataset (TuPy) is an annotated corpus designed to facilitate the development of advanced hate speech detection models using machine learning (ML) and natural language processing (NLP) techniques. TuPy is formed by 10000 thousand unpublished annotated tweets collected in 2023.
This repository is organized as follows:
root.
├── annotations : classification given by annotators
├── raw corpus : dataset before… See the full description on the dataset page: https://huggingface.co/datasets/victoriadreis/TuPY_dataset_binary.
