datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
wikipedia-2023-11-embed-multilingual-v3-int8-binary
Multilingual Embeddings for Wikipedia in 300+ Languages (int8 & binary embeddings)
This dataset contains the wikimedia/wikipedia dataset dump from 2023-11-01 from Wikipedia in all 300+ languages. The embeddings are provided as int8 and ubinary that allow quick search and reduction of your vector index size up to 32. For more details, see Cohere int8 & binary Embeddings
The individual articles have been chunked and embedded with the state-of-the-art multilingual Cohere Embed V3… See the full description on the dataset page: https://huggingface.co/datasets/CohereLabs/wikipedia-2023-11-embed-multilingual-v3-int8-binary.mb-crater_binary_seg
mb-crater_binary_seg
A segmentation dataset for planetary science applications.
Dataset Metadata
License: CC-BY-4.0 (Creative Commons Attribution 4.0 International)
Version: 1.0
Date Published: 2025-05-15
Cite As: TBD
Classes
This dataset contains the following classes:
0: Background
1: Crater
Directory Structure
The dataset follows this structure:
dataset/
├── train/
│ ├── images/ # Image files
│ └── masks/ # Segmentation masks… See the full description on the dataset page: https://huggingface.co/datasets/Mirali33/mb-crater_binary_seg.arxiv_abstract_embedding_mxbai_large_v1_milvus_binaryThis repo serves as the dataset backup for PaperMatch. A semantic similarity search engine.
For more information, visit the blog: Behind PaperMatch
binary-30k
Binary-30K: Cross-Platform Binary Dataset with Stratified Splits
Paper | Code
🔗 Original Dataset (no splits): mjbommar/binary-30k-tokenized
This is the stratified train/validation/test split version of the Binary-30K dataset, containing 29,793 unique cross-platform binaries with pre-computed tokenization. This version provides standardized splits for reproducible machine learning research.
🎯 Key Features
✅ Stratified 70/15/15 splits maintaining class balance across… See the full description on the dataset page: https://huggingface.co/datasets/mjbommar/binary-30k.binary-30k-tokenized
Dataset Card for Binary-30K
Dataset Summary
Binary-30K is a comprehensive, multi-platform binary executable dataset designed for machine learning research in binary analysis, malware detection, and program understanding. The dataset contains 38,467 records representing ~30,000 unique binary executables totaling ~33.41 GB, collected from diverse sources including Linux distributions, Windows operating systems, SOREL-20M malware dataset, and Malware Bazaar collection.
Note… See the full description on the dataset page: https://huggingface.co/datasets/mjbommar/binary-30k-tokenized.voxceleb2-mp4-binarywikipedia-2023-11-en-embed-mxbai-int8-binaryThis dataset is an extension of the krasserm/wikipedia-2023-11-en-text
dataset, with additional columns containing ubinary and int8 embeddings of the text, created with the mixedbread-ai/mxbai-embed-large-v1
embedding model. The dataset has the following columns:
_id: unique identifier of the Wikipedia text chunk
title: title of the Wikipedia article
url: URL of the Wikipedia article
text: text chunk of the Wikipedia article
emb_ubinary: binary embeddings of the Wikipedia text chunk… See the full description on the dataset page: https://huggingface.co/datasets/krasserm/wikipedia-2023-11-en-embed-mxbai-int8-binary.ukr-emotions-binary
EmoBench-UA: Emotions Detection Dataset in Ukrainian Texts
EmoBench-UA: the first of its kind emotions detection dataset in Ukrainian texts. This dataset covers the detection of basic emotions: Joy, Anger, Fear, Disgust, Surprise, Sadness, or None.
Any text can contain any amount of emotion -- only one, several, or none at all. The texts with None emotions are the ones where the labels per emotions classes are 0.
Binary: specifically this dataset contains binary labels… See the full description on the dataset page: https://huggingface.co/datasets/ukr-detect/ukr-emotions-binary.FOTBCD-Binary
FOTBCD-Binary
A large-scale building change detection benchmark from French orthophotos and topographic data.
Dataset Description
Property
Value
Departments
28 (25 train / 3 eval)
Image pairs
~28k
Patch size
512×512
Resolution
0.2m
Annotation
Binary mask
Splits
Split
Examples
train
~26k
val
~1k
test
~1k
Features
Field
Type
Description
image_id
string
Unique identifier
image_before
Image
Before… See the full description on the dataset page: https://huggingface.co/datasets/retgenai/FOTBCD-Binary.ethos_binaryThis is the binary split of ethos, split into train and test.
It contains comments annotated for hate speech or not.
truthful_qa_binaryTruthfulQA-Binary is a benchmark to measure whether a language model is truthful in
generating answers to questions. The benchmark comprises 817 questions that
span 38 categories, including health, law, finance and politics. Questions are
crafted so that some humans would answer falsely due to a false belief or
misconception. To perform well, models must avoid generating false answers
learned from imitating human texts.task066_timetravel_binary_consistency_classification
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task066_timetravel_binary_consistency_classification
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task066_timetravel_binary_consistency_classification.ko-voicephishing-binary-classificationcambench_binary_eval
CameraBench Binary Evaluation Dataset
A balanced VQA dataset for evaluating camera motion understanding in videos.
📊 Dataset Statistics
Total Questions: 384
Unique Videos: 119
Unique Questions: 31
Yes Answers: 192 (50.0%)
No Answers: 192 (50.0%)
Balance Ratio: 1.00
Total Size: 126.16 MB (0.12 GB)
Average Video Size: 1.06 MB
🎯 Task Categories
This dataset covers various camera motion tasks including:
Static: 42 questions
Move In: 29 questions
Pan Left: 24… See the full description on the dataset page: https://huggingface.co/datasets/tuhink/cambench_binary_eval.task022_cosmosqa_passage_inappropriate_binary
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task022_cosmosqa_passage_inappropriate_binary
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task022_cosmosqa_passage_inappropriate_binary.SignalP_Binary
SignalP_Binary
Binary benchmark for signal peptide prediction from protein sequences, adapted from the ProteinBERT benchmark collection and SignalP data.
Source
This dataset is sourced from the ProteinBERT benchmark repository:
https://github.com/nadavbra/protein_bert/tree/master/protein_benchmarks
Curator Attribution
This Hugging Face dataset packaging, curation, and publication was prepared by Dan Ofer.
Splits and Schema
Splits follow the… See the full description on the dataset page: https://huggingface.co/datasets/GrimSqueaker/SignalP_Binary.Amazon_Reviews_Binary_for_Sentiment_Analysis
Dataset Card for Dataset Name
The Amazon reviews polarity dataset is constructed by taking review score 1 and 2 as negative, and 4 and 5 as positive. Samples of score 3 is ignored. In the dataset, class 1 is the negative and class 2 is the positive. Each class has 1,800,000 training samples and 200,000 testing samples.
Dataset Details
Dataset Description
The files train.csv and test.csv contain all the training samples as comma-sparated values. There are 3… See the full description on the dataset page: https://huggingface.co/datasets/yassiracharki/Amazon_Reviews_Binary_for_Sentiment_Analysis.task1559_blimp_binary_classification
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1559_blimp_binary_classification
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1559_blimp_binary_classification.biomap-research-contact_prediction_binary
contact_prediction_binary
Sourced from biomap-research/contact_prediction_binary and prepared for Hugging Face datasets usage.
Data files
Parquet files are stored under data/ using Hugging Face split naming conventions
(train-*, validation-*, test-*).
Preparation
Preprocess mode: minimal.
Seed: 1957723.
No max sequence length filter was applied.
Renamed source columns: label -> targets, seq -> sequence.
Columns: id, sequence, targets, split.… See the full description on the dataset page: https://huggingface.co/datasets/swhitfield/biomap-research-contact_prediction_binary.toxicity-multilingual-binary-classification-datasetThis dataset is a comprehensive collection designed to aid in the development of robust and nuanced models for identifying toxic language across multiple languages, while critically distinguishing it from expressions related to mental health, specifically depression. It synthesizes content from three existing public datasets (ToxiGen, TextDetox, and Mental Health - Depression) with a newly generated synthetic dataset (ToxiLLaMA). The creation process involved careful collection, extensive… See the full description on the dataset page: https://huggingface.co/datasets/malexandersalazar/toxicity-multilingual-binary-classification-dataset.task609_sbic_potentially_offense_binary_classification
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task609_sbic_potentially_offense_binary_classification
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task609_sbic_potentially_offense_binary_classification.binary-30k
Binary-30K: Cross-Platform Binary Dataset with Stratified Splits
Paper | Code
🔗 Original Dataset (no splits): mjbommar/binary-30k-tokenized
This is the stratified train/validation/test split version of the Binary-30K dataset, containing 29,793 unique cross-platform binaries with pre-computed tokenization. This version provides standardized splits for reproducible machine learning research.
🎯 Key Features
✅ Stratified 70/15/15 splits maintaining class balance across… See the full description on the dataset page: https://huggingface.co/datasets/HandsomeWin/binary-30k.binary-30k-tokenized
Dataset Card for Binary-30K
Dataset Summary
Binary-30K is a comprehensive, multi-platform binary executable dataset designed for machine learning research in binary analysis, malware detection, and program understanding. The dataset contains 38,467 records representing ~30,000 unique binary executables totaling ~33.41 GB, collected from diverse sources including Linux distributions, Windows operating systems, SOREL-20M malware dataset, and Malware Bazaar collection.
Note… See the full description on the dataset page: https://huggingface.co/datasets/fwufbhiwuhf/binary-30k-tokenized.binary-nepali-ged-datasettask607_sbic_intentional_offense_binary_classification
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task607_sbic_intentional_offense_binary_classification
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task607_sbic_intentional_offense_binary_classification.binary-classifier-birdnet
Binary BirdNet Classifier
Contiene anotaciones y audios de 3s y 5s para clasificación binaria con rutas relativas.
binary-10IQR-secu
Dataset Card for "binary-10IQR-secu"
More Information needed
steam-reviews-constructiveness-binary-label-annotations-1.5k
1.5K Steam Reviews Binary Labeled for Constructiveness
Dataset Summary
This dataset contains 1,461 Steam reviews from 10 of the most reviewed games. Each game has about the same amount of reviews. Each review is annotated with a binary label indicating whether the review is constructive or not. The dataset is designed to support tasks related to text classification, particularly constructiveness detection tasks in the gaming domain.
Also available as… See the full description on the dataset page: https://huggingface.co/datasets/abullard1/steam-reviews-constructiveness-binary-label-annotations-1.5k.terminal_bench_2_a3_rl_DCAgent_mix_h4_binary_easy_50_8B_20260828_180443swebench_verified_random_100_folders_a3_rl_DCAgent_mix_h4_binary_easy_50_8B_20260897d58773
