CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01CohereLabs /wikipedia-2023-11-embed-multilingual-v3-int8-binary Multilingual Embeddings for Wikipedia in 300+ Languages (int8 & binary embeddings) This dataset contains the wikimedia/wikipedia dataset dump from 2023-11-01 from Wikipedia in all 300+ languages. The embeddings are provided as int8 and ubinary that allow quick search and reduction of your vector index size up to 32. For more details, see Cohere int8 & binary Embeddings The individual articles have been chunked and embedded with the state-of-the-art multilingual Cohere Embed V3… See the full description on the dataset page: https://huggingface.co/datasets/CohereLabs/wikipedia-2023-11-embed-multilingual-v3-int8-binary.text100M<n<1B49 likes2.7k downloads6mo agoHugging Face02Mirali33 /mb-crater_binary_seg mb-crater_binary_seg A segmentation dataset for planetary science applications. Dataset Metadata License: CC-BY-4.0 (Creative Commons Attribution 4.0 International) Version: 1.0 Date Published: 2025-05-15 Cite As: TBD Classes This dataset contains the following classes: 0: Background 1: Crater Directory Structure The dataset follows this structure: dataset/ ├── train/ │ ├── images/ # Image files │ └── masks/ # Segmentation masks… See the full description on the dataset page: https://huggingface.co/datasets/Mirali33/mb-crater_binary_seg.imageimage-segmentation1K<n<10K0 likes887 downloads11mo agoHugging Face03bluuebunny /arxiv_abstract_embedding_mxbai_large_v1_milvus_binaryThis repo serves as the dataset backup for PaperMatch. A semantic similarity search engine. For more information, visit the blog: Behind PaperMatch textsentence-similarity1M<n<10M4 likes632 downloads2d agoHugging Face04mjbommar /binary-30k Binary-30K: Cross-Platform Binary Dataset with Stratified Splits Paper | Code 🔗 Original Dataset (no splits): mjbommar/binary-30k-tokenized This is the stratified train/validation/test split version of the Binary-30K dataset, containing 29,793 unique cross-platform binaries with pre-computed tokenization. This version provides standardized splits for reproducible machine learning research. 🎯 Key Features ✅ Stratified 70/15/15 splits maintaining class balance across… See the full description on the dataset page: https://huggingface.co/datasets/mjbommar/binary-30k.tabulartext-classification10K<n<100K5 likes505 downloads10mo agoHugging Face05mjbommar /binary-30k-tokenized Dataset Card for Binary-30K Dataset Summary Binary-30K is a comprehensive, multi-platform binary executable dataset designed for machine learning research in binary analysis, malware detection, and program understanding. The dataset contains 38,467 records representing ~30,000 unique binary executables totaling ~33.41 GB, collected from diverse sources including Linux distributions, Windows operating systems, SOREL-20M malware dataset, and Malware Bazaar collection. Note… See the full description on the dataset page: https://huggingface.co/datasets/mjbommar/binary-30k-tokenized.tabularother10K<n<100K1 likes407 downloads10mo agoHugging Face06blueskyheaven /voxceleb2-mp4-binarytext1M<n<10M0 likes345 downloads1y agoHugging Face07krasserm /wikipedia-2023-11-en-embed-mxbai-int8-binaryThis dataset is an extension of the krasserm/wikipedia-2023-11-en-text dataset, with additional columns containing ubinary and int8 embeddings of the text, created with the mixedbread-ai/mxbai-embed-large-v1 embedding model. The dataset has the following columns: _id: unique identifier of the Wikipedia text chunk title: title of the Wikipedia article url: URL of the Wikipedia article text: text chunk of the Wikipedia article emb_ubinary: binary embeddings of the Wikipedia text chunk… See the full description on the dataset page: https://huggingface.co/datasets/krasserm/wikipedia-2023-11-en-embed-mxbai-int8-binary.text10M<n<100M0 likes288 downloads2y agoHugging Face08ukr-detect /ukr-emotions-binary EmoBench-UA: Emotions Detection Dataset in Ukrainian Texts EmoBench-UA: the first of its kind emotions detection dataset in Ukrainian texts. This dataset covers the detection of basic emotions: Joy, Anger, Fear, Disgust, Surprise, Sadness, or None. Any text can contain any amount of emotion -- only one, several, or none at all. The texts with None emotions are the ones where the labels per emotions classes are 0. Binary: specifically this dataset contains binary labels… See the full description on the dataset page: https://huggingface.co/datasets/ukr-detect/ukr-emotions-binary.imagetext-classification1K<n<10K0 likes261 downloads2mo agoHugging Face09retgenai /FOTBCD-Binary FOTBCD-Binary A large-scale building change detection benchmark from French orthophotos and topographic data. Dataset Description Property Value Departments 28 (25 train / 3 eval) Image pairs ~28k Patch size 512×512 Resolution 0.2m Annotation Binary mask Splits Split Examples train ~26k val ~1k test ~1k Features Field Type Description image_id string Unique identifier image_before Image Before… See the full description on the dataset page: https://huggingface.co/datasets/retgenai/FOTBCD-Binary.imagemask-generation10K<n<100K1 likes261 downloads7mo agoHugging Face10SetFit /ethos_binaryThis is the binary split of ethos, split into train and test. It contains comments annotated for hate speech or not. textn<1K1 likes230 downloads5y agoHugging Face11EleutherAI /truthful_qa_binaryTruthfulQA-Binary is a benchmark to measure whether a language model is truthful in generating answers to questions. The benchmark comprises 817 questions that span 38 categories, including health, law, finance and politics. Questions are crafted so that some humans would answer falsely due to a false belief or misconception. To perform well, models must avoid generating false answers learned from imitating human texts.textmultiple-choicen<1K2 likes208 downloads3y agoHugging Face12Lots-of-LoRAs /task066_timetravel_binary_consistency_classification Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task066_timetravel_binary_consistency_classification Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative Instructions on… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task066_timetravel_binary_consistency_classification.texttext-generation1K<n<10K0 likes206 downloads2y agoHugging Face13HyaDoo /ko-voicephishing-binary-classificationtabular1K<n<10K0 likes153 downloads2y agoHugging Face14tuhink /cambench_binary_eval CameraBench Binary Evaluation Dataset A balanced VQA dataset for evaluating camera motion understanding in videos. 📊 Dataset Statistics Total Questions: 384 Unique Videos: 119 Unique Questions: 31 Yes Answers: 192 (50.0%) No Answers: 192 (50.0%) Balance Ratio: 1.00 Total Size: 126.16 MB (0.12 GB) Average Video Size: 1.06 MB 🎯 Task Categories This dataset covers various camera motion tasks including: Static: 42 questions Move In: 29 questions Pan Left: 24… See the full description on the dataset page: https://huggingface.co/datasets/tuhink/cambench_binary_eval.imagevisual-question-answeringn<1K0 likes151 downloads11mo agoHugging Face15Lots-of-LoRAs /task022_cosmosqa_passage_inappropriate_binary Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task022_cosmosqa_passage_inappropriate_binary Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task022_cosmosqa_passage_inappropriate_binary.texttext-generationn<1K0 likes142 downloads2y agoHugging Face16GrimSqueaker /SignalP_Binary SignalP_Binary Binary benchmark for signal peptide prediction from protein sequences, adapted from the ProteinBERT benchmark collection and SignalP data. Source This dataset is sourced from the ProteinBERT benchmark repository: https://github.com/nadavbra/protein_bert/tree/master/protein_benchmarks Curator Attribution This Hugging Face dataset packaging, curation, and publication was prepared by Dan Ofer. Splits and Schema Splits follow the… See the full description on the dataset page: https://huggingface.co/datasets/GrimSqueaker/SignalP_Binary.texttext-classification10K<n<100K2 likes138 downloads6mo agoHugging Face17yassiracharki /Amazon_Reviews_Binary_for_Sentiment_Analysis Dataset Card for Dataset Name The Amazon reviews polarity dataset is constructed by taking review score 1 and 2 as negative, and 4 and 5 as positive. Samples of score 3 is ignored. In the dataset, class 1 is the negative and class 2 is the positive. Each class has 1,800,000 training samples and 200,000 testing samples. Dataset Details Dataset Description The files train.csv and test.csv contain all the training samples as comma-sparated values. There are 3… See the full description on the dataset page: https://huggingface.co/datasets/yassiracharki/Amazon_Reviews_Binary_for_Sentiment_Analysis.texttext-classification1M<n<10M0 likes134 downloads2y agoHugging Face18Lots-of-LoRAs /task1559_blimp_binary_classification Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1559_blimp_binary_classification Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1559_blimp_binary_classification.texttext-generation1K<n<10K0 likes133 downloads2y agoHugging Face19swhitfield /biomap-research-contact_prediction_binary contact_prediction_binary Sourced from biomap-research/contact_prediction_binary and prepared for Hugging Face datasets usage. Data files Parquet files are stored under data/ using Hugging Face split naming conventions (train-*, validation-*, test-*). Preparation Preprocess mode: minimal. Seed: 1957723. No max sequence length filter was applied. Renamed source columns: label -> targets, seq -> sequence. Columns: id, sequence, targets, split.… See the full description on the dataset page: https://huggingface.co/datasets/swhitfield/biomap-research-contact_prediction_binary.text10K<n<100K0 likes116 downloads18d agoHugging Face20malexandersalazar /toxicity-multilingual-binary-classification-datasetThis dataset is a comprehensive collection designed to aid in the development of robust and nuanced models for identifying toxic language across multiple languages, while critically distinguishing it from expressions related to mental health, specifically depression. It synthesizes content from three existing public datasets (ToxiGen, TextDetox, and Mental Health - Depression) with a newly generated synthetic dataset (ToxiLLaMA). The creation process involved careful collection, extensive… See the full description on the dataset page: https://huggingface.co/datasets/malexandersalazar/toxicity-multilingual-binary-classification-dataset.text100K<n<1M1 likes109 downloads1y agoHugging Face21Lots-of-LoRAs /task609_sbic_potentially_offense_binary_classification Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task609_sbic_potentially_offense_binary_classification Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative Instructions on… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task609_sbic_potentially_offense_binary_classification.texttext-generation1K<n<10K0 likes107 downloads2y agoHugging Face22HandsomeWin /binary-30k Binary-30K: Cross-Platform Binary Dataset with Stratified Splits Paper | Code 🔗 Original Dataset (no splits): mjbommar/binary-30k-tokenized This is the stratified train/validation/test split version of the Binary-30K dataset, containing 29,793 unique cross-platform binaries with pre-computed tokenization. This version provides standardized splits for reproducible machine learning research. 🎯 Key Features ✅ Stratified 70/15/15 splits maintaining class balance across… See the full description on the dataset page: https://huggingface.co/datasets/HandsomeWin/binary-30k.tabulartext-classification10K<n<100K0 likes107 downloads6mo agoHugging Face23fwufbhiwuhf /binary-30k-tokenized Dataset Card for Binary-30K Dataset Summary Binary-30K is a comprehensive, multi-platform binary executable dataset designed for machine learning research in binary analysis, malware detection, and program understanding. The dataset contains 38,467 records representing ~30,000 unique binary executables totaling ~33.41 GB, collected from diverse sources including Linux distributions, Windows operating systems, SOREL-20M malware dataset, and Malware Bazaar collection. Note… See the full description on the dataset page: https://huggingface.co/datasets/fwufbhiwuhf/binary-30k-tokenized.tabularother10K<n<100K0 likes103 downloads5mo agoHugging Face24DipeshChaudhary /binary-nepali-ged-datasettext10M<n<100M0 likes101 downloads11mo agoHugging Face25Lots-of-LoRAs /task607_sbic_intentional_offense_binary_classification Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task607_sbic_intentional_offense_binary_classification Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative Instructions on… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task607_sbic_intentional_offense_binary_classification.texttext-generation1K<n<10K0 likes98 downloads2y agoHugging Face26capa2000 /binary-classifier-birdnet Binary BirdNet Classifier Contiene anotaciones y audios de 3s y 5s para clasificación binaria con rutas relativas. audion<1K0 likes95 downloads1y agoHugging Face27karths /binary-10IQR-secu Dataset Card for "binary-10IQR-secu" More Information needed tabular100K<n<1M0 likes94 downloads2y agoHugging Face28abullard1 /steam-reviews-constructiveness-binary-label-annotations-1.5k 1.5K Steam Reviews Binary Labeled for Constructiveness Dataset Summary This dataset contains 1,461 Steam reviews from 10 of the most reviewed games. Each game has about the same amount of reviews. Each review is annotated with a binary label indicating whether the review is constructive or not. The dataset is designed to support tasks related to text classification, particularly constructiveness detection tasks in the gaming domain. Also available as… See the full description on the dataset page: https://huggingface.co/datasets/abullard1/steam-reviews-constructiveness-binary-label-annotations-1.5k.tabulartext-classification1K<n<10K2 likes91 downloads2y agoHugging Face29laion /terminal_bench_2_a3_rl_DCAgent_mix_h4_binary_easy_50_8B_20260828_180443text1K<n<10K0 likes91 downloads23d agoHugging Face30laion /swebench_verified_random_100_folders_a3_rl_DCAgent_mix_h4_binary_easy_50_8B_20260897d58773text10K<n<100K0 likes91 downloads21d agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.