CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01sentence-transformers /msmarco-distilbert-margin-mse-mean-dot-v1 MS MARCO with hard negatives from distilbert-margin-mse-mean-dot-v1 MS MARCO is a large scale information retrieval corpus that was created based on real user search queries using the Bing search engine. For each query and gold positive passage, the 50 most similar paragraphs were mined using 13 different models. The resulting data can be used to train Sentence Transformer models. Related Datasets These are the datasets generated using the 13 different models:… See the full description on the dataset page: https://huggingface.co/datasets/sentence-transformers/msmarco-distilbert-margin-mse-mean-dot-v1.tabularfeature-extraction10M<n<100M2 likes3.2k downloads2y agoHugging Face02sentence-transformers /msmarco-msmarco-distilbert-base-v3 MS MARCO with hard negatives from msmarco-distilbert-base-v3 MS MARCO is a large scale information retrieval corpus that was created based on real user search queries using the Bing search engine. For each query and gold positive passage, the 50 most similar paragraphs were mined using 13 different models. The resulting data can be used to train Sentence Transformer models. Related Datasets These are the datasets generated using the 13 different models: msmarco-bm25… See the full description on the dataset page: https://huggingface.co/datasets/sentence-transformers/msmarco-msmarco-distilbert-base-v3.tabularfeature-extraction10M<n<100M5 likes2.3k downloads2y agoHugging Face03sentence-transformers /msmarco-distilbert-margin-mse-sym-mnrl-mean-v2 MS MARCO with hard negatives from distilbert-margin-mse-sym-mnrl-mean-v2 MS MARCO is a large scale information retrieval corpus that was created based on real user search queries using the Bing search engine. For each query and gold positive passage, the 50 most similar paragraphs were mined using 13 different models. The resulting data can be used to train Sentence Transformer models. Related Datasets These are the datasets generated using the 13 different models:… See the full description on the dataset page: https://huggingface.co/datasets/sentence-transformers/msmarco-distilbert-margin-mse-sym-mnrl-mean-v2.tabularfeature-extraction10M<n<100M1 likes2.2k downloads2y agoHugging Face04sentence-transformers /msmarco-distilbert-margin-mse-cls-dot-v1 MS MARCO with hard negatives from distilbert-margin-mse-cls-dot-v1 MS MARCO is a large scale information retrieval corpus that was created based on real user search queries using the Bing search engine. For each query and gold positive passage, the 50 most similar paragraphs were mined using 13 different models. The resulting data can be used to train Sentence Transformer models. Related Datasets These are the datasets generated using the 13 different models:… See the full description on the dataset page: https://huggingface.co/datasets/sentence-transformers/msmarco-distilbert-margin-mse-cls-dot-v1.tabularfeature-extraction10M<n<100M0 likes1.7k downloads2y agoHugging Face05sentence-transformers /msmarco-msmarco-distilbert-base-tas-b MS MARCO with hard negatives from msmarco-distilbert-base-tas-b MS MARCO is a large scale information retrieval corpus that was created based on real user search queries using the Bing search engine. For each query and gold positive passage, the 50 most similar paragraphs were mined using 13 different models. The resulting data can be used to train Sentence Transformer models. Related Datasets These are the datasets generated using the 13 different models:… See the full description on the dataset page: https://huggingface.co/datasets/sentence-transformers/msmarco-msmarco-distilbert-base-tas-b.tabularfeature-extraction10M<n<100M5 likes1.7k downloads2y agoHugging Face06sentence-transformers /msmarco-distilbert-margin-mse-sym-mnrl-mean-v1 MS MARCO with hard negatives from distilbert-margin-mse-sym-mnrl-mean-v1 MS MARCO is a large scale information retrieval corpus that was created based on real user search queries using the Bing search engine. For each query and gold positive passage, the 50 most similar paragraphs were mined using 13 different models. The resulting data can be used to train Sentence Transformer models. Related Datasets These are the datasets generated using the 13 different models:… See the full description on the dataset page: https://huggingface.co/datasets/sentence-transformers/msmarco-distilbert-margin-mse-sym-mnrl-mean-v1.tabularfeature-extraction10M<n<100M1 likes1.4k downloads2y agoHugging Face07sentence-transformers /msmarco-distilbert-margin-mse-mnrl-mean-v1 MS MARCO with hard negatives from distilbert-margin-mse-mnrl-mean-v1 MS MARCO is a large scale information retrieval corpus that was created based on real user search queries using the Bing search engine. For each query and gold positive passage, the 50 most similar paragraphs were mined using 13 different models. The resulting data can be used to train Sentence Transformer models. Related Datasets These are the datasets generated using the 13 different models:… See the full description on the dataset page: https://huggingface.co/datasets/sentence-transformers/msmarco-distilbert-margin-mse-mnrl-mean-v1.tabularfeature-extraction10M<n<100M0 likes833 downloads2y agoHugging Face08lsr42 /msmarco-psgs-distilbert-dot-v5text1M<n<10M0 likes641 downloads2y agoHugging Face09WendyHoang /news-ka-shuffled-DISTILBERTtext1M<n<10M0 likes264 downloads2y agoHugging Face10sentence-transformers /msmarco-distilbert-margin-mse-cls-dot-v2 MS MARCO with hard negatives from distilbert-margin-mse-cls-dot-v2 MS MARCO is a large scale information retrieval corpus that was created based on real user search queries using the Bing search engine. For each query and gold positive passage, the 50 most similar paragraphs were mined using 13 different models. The resulting data can be used to train Sentence Transformer models. Related Datasets These are the datasets generated using the 13 different models:… See the full description on the dataset page: https://huggingface.co/datasets/sentence-transformers/msmarco-distilbert-margin-mse-cls-dot-v2.tabularfeature-extraction10M<n<100M2 likes223 downloads2y agoHugging Face11WendyHoang /news-ka-small-DISTILBERTtext100K<n<1M0 likes80 downloads2y agoHugging Face12feyninc /chonkiepedia-distilbert-tokenized1M<n<10M0 likes54 downloads1y agoHugging Face13baharehansari1 /distilbert_spelling_dataset-v2text10K<n<100K0 likes46 downloads25d agoHugging Face14baharehansari1 /distilbert_spelling_dataset-v3text10K<n<100K0 likes44 downloads25d agoHugging Face15baharehansari1 /distilbert_spelling_datasettext10K<n<100K0 likes43 downloads28d agoHugging Face16Khalyie /sst2-distilbert-data SST-2 (GLUE) — Raw Splits Used for Khalyie/sst2-distilbert Same train/validation/test splits used to fine-tune Khalyie/sst2-distilbert. Note: the test split's label column is -1 for every row — GLUE withholds official test labels. Use validation for labeled evaluation. Reload with: from datasets import load_dataset ds = load_dataset("Khalyie/sst2-distilbert-data") text10K<n<100K0 likes38 downloads12d agoHugging Face17bzhao18 /hyperpartisan-news-distilbert-tokenstext100K<n<1M0 likes29 downloads2y agoHugging Face18AmjaadXX /sentiment-analysis-distilbert-rotten-tomatoes Sentiment Analysis NLP Pipeline (Rotten Tomatoes) This project builds a complete NLP data processing pipeline for sentiment analysis using the Rotten Tomatoes dataset. The workflow focuses on: Data exploration Data cleaning Feature engineering Tokenization and model preparation No model training has been performed yet. This project prepares the dataset for training transformer-based models such as DistilBERT. Dataset We use the Rotten Tomatoes dataset from… See the full description on the dataset page: https://huggingface.co/datasets/AmjaadXX/sentiment-analysis-distilbert-rotten-tomatoes.text10K<n<100K0 likes28 downloads5mo agoHugging Face19snyamson /covid-tweet-sentiment-analyzer-distilbert-data Dataset Card for "covid-tweet-sentiment-analyzer-distilbert-data" More Information needed 1K<n<10K1 likes26 downloads3y agoHugging Face20nikchar /retrieval_verification_bm25_distilbert Dataset Card for "retrieval_verification_bm25_distilbert" More Information needed tabular10K<n<100K0 likes25 downloads3y agoHugging Face21nikchar /retrieval_verification_distilbert Dataset Card for "retrieval_verification_distilbert" More Information needed tabular10K<n<100K0 likes23 downloads3y agoHugging Face22bambadij /Tweet_sentiment_analysis_Distilbert Dataset Card for "Tweet_sentiment_analysis_Distilbert" More Information needed 1K<n<10K0 likes19 downloads3y agoHugging Face23fantasticrambo /covid-tweet-sentiment-analyzer-distilbert-data1K<n<10K1 likes18 downloads3y agoHugging Face24sanjin7 /embedding_dataset_distilbert_base_uncased_ad_subwords Dataset Card for "embedding_dataset_distilbert_base_uncased_ad_subwords" More Information needed tabular1K<n<10K0 likes15 downloads4y agoHugging Face25sukantan /nyaya-ae-msmarco-distilbert-base-tas-b Dataset Card for "nyaya-ae-msmarco-distilbert-base-tas-b" More Information needed tabular10K<n<100K0 likes15 downloads3y agoHugging Face26monostate /fintech-sentiment-distilbert-balanced-v2 Dataset Card for monostate/fintech-sentiment-distilbert-balanced-v2 Dataset Description This dataset was generated using Vibe Data Director, a tool for creating and curating text classification datasets. Dataset Summary Session ID: session_f15db25a Generated: 2026-03-16T13:09:50.740014 Total Samples: 481 Classes: negative, positive, neutral Styles: none Dataset Structure Data Fields text (string): The text content of the sample class… See the full description on the dataset page: https://huggingface.co/datasets/monostate/fintech-sentiment-distilbert-balanced-v2.texttext-classificationn<1K0 likes15 downloads6mo agoHugging Face27johannes-garstenauer /embeddings_from_distilbert_class_heaps_and_eval_part0 Dataset Card for "embeddings_from_distilbert_class_heaps_and_eval_part0" More Information needed tabular100K<n<1M0 likes14 downloads3y agoHugging Face28KongYang /distilbert-base-uncased-finetuned-imdb-accelerator10K<n<100K0 likes14 downloads2y agoHugging Face29scademy /autotrain-data-DistilBert-500-500 Dataset Card for "autotrain-data-DistilBert-500-500" More Information needed text1K<n<10K0 likes13 downloads3y agoHugging Face30johannes-garstenauer /embeddings_from_distilbert_class_heaps Dataset Card for "embeddings_from_distilbert_class_heaps" Dataset created for thesis: "Generating Robust Representations of Structures in OpenSSH Heap Dumps" by Johannes Garstenauer. This dataset contains representations of heap data structures along with their labels and the predicted label. The representations are the [CLS] token embeddings of the last 3 layers of the DistilBERT model. The representation-generating model is:… See the full description on the dataset page: https://huggingface.co/datasets/johannes-garstenauer/embeddings_from_distilbert_class_heaps.tabular100K<n<1M1 likes12 downloads3y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.