CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01sentence-transformers /msmarco-distilbert-margin-mse-mean-dot-v1 MS MARCO with hard negatives from distilbert-margin-mse-mean-dot-v1 MS MARCO is a large scale information retrieval corpus that was created based on real user search queries using the Bing search engine. For each query and gold positive passage, the 50 most similar paragraphs were mined using 13 different models. The resulting data can be used to train Sentence Transformer models. Related Datasets These are the datasets generated using the 13 different models:… See the full description on the dataset page: https://huggingface.co/datasets/sentence-transformers/msmarco-distilbert-margin-mse-mean-dot-v1.tabularfeature-extraction10M<n<100M2 likes3.4k downloads2y agoHugging Face02sentence-transformers /msmarco-msmarco-distilbert-base-v3 MS MARCO with hard negatives from msmarco-distilbert-base-v3 MS MARCO is a large scale information retrieval corpus that was created based on real user search queries using the Bing search engine. For each query and gold positive passage, the 50 most similar paragraphs were mined using 13 different models. The resulting data can be used to train Sentence Transformer models. Related Datasets These are the datasets generated using the 13 different models: msmarco-bm25… See the full description on the dataset page: https://huggingface.co/datasets/sentence-transformers/msmarco-msmarco-distilbert-base-v3.tabularfeature-extraction10M<n<100M5 likes2.4k downloads2y agoHugging Face03sentence-transformers /msmarco-distilbert-margin-mse-sym-mnrl-mean-v2 MS MARCO with hard negatives from distilbert-margin-mse-sym-mnrl-mean-v2 MS MARCO is a large scale information retrieval corpus that was created based on real user search queries using the Bing search engine. For each query and gold positive passage, the 50 most similar paragraphs were mined using 13 different models. The resulting data can be used to train Sentence Transformer models. Related Datasets These are the datasets generated using the 13 different models:… See the full description on the dataset page: https://huggingface.co/datasets/sentence-transformers/msmarco-distilbert-margin-mse-sym-mnrl-mean-v2.tabularfeature-extraction10M<n<100M1 likes2.3k downloads2y agoHugging Face04sentence-transformers /msmarco-msmarco-distilbert-base-tas-b MS MARCO with hard negatives from msmarco-distilbert-base-tas-b MS MARCO is a large scale information retrieval corpus that was created based on real user search queries using the Bing search engine. For each query and gold positive passage, the 50 most similar paragraphs were mined using 13 different models. The resulting data can be used to train Sentence Transformer models. Related Datasets These are the datasets generated using the 13 different models:… See the full description on the dataset page: https://huggingface.co/datasets/sentence-transformers/msmarco-msmarco-distilbert-base-tas-b.tabularfeature-extraction10M<n<100M5 likes1.8k downloads2y agoHugging Face05sentence-transformers /msmarco-distilbert-margin-mse-cls-dot-v1 MS MARCO with hard negatives from distilbert-margin-mse-cls-dot-v1 MS MARCO is a large scale information retrieval corpus that was created based on real user search queries using the Bing search engine. For each query and gold positive passage, the 50 most similar paragraphs were mined using 13 different models. The resulting data can be used to train Sentence Transformer models. Related Datasets These are the datasets generated using the 13 different models:… See the full description on the dataset page: https://huggingface.co/datasets/sentence-transformers/msmarco-distilbert-margin-mse-cls-dot-v1.tabularfeature-extraction10M<n<100M0 likes1.5k downloads2y agoHugging Face06sentence-transformers /msmarco-distilbert-margin-mse-sym-mnrl-mean-v1 MS MARCO with hard negatives from distilbert-margin-mse-sym-mnrl-mean-v1 MS MARCO is a large scale information retrieval corpus that was created based on real user search queries using the Bing search engine. For each query and gold positive passage, the 50 most similar paragraphs were mined using 13 different models. The resulting data can be used to train Sentence Transformer models. Related Datasets These are the datasets generated using the 13 different models:… See the full description on the dataset page: https://huggingface.co/datasets/sentence-transformers/msmarco-distilbert-margin-mse-sym-mnrl-mean-v1.tabularfeature-extraction10M<n<100M1 likes1.5k downloads2y agoHugging Face07sentence-transformers /msmarco-distilbert-margin-mse-mnrl-mean-v1 MS MARCO with hard negatives from distilbert-margin-mse-mnrl-mean-v1 MS MARCO is a large scale information retrieval corpus that was created based on real user search queries using the Bing search engine. For each query and gold positive passage, the 50 most similar paragraphs were mined using 13 different models. The resulting data can be used to train Sentence Transformer models. Related Datasets These are the datasets generated using the 13 different models:… See the full description on the dataset page: https://huggingface.co/datasets/sentence-transformers/msmarco-distilbert-margin-mse-mnrl-mean-v1.tabularfeature-extraction10M<n<100M0 likes857 downloads2y agoHugging Face08lsr42 /msmarco-psgs-distilbert-dot-v5text1M<n<10M0 likes451 downloads2y agoHugging Face09WendyHoang /news-ka-shuffled-DISTILBERTtext1M<n<10M0 likes420 downloads2y agoHugging Face10sentence-transformers /msmarco-distilbert-margin-mse-cls-dot-v2 MS MARCO with hard negatives from distilbert-margin-mse-cls-dot-v2 MS MARCO is a large scale information retrieval corpus that was created based on real user search queries using the Bing search engine. For each query and gold positive passage, the 50 most similar paragraphs were mined using 13 different models. The resulting data can be used to train Sentence Transformer models. Related Datasets These are the datasets generated using the 13 different models:… See the full description on the dataset page: https://huggingface.co/datasets/sentence-transformers/msmarco-distilbert-margin-mse-cls-dot-v2.tabularfeature-extraction10M<n<100M2 likes237 downloads2y agoHugging Face11WendyHoang /news-ka-small-DISTILBERTtext100K<n<1M0 likes104 downloads2y agoHugging Face12feyninc /chonkiepedia-distilbert-tokenized1M<n<10M0 likes49 downloads1y agoHugging Face13baharehansari1 /distilbert_spelling_dataset-v2text10K<n<100K0 likes47 downloads27d agoHugging Face14baharehansari1 /distilbert_spelling_dataset-v3text10K<n<100K0 likes45 downloads27d agoHugging Face15Khalyie /sst2-distilbert-data SST-2 (GLUE) — Raw Splits Used for Khalyie/sst2-distilbert Same train/validation/test splits used to fine-tune Khalyie/sst2-distilbert. Note: the test split's label column is -1 for every row — GLUE withholds official test labels. Use validation for labeled evaluation. Reload with: from datasets import load_dataset ds = load_dataset("Khalyie/sst2-distilbert-data") text10K<n<100K0 likes39 downloads14d agoHugging Face16bzhao18 /hyperpartisan-news-distilbert-tokenstext100K<n<1M0 likes30 downloads2y agoHugging Face17snyamson /covid-tweet-sentiment-analyzer-distilbert-data Dataset Card for "covid-tweet-sentiment-analyzer-distilbert-data" More Information needed 1K<n<10K1 likes25 downloads3y agoHugging Face18baharehansari1 /distilbert_spelling_datasettext10K<n<100K0 likes24 downloads1mo agoHugging Face19nikchar /retrieval_verification_distilbert Dataset Card for "retrieval_verification_distilbert" More Information needed tabular10K<n<100K0 likes22 downloads3y agoHugging Face20AmjaadXX /sentiment-analysis-distilbert-rotten-tomatoes Sentiment Analysis NLP Pipeline (Rotten Tomatoes) This project builds a complete NLP data processing pipeline for sentiment analysis using the Rotten Tomatoes dataset. The workflow focuses on: Data exploration Data cleaning Feature engineering Tokenization and model preparation No model training has been performed yet. This project prepares the dataset for training transformer-based models such as DistilBERT. Dataset We use the Rotten Tomatoes dataset from… See the full description on the dataset page: https://huggingface.co/datasets/AmjaadXX/sentiment-analysis-distilbert-rotten-tomatoes.text10K<n<100K0 likes21 downloads5mo agoHugging Face21bambadij /Tweet_sentiment_analysis_Distilbert Dataset Card for "Tweet_sentiment_analysis_Distilbert" More Information needed 1K<n<10K0 likes20 downloads3y agoHugging Face22nikchar /retrieval_verification_bm25_distilbert Dataset Card for "retrieval_verification_bm25_distilbert" More Information needed tabular10K<n<100K0 likes19 downloads3y agoHugging Face23fantasticrambo /covid-tweet-sentiment-analyzer-distilbert-data1K<n<10K1 likes17 downloads3y agoHugging Face24sanjin7 /embedding_dataset_distilbert_base_uncased_ad_subwords Dataset Card for "embedding_dataset_distilbert_base_uncased_ad_subwords" More Information needed tabular1K<n<10K0 likes15 downloads4y agoHugging Face25sukantan /nyaya-ae-msmarco-distilbert-base-tas-b Dataset Card for "nyaya-ae-msmarco-distilbert-base-tas-b" More Information needed tabular10K<n<100K0 likes15 downloads3y agoHugging Face26monostate /fintech-sentiment-distilbert-balanced-v2 Dataset Card for monostate/fintech-sentiment-distilbert-balanced-v2 Dataset Description This dataset was generated using Vibe Data Director, a tool for creating and curating text classification datasets. Dataset Summary Session ID: session_f15db25a Generated: 2026-03-16T13:09:50.740014 Total Samples: 481 Classes: negative, positive, neutral Styles: none Dataset Structure Data Fields text (string): The text content of the sample class… See the full description on the dataset page: https://huggingface.co/datasets/monostate/fintech-sentiment-distilbert-balanced-v2.texttext-classificationn<1K0 likes15 downloads6mo agoHugging Face27scademy /autotrain-data-DistilBert-500-500 Dataset Card for "autotrain-data-DistilBert-500-500" More Information needed text1K<n<10K0 likes14 downloads3y agoHugging Face28monostate /fintech-sentiment-distilbert-ready Dataset Card for monostate/fintech-sentiment-distilbert-ready Dataset Description This dataset was generated using Vibe Data Director, a tool for creating and curating text classification datasets. Dataset Summary Session ID: session_7987dfd2 Generated: 2026-03-16T12:59:16.439137 Total Samples: 406 Classes: positive, negative, neutral Styles: none Dataset Structure Data Fields text (string): The text content of the sample class… See the full description on the dataset page: https://huggingface.co/datasets/monostate/fintech-sentiment-distilbert-ready.texttext-classificationn<1K0 likes14 downloads6mo agoHugging Face29jhan21 /amazon-reviews-tokenized-distilbert-balanced-3labels100K<n<1M0 likes13 downloads1y agoHugging Face30burkelive /distilbert-base-uncased-pii-200_datasetn<1K0 likes12 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.