CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01sentence-transformers /msmarco-distilbert-margin-mse-mean-dot-v1 MS MARCO with hard negatives from distilbert-margin-mse-mean-dot-v1 MS MARCO is a large scale information retrieval corpus that was created based on real user search queries using the Bing search engine. For each query and gold positive passage, the 50 most similar paragraphs were mined using 13 different models. The resulting data can be used to train Sentence Transformer models. Related Datasets These are the datasets generated using the 13 different models:… See the full description on the dataset page: https://huggingface.co/datasets/sentence-transformers/msmarco-distilbert-margin-mse-mean-dot-v1.tabularfeature-extraction10M<n<100M2 likes3.4k downloads2y agoHugging Face02sentence-transformers /msmarco-msmarco-distilbert-base-v3 MS MARCO with hard negatives from msmarco-distilbert-base-v3 MS MARCO is a large scale information retrieval corpus that was created based on real user search queries using the Bing search engine. For each query and gold positive passage, the 50 most similar paragraphs were mined using 13 different models. The resulting data can be used to train Sentence Transformer models. Related Datasets These are the datasets generated using the 13 different models: msmarco-bm25… See the full description on the dataset page: https://huggingface.co/datasets/sentence-transformers/msmarco-msmarco-distilbert-base-v3.tabularfeature-extraction10M<n<100M5 likes2.4k downloads2y agoHugging Face03sentence-transformers /msmarco-distilbert-margin-mse-sym-mnrl-mean-v2 MS MARCO with hard negatives from distilbert-margin-mse-sym-mnrl-mean-v2 MS MARCO is a large scale information retrieval corpus that was created based on real user search queries using the Bing search engine. For each query and gold positive passage, the 50 most similar paragraphs were mined using 13 different models. The resulting data can be used to train Sentence Transformer models. Related Datasets These are the datasets generated using the 13 different models:… See the full description on the dataset page: https://huggingface.co/datasets/sentence-transformers/msmarco-distilbert-margin-mse-sym-mnrl-mean-v2.tabularfeature-extraction10M<n<100M1 likes2.2k downloads2y agoHugging Face04sentence-transformers /msmarco-distilbert-margin-mse-cls-dot-v1 MS MARCO with hard negatives from distilbert-margin-mse-cls-dot-v1 MS MARCO is a large scale information retrieval corpus that was created based on real user search queries using the Bing search engine. For each query and gold positive passage, the 50 most similar paragraphs were mined using 13 different models. The resulting data can be used to train Sentence Transformer models. Related Datasets These are the datasets generated using the 13 different models:… See the full description on the dataset page: https://huggingface.co/datasets/sentence-transformers/msmarco-distilbert-margin-mse-cls-dot-v1.tabularfeature-extraction10M<n<100M0 likes1.7k downloads2y agoHugging Face05sentence-transformers /msmarco-msmarco-distilbert-base-tas-b MS MARCO with hard negatives from msmarco-distilbert-base-tas-b MS MARCO is a large scale information retrieval corpus that was created based on real user search queries using the Bing search engine. For each query and gold positive passage, the 50 most similar paragraphs were mined using 13 different models. The resulting data can be used to train Sentence Transformer models. Related Datasets These are the datasets generated using the 13 different models:… See the full description on the dataset page: https://huggingface.co/datasets/sentence-transformers/msmarco-msmarco-distilbert-base-tas-b.tabularfeature-extraction10M<n<100M5 likes1.7k downloads2y agoHugging Face06sentence-transformers /msmarco-distilbert-margin-mse-sym-mnrl-mean-v1 MS MARCO with hard negatives from distilbert-margin-mse-sym-mnrl-mean-v1 MS MARCO is a large scale information retrieval corpus that was created based on real user search queries using the Bing search engine. For each query and gold positive passage, the 50 most similar paragraphs were mined using 13 different models. The resulting data can be used to train Sentence Transformer models. Related Datasets These are the datasets generated using the 13 different models:… See the full description on the dataset page: https://huggingface.co/datasets/sentence-transformers/msmarco-distilbert-margin-mse-sym-mnrl-mean-v1.tabularfeature-extraction10M<n<100M1 likes1.4k downloads2y agoHugging Face07sentence-transformers /msmarco-distilbert-margin-mse-mnrl-mean-v1 MS MARCO with hard negatives from distilbert-margin-mse-mnrl-mean-v1 MS MARCO is a large scale information retrieval corpus that was created based on real user search queries using the Bing search engine. For each query and gold positive passage, the 50 most similar paragraphs were mined using 13 different models. The resulting data can be used to train Sentence Transformer models. Related Datasets These are the datasets generated using the 13 different models:… See the full description on the dataset page: https://huggingface.co/datasets/sentence-transformers/msmarco-distilbert-margin-mse-mnrl-mean-v1.tabularfeature-extraction10M<n<100M0 likes822 downloads2y agoHugging Face08lsr42 /msmarco-psgs-distilbert-dot-v5text1M<n<10M0 likes636 downloads2y agoHugging Face09WendyHoang /news-ka-shuffled-DISTILBERTtext1M<n<10M0 likes254 downloads2y agoHugging Face10sentence-transformers /msmarco-distilbert-margin-mse-cls-dot-v2 MS MARCO with hard negatives from distilbert-margin-mse-cls-dot-v2 MS MARCO is a large scale information retrieval corpus that was created based on real user search queries using the Bing search engine. For each query and gold positive passage, the 50 most similar paragraphs were mined using 13 different models. The resulting data can be used to train Sentence Transformer models. Related Datasets These are the datasets generated using the 13 different models:… See the full description on the dataset page: https://huggingface.co/datasets/sentence-transformers/msmarco-distilbert-margin-mse-cls-dot-v2.tabularfeature-extraction10M<n<100M2 likes217 downloads2y agoHugging Face11WendyHoang /news-ka-small-DISTILBERTtext100K<n<1M0 likes80 downloads2y agoHugging Face12open-llm-leaderboard /distilbert__distilgpt2-detailsgated Dataset Card for Evaluation run of distilbert/distilgpt2 Dataset automatically created during the evaluation run of model distilbert/distilgpt2 The dataset is composed of 44 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results. An additional… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/distilbert__distilgpt2-details.tabular10K<n<100K1 likes54 downloads2y agoHugging Face13baharehansari1 /distilbert_spelling_dataset-v2text10K<n<100K0 likes46 downloads24d agoHugging Face14baharehansari1 /distilbert_spelling_dataset-v3text10K<n<100K0 likes44 downloads24d agoHugging Face15baharehansari1 /distilbert_spelling_datasettext10K<n<100K0 likes43 downloads27d agoHugging Face16Khalyie /sst2-distilbert-data SST-2 (GLUE) — Raw Splits Used for Khalyie/sst2-distilbert Same train/validation/test splits used to fine-tune Khalyie/sst2-distilbert. Note: the test split's label column is -1 for every row — GLUE withholds official test labels. Use validation for labeled evaluation. Reload with: from datasets import load_dataset ds = load_dataset("Khalyie/sst2-distilbert-data") text10K<n<100K0 likes38 downloads11d agoHugging Face17bzhao18 /hyperpartisan-news-distilbert-tokenstext100K<n<1M0 likes28 downloads2y agoHugging Face18AmjaadXX /sentiment-analysis-distilbert-rotten-tomatoes Sentiment Analysis NLP Pipeline (Rotten Tomatoes) This project builds a complete NLP data processing pipeline for sentiment analysis using the Rotten Tomatoes dataset. The workflow focuses on: Data exploration Data cleaning Feature engineering Tokenization and model preparation No model training has been performed yet. This project prepares the dataset for training transformer-based models such as DistilBERT. Dataset We use the Rotten Tomatoes dataset from… See the full description on the dataset page: https://huggingface.co/datasets/AmjaadXX/sentiment-analysis-distilbert-rotten-tomatoes.text10K<n<100K0 likes28 downloads5mo agoHugging Face19nikchar /retrieval_verification_bm25_distilbert Dataset Card for "retrieval_verification_bm25_distilbert" More Information needed tabular10K<n<100K0 likes24 downloads3y agoHugging Face20nikchar /retrieval_verification_distilbert Dataset Card for "retrieval_verification_distilbert" More Information needed tabular10K<n<100K0 likes22 downloads3y agoHugging Face21monostate /fintech-sentiment-distilbert-balanced-v2 Dataset Card for monostate/fintech-sentiment-distilbert-balanced-v2 Dataset Description This dataset was generated using Vibe Data Director, a tool for creating and curating text classification datasets. Dataset Summary Session ID: session_f15db25a Generated: 2026-03-16T13:09:50.740014 Total Samples: 481 Classes: negative, positive, neutral Styles: none Dataset Structure Data Fields text (string): The text content of the sample class… See the full description on the dataset page: https://huggingface.co/datasets/monostate/fintech-sentiment-distilbert-balanced-v2.texttext-classificationn<1K0 likes15 downloads6mo agoHugging Face22johannes-garstenauer /embeddings_from_distilbert_class_heaps_and_eval_part0 Dataset Card for "embeddings_from_distilbert_class_heaps_and_eval_part0" More Information needed tabular100K<n<1M0 likes13 downloads3y agoHugging Face23scademy /autotrain-data-DistilBert-500-500 Dataset Card for "autotrain-data-DistilBert-500-500" More Information needed text1K<n<10K0 likes12 downloads3y agoHugging Face24johannes-garstenauer /embeddings_from_distilbert_class_heaps Dataset Card for "embeddings_from_distilbert_class_heaps" Dataset created for thesis: "Generating Robust Representations of Structures in OpenSSH Heap Dumps" by Johannes Garstenauer. This dataset contains representations of heap data structures along with their labels and the predicted label. The representations are the [CLS] token embeddings of the last 3 layers of the DistilBERT model. The representation-generating model is:… See the full description on the dataset page: https://huggingface.co/datasets/johannes-garstenauer/embeddings_from_distilbert_class_heaps.tabular100K<n<1M1 likes11 downloads3y agoHugging Face25monostate /fintech-sentiment-distilbert-ready Dataset Card for monostate/fintech-sentiment-distilbert-ready Dataset Description This dataset was generated using Vibe Data Director, a tool for creating and curating text classification datasets. Dataset Summary Session ID: session_7987dfd2 Generated: 2026-03-16T12:59:16.439137 Total Samples: 406 Classes: positive, negative, neutral Styles: none Dataset Structure Data Fields text (string): The text content of the sample class… See the full description on the dataset page: https://huggingface.co/datasets/monostate/fintech-sentiment-distilbert-ready.texttext-classificationn<1K0 likes11 downloads6mo agoHugging Face26johannes-garstenauer /embeddings_from_distilbert_masking_heaps_and_eval_part0 Dataset Card for "embeddings_from_distilbert_masking_heaps_and_eval_part0" More Information needed tabular100K<n<1M0 likes10 downloads3y agoHugging Face27johannes-garstenauer /embeddings_from_distilbert_masking_heaps Dataset Card for "embeddings_from_distilbert_masking_heaps" Dataset created for thesis: "Generating Robust Representations of Structures in OpenSSH Heap Dumps" by Johannes Garstenauer. This dataset contains representations of heap data structures along with their labels and the predicted label. The representations are the [CLS] token embeddings of the last 3 layers of the DistilBERT model. The representation-generating model is:… See the full description on the dataset page: https://huggingface.co/datasets/johannes-garstenauer/embeddings_from_distilbert_masking_heaps.tabular100K<n<1M1 likes10 downloads3y agoHugging Face28interneuronai /companyx_customer_support_ticket_routing_distilbert_dataset CompanyX Customer Support Ticket Routing Description: Automatically route customer support tickets to relevant teams based on issue descriptions, speeding up resolution time and enhancing customer experience. How to Use Here is how to use this model to classify text into different categories: from transformers import AutoModelForSequenceClassification, AutoTokenizer model_name = "interneuronai/companyx_customer_support_ticket_routing_distilbert" model =… See the full description on the dataset page: https://huggingface.co/datasets/interneuronai/companyx_customer_support_ticket_routing_distilbert_dataset.tabular1K<n<10K0 likes10 downloads2y agoHugging Face29rebego /imdb-distilbert-tokenized path: data/unsupervised-* Descripción Este dataset contiene reseñas de películas tomadas del dataset de IMDB. Las reseñas están clasificadas en dos categorías: positivas y negativas. El dataset ha sido procesado y tokenizado utilizando el modelo preentrenado distilbert-base-uncased-finetuned-sst-2-english con la librería Transformers. Este dataset tokenizado es adecuado para entrenar modelos de aprendizaje automático, en tareas de análisis de sentimientos. text100K<n<1M0 likes10 downloads2y agoHugging Face30eliodecolli /distilbert-learning-feedbacktabularn<1K0 likes10 downloads1mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.