CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01sentence-transformers /msmarco-distilbert-margin-mse-mean-dot-v1 MS MARCO with hard negatives from distilbert-margin-mse-mean-dot-v1 MS MARCO is a large scale information retrieval corpus that was created based on real user search queries using the Bing search engine. For each query and gold positive passage, the 50 most similar paragraphs were mined using 13 different models. The resulting data can be used to train Sentence Transformer models. Related Datasets These are the datasets generated using the 13 different models:… See the full description on the dataset page: https://huggingface.co/datasets/sentence-transformers/msmarco-distilbert-margin-mse-mean-dot-v1.tabularfeature-extraction10M<n<100M2 likes3.4k downloads2y agoHugging Face02sentence-transformers /msmarco-msmarco-distilbert-base-v3 MS MARCO with hard negatives from msmarco-distilbert-base-v3 MS MARCO is a large scale information retrieval corpus that was created based on real user search queries using the Bing search engine. For each query and gold positive passage, the 50 most similar paragraphs were mined using 13 different models. The resulting data can be used to train Sentence Transformer models. Related Datasets These are the datasets generated using the 13 different models: msmarco-bm25… See the full description on the dataset page: https://huggingface.co/datasets/sentence-transformers/msmarco-msmarco-distilbert-base-v3.tabularfeature-extraction10M<n<100M5 likes2.4k downloads2y agoHugging Face03sentence-transformers /msmarco-distilbert-margin-mse-sym-mnrl-mean-v2 MS MARCO with hard negatives from distilbert-margin-mse-sym-mnrl-mean-v2 MS MARCO is a large scale information retrieval corpus that was created based on real user search queries using the Bing search engine. For each query and gold positive passage, the 50 most similar paragraphs were mined using 13 different models. The resulting data can be used to train Sentence Transformer models. Related Datasets These are the datasets generated using the 13 different models:… See the full description on the dataset page: https://huggingface.co/datasets/sentence-transformers/msmarco-distilbert-margin-mse-sym-mnrl-mean-v2.tabularfeature-extraction10M<n<100M1 likes2.3k downloads2y agoHugging Face04sentence-transformers /msmarco-msmarco-distilbert-base-tas-b MS MARCO with hard negatives from msmarco-distilbert-base-tas-b MS MARCO is a large scale information retrieval corpus that was created based on real user search queries using the Bing search engine. For each query and gold positive passage, the 50 most similar paragraphs were mined using 13 different models. The resulting data can be used to train Sentence Transformer models. Related Datasets These are the datasets generated using the 13 different models:… See the full description on the dataset page: https://huggingface.co/datasets/sentence-transformers/msmarco-msmarco-distilbert-base-tas-b.tabularfeature-extraction10M<n<100M5 likes1.8k downloads2y agoHugging Face05sentence-transformers /msmarco-distilbert-margin-mse-cls-dot-v1 MS MARCO with hard negatives from distilbert-margin-mse-cls-dot-v1 MS MARCO is a large scale information retrieval corpus that was created based on real user search queries using the Bing search engine. For each query and gold positive passage, the 50 most similar paragraphs were mined using 13 different models. The resulting data can be used to train Sentence Transformer models. Related Datasets These are the datasets generated using the 13 different models:… See the full description on the dataset page: https://huggingface.co/datasets/sentence-transformers/msmarco-distilbert-margin-mse-cls-dot-v1.tabularfeature-extraction10M<n<100M0 likes1.5k downloads2y agoHugging Face06sentence-transformers /msmarco-distilbert-margin-mse-sym-mnrl-mean-v1 MS MARCO with hard negatives from distilbert-margin-mse-sym-mnrl-mean-v1 MS MARCO is a large scale information retrieval corpus that was created based on real user search queries using the Bing search engine. For each query and gold positive passage, the 50 most similar paragraphs were mined using 13 different models. The resulting data can be used to train Sentence Transformer models. Related Datasets These are the datasets generated using the 13 different models:… See the full description on the dataset page: https://huggingface.co/datasets/sentence-transformers/msmarco-distilbert-margin-mse-sym-mnrl-mean-v1.tabularfeature-extraction10M<n<100M1 likes1.5k downloads2y agoHugging Face07sentence-transformers /msmarco-distilbert-margin-mse-mnrl-mean-v1 MS MARCO with hard negatives from distilbert-margin-mse-mnrl-mean-v1 MS MARCO is a large scale information retrieval corpus that was created based on real user search queries using the Bing search engine. For each query and gold positive passage, the 50 most similar paragraphs were mined using 13 different models. The resulting data can be used to train Sentence Transformer models. Related Datasets These are the datasets generated using the 13 different models:… See the full description on the dataset page: https://huggingface.co/datasets/sentence-transformers/msmarco-distilbert-margin-mse-mnrl-mean-v1.tabularfeature-extraction10M<n<100M0 likes857 downloads2y agoHugging Face08lsr42 /msmarco-psgs-distilbert-dot-v5text1M<n<10M0 likes451 downloads2y agoHugging Face09WendyHoang /news-ka-shuffled-DISTILBERTtext1M<n<10M0 likes420 downloads2y agoHugging Face10sentence-transformers /msmarco-distilbert-margin-mse-cls-dot-v2 MS MARCO with hard negatives from distilbert-margin-mse-cls-dot-v2 MS MARCO is a large scale information retrieval corpus that was created based on real user search queries using the Bing search engine. For each query and gold positive passage, the 50 most similar paragraphs were mined using 13 different models. The resulting data can be used to train Sentence Transformer models. Related Datasets These are the datasets generated using the 13 different models:… See the full description on the dataset page: https://huggingface.co/datasets/sentence-transformers/msmarco-distilbert-margin-mse-cls-dot-v2.tabularfeature-extraction10M<n<100M2 likes237 downloads2y agoHugging Face11WendyHoang /news-ka-small-DISTILBERTtext100K<n<1M0 likes104 downloads2y agoHugging Face12open-llm-leaderboard /distilbert__distilgpt2-detailsgated Dataset Card for Evaluation run of distilbert/distilgpt2 Dataset automatically created during the evaluation run of model distilbert/distilgpt2 The dataset is composed of 44 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results. An additional… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/distilbert__distilgpt2-details.tabular10K<n<100K1 likes67 downloads2y agoHugging Face13baharehansari1 /distilbert_spelling_dataset-v2text10K<n<100K0 likes47 downloads27d agoHugging Face14baharehansari1 /distilbert_spelling_dataset-v3text10K<n<100K0 likes45 downloads27d agoHugging Face15Khalyie /sst2-distilbert-data SST-2 (GLUE) — Raw Splits Used for Khalyie/sst2-distilbert Same train/validation/test splits used to fine-tune Khalyie/sst2-distilbert. Note: the test split's label column is -1 for every row — GLUE withholds official test labels. Use validation for labeled evaluation. Reload with: from datasets import load_dataset ds = load_dataset("Khalyie/sst2-distilbert-data") text10K<n<100K0 likes39 downloads14d agoHugging Face16bzhao18 /hyperpartisan-news-distilbert-tokenstext100K<n<1M0 likes30 downloads2y agoHugging Face17baharehansari1 /distilbert_spelling_datasettext10K<n<100K0 likes24 downloads1mo agoHugging Face18nikchar /retrieval_verification_distilbert Dataset Card for "retrieval_verification_distilbert" More Information needed tabular10K<n<100K0 likes22 downloads3y agoHugging Face19AmjaadXX /sentiment-analysis-distilbert-rotten-tomatoes Sentiment Analysis NLP Pipeline (Rotten Tomatoes) This project builds a complete NLP data processing pipeline for sentiment analysis using the Rotten Tomatoes dataset. The workflow focuses on: Data exploration Data cleaning Feature engineering Tokenization and model preparation No model training has been performed yet. This project prepares the dataset for training transformer-based models such as DistilBERT. Dataset We use the Rotten Tomatoes dataset from… See the full description on the dataset page: https://huggingface.co/datasets/AmjaadXX/sentiment-analysis-distilbert-rotten-tomatoes.text10K<n<100K0 likes21 downloads5mo agoHugging Face20nikchar /retrieval_verification_bm25_distilbert Dataset Card for "retrieval_verification_bm25_distilbert" More Information needed tabular10K<n<100K0 likes19 downloads3y agoHugging Face21monostate /fintech-sentiment-distilbert-balanced-v2 Dataset Card for monostate/fintech-sentiment-distilbert-balanced-v2 Dataset Description This dataset was generated using Vibe Data Director, a tool for creating and curating text classification datasets. Dataset Summary Session ID: session_f15db25a Generated: 2026-03-16T13:09:50.740014 Total Samples: 481 Classes: negative, positive, neutral Styles: none Dataset Structure Data Fields text (string): The text content of the sample class… See the full description on the dataset page: https://huggingface.co/datasets/monostate/fintech-sentiment-distilbert-balanced-v2.texttext-classificationn<1K0 likes15 downloads6mo agoHugging Face22scademy /autotrain-data-DistilBert-500-500 Dataset Card for "autotrain-data-DistilBert-500-500" More Information needed text1K<n<10K0 likes14 downloads3y agoHugging Face23monostate /fintech-sentiment-distilbert-ready Dataset Card for monostate/fintech-sentiment-distilbert-ready Dataset Description This dataset was generated using Vibe Data Director, a tool for creating and curating text classification datasets. Dataset Summary Session ID: session_7987dfd2 Generated: 2026-03-16T12:59:16.439137 Total Samples: 406 Classes: positive, negative, neutral Styles: none Dataset Structure Data Fields text (string): The text content of the sample class… See the full description on the dataset page: https://huggingface.co/datasets/monostate/fintech-sentiment-distilbert-ready.texttext-classificationn<1K0 likes14 downloads6mo agoHugging Face24RVMadhu /distil_berttexttext-classification0 likes11 downloads3y agoHugging Face25johannes-garstenauer /embeddings_from_distilbert_class_heaps_and_eval_part0 Dataset Card for "embeddings_from_distilbert_class_heaps_and_eval_part0" More Information needed tabular100K<n<1M0 likes11 downloads3y agoHugging Face26johannes-garstenauer /embeddings_from_distilbert_class_heaps Dataset Card for "embeddings_from_distilbert_class_heaps" Dataset created for thesis: "Generating Robust Representations of Structures in OpenSSH Heap Dumps" by Johannes Garstenauer. This dataset contains representations of heap data structures along with their labels and the predicted label. The representations are the [CLS] token embeddings of the last 3 layers of the DistilBERT model. The representation-generating model is:… See the full description on the dataset page: https://huggingface.co/datasets/johannes-garstenauer/embeddings_from_distilbert_class_heaps.tabular100K<n<1M1 likes11 downloads3y agoHugging Face27rebego /imdb-distilbert-tokenized path: data/unsupervised-* Descripción Este dataset contiene reseñas de películas tomadas del dataset de IMDB. Las reseñas están clasificadas en dos categorías: positivas y negativas. El dataset ha sido procesado y tokenizado utilizando el modelo preentrenado distilbert-base-uncased-finetuned-sst-2-english con la librería Transformers. Este dataset tokenizado es adecuado para entrenar modelos de aprendizaje automático, en tareas de análisis de sentimientos. text100K<n<1M0 likes11 downloads2y agoHugging Face28Parth1612 /pp_distilbert_ft_emotionstext10K<n<100K0 likes10 downloads3y agoHugging Face29tsch00001 /news-ka-shuffled-DISTILBERT-testtextn<1K0 likes10 downloads2y agoHugging Face30eliodecolli /distilbert-learning-feedbacktabularn<1K0 likes10 downloads1mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.