CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01sentence-transformers /msmarco-distilbert-margin-mse-mean-dot-v1 MS MARCO with hard negatives from distilbert-margin-mse-mean-dot-v1 MS MARCO is a large scale information retrieval corpus that was created based on real user search queries using the Bing search engine. For each query and gold positive passage, the 50 most similar paragraphs were mined using 13 different models. The resulting data can be used to train Sentence Transformer models. Related Datasets These are the datasets generated using the 13 different models:… See the full description on the dataset page: https://huggingface.co/datasets/sentence-transformers/msmarco-distilbert-margin-mse-mean-dot-v1.tabularfeature-extraction10M<n<100M2 likes3.4k downloads2y agoHugging Face02sentence-transformers /msmarco-msmarco-distilbert-base-v3 MS MARCO with hard negatives from msmarco-distilbert-base-v3 MS MARCO is a large scale information retrieval corpus that was created based on real user search queries using the Bing search engine. For each query and gold positive passage, the 50 most similar paragraphs were mined using 13 different models. The resulting data can be used to train Sentence Transformer models. Related Datasets These are the datasets generated using the 13 different models: msmarco-bm25… See the full description on the dataset page: https://huggingface.co/datasets/sentence-transformers/msmarco-msmarco-distilbert-base-v3.tabularfeature-extraction10M<n<100M5 likes2.4k downloads2y agoHugging Face03sentence-transformers /msmarco-distilbert-margin-mse-sym-mnrl-mean-v2 MS MARCO with hard negatives from distilbert-margin-mse-sym-mnrl-mean-v2 MS MARCO is a large scale information retrieval corpus that was created based on real user search queries using the Bing search engine. For each query and gold positive passage, the 50 most similar paragraphs were mined using 13 different models. The resulting data can be used to train Sentence Transformer models. Related Datasets These are the datasets generated using the 13 different models:… See the full description on the dataset page: https://huggingface.co/datasets/sentence-transformers/msmarco-distilbert-margin-mse-sym-mnrl-mean-v2.tabularfeature-extraction10M<n<100M1 likes2.2k downloads2y agoHugging Face04sentence-transformers /msmarco-distilbert-margin-mse-cls-dot-v1 MS MARCO with hard negatives from distilbert-margin-mse-cls-dot-v1 MS MARCO is a large scale information retrieval corpus that was created based on real user search queries using the Bing search engine. For each query and gold positive passage, the 50 most similar paragraphs were mined using 13 different models. The resulting data can be used to train Sentence Transformer models. Related Datasets These are the datasets generated using the 13 different models:… See the full description on the dataset page: https://huggingface.co/datasets/sentence-transformers/msmarco-distilbert-margin-mse-cls-dot-v1.tabularfeature-extraction10M<n<100M0 likes1.7k downloads2y agoHugging Face05sentence-transformers /msmarco-msmarco-distilbert-base-tas-b MS MARCO with hard negatives from msmarco-distilbert-base-tas-b MS MARCO is a large scale information retrieval corpus that was created based on real user search queries using the Bing search engine. For each query and gold positive passage, the 50 most similar paragraphs were mined using 13 different models. The resulting data can be used to train Sentence Transformer models. Related Datasets These are the datasets generated using the 13 different models:… See the full description on the dataset page: https://huggingface.co/datasets/sentence-transformers/msmarco-msmarco-distilbert-base-tas-b.tabularfeature-extraction10M<n<100M5 likes1.7k downloads2y agoHugging Face06sentence-transformers /msmarco-distilbert-margin-mse-sym-mnrl-mean-v1 MS MARCO with hard negatives from distilbert-margin-mse-sym-mnrl-mean-v1 MS MARCO is a large scale information retrieval corpus that was created based on real user search queries using the Bing search engine. For each query and gold positive passage, the 50 most similar paragraphs were mined using 13 different models. The resulting data can be used to train Sentence Transformer models. Related Datasets These are the datasets generated using the 13 different models:… See the full description on the dataset page: https://huggingface.co/datasets/sentence-transformers/msmarco-distilbert-margin-mse-sym-mnrl-mean-v1.tabularfeature-extraction10M<n<100M1 likes1.4k downloads2y agoHugging Face07sentence-transformers /msmarco-distilbert-margin-mse-mnrl-mean-v1 MS MARCO with hard negatives from distilbert-margin-mse-mnrl-mean-v1 MS MARCO is a large scale information retrieval corpus that was created based on real user search queries using the Bing search engine. For each query and gold positive passage, the 50 most similar paragraphs were mined using 13 different models. The resulting data can be used to train Sentence Transformer models. Related Datasets These are the datasets generated using the 13 different models:… See the full description on the dataset page: https://huggingface.co/datasets/sentence-transformers/msmarco-distilbert-margin-mse-mnrl-mean-v1.tabularfeature-extraction10M<n<100M0 likes822 downloads2y agoHugging Face08lsr42 /msmarco-psgs-distilbert-dot-v5text1M<n<10M0 likes636 downloads2y agoHugging Face09WendyHoang /news-ka-shuffled-DISTILBERTtext1M<n<10M0 likes254 downloads2y agoHugging Face10sentence-transformers /msmarco-distilbert-margin-mse-cls-dot-v2 MS MARCO with hard negatives from distilbert-margin-mse-cls-dot-v2 MS MARCO is a large scale information retrieval corpus that was created based on real user search queries using the Bing search engine. For each query and gold positive passage, the 50 most similar paragraphs were mined using 13 different models. The resulting data can be used to train Sentence Transformer models. Related Datasets These are the datasets generated using the 13 different models:… See the full description on the dataset page: https://huggingface.co/datasets/sentence-transformers/msmarco-distilbert-margin-mse-cls-dot-v2.tabularfeature-extraction10M<n<100M2 likes217 downloads2y agoHugging Face11WendyHoang /news-ka-small-DISTILBERTtext100K<n<1M0 likes80 downloads2y agoHugging Face12open-llm-leaderboard /distilbert__distilgpt2-detailsgated Dataset Card for Evaluation run of distilbert/distilgpt2 Dataset automatically created during the evaluation run of model distilbert/distilgpt2 The dataset is composed of 44 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results. An additional… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/distilbert__distilgpt2-details.tabular10K<n<100K1 likes54 downloads2y agoHugging Face13feyninc /chonkiepedia-distilbert-tokenized1M<n<10M0 likes52 downloads1y agoHugging Face14broadfield-dev /model-atlas-distilbert-base-uncased0 likes48 downloads1y agoHugging Face15baharehansari1 /distilbert_spelling_dataset-v2text10K<n<100K0 likes46 downloads23d agoHugging Face16baharehansari1 /distilbert_spelling_dataset-v3text10K<n<100K0 likes44 downloads23d agoHugging Face17baharehansari1 /distilbert_spelling_datasettext10K<n<100K0 likes43 downloads26d agoHugging Face18Merlin041 /imdb-distilbert-features IMDb Movie Reviews - DistilBERT Contextual Embedding Cache This dataset contains pre-extracted contextual embedding features of the standard IMDb Movie Reviews dataset (Sentiment Analysis). The features were extracted using a frozen DistilBERT (distilbert-base-uncased) encoder. By caching these high-dimensional embeddings, you can train downstream classifiers (like LSTMs, GRUs, Attention heads, or custom Transformers) in seconds on a local GPU or CPU, bypassing the massive… See the full description on the dataset page: https://huggingface.co/datasets/Merlin041/imdb-distilbert-features.0 likes38 downloads2mo agoHugging Face19Khalyie /sst2-distilbert-data SST-2 (GLUE) — Raw Splits Used for Khalyie/sst2-distilbert Same train/validation/test splits used to fine-tune Khalyie/sst2-distilbert. Note: the test split's label column is -1 for every row — GLUE withholds official test labels. Use validation for labeled evaluation. Reload with: from datasets import load_dataset ds = load_dataset("Khalyie/sst2-distilbert-data") text10K<n<100K0 likes38 downloads10d agoHugging Face20Mmoorejoseph /distilbert-translate dataloader.py Dataset Summary A speech dataset with audio video modality, stored in lmdb format. Preprocessing & Augmentation Preprocessing: minimal Augmentation: mixup cutmix Splits & Sampling Split strategy: stratified 90 10 Sampling: contrastive Quality & Labeling Quality filtering: strict Labeling: pseudo label Files dataloader.py — main artifact of this repository License See… See the full description on the dataset page: https://huggingface.co/datasets/Mmoorejoseph/distilbert-translate.0 likes33 downloads28d agoHugging Face21Daffasari /distilbert-parser-finetune preprocess.py Dataset Summary A agriculture dataset with video text modality, stored in csv format. Preprocessing & Augmentation Preprocessing: progressive Augmentation: light Splits & Sampling Split strategy: random 90 10 Sampling: weighted Quality & Labeling Quality filtering: adaptive Labeling: pseudo label Files preprocess.py — main artifact of this repository License See the… See the full description on the dataset page: https://huggingface.co/datasets/Daffasari/distilbert-parser-finetune.0 likes31 downloads28d agoHugging Face22bzhao18 /hyperpartisan-news-distilbert-tokenstext100K<n<1M0 likes28 downloads2y agoHugging Face23AmjaadXX /sentiment-analysis-distilbert-rotten-tomatoes Sentiment Analysis NLP Pipeline (Rotten Tomatoes) This project builds a complete NLP data processing pipeline for sentiment analysis using the Rotten Tomatoes dataset. The workflow focuses on: Data exploration Data cleaning Feature engineering Tokenization and model preparation No model training has been performed yet. This project prepares the dataset for training transformer-based models such as DistilBERT. Dataset We use the Rotten Tomatoes dataset from… See the full description on the dataset page: https://huggingface.co/datasets/AmjaadXX/sentiment-analysis-distilbert-rotten-tomatoes.text10K<n<100K0 likes28 downloads5mo agoHugging Face24snyamson /covid-tweet-sentiment-analyzer-distilbert-data Dataset Card for "covid-tweet-sentiment-analyzer-distilbert-data" More Information needed 1K<n<10K1 likes25 downloads3y agoHugging Face25nikchar /retrieval_verification_bm25_distilbert Dataset Card for "retrieval_verification_bm25_distilbert" More Information needed tabular10K<n<100K0 likes24 downloads3y agoHugging Face26nikchar /retrieval_verification_distilbert Dataset Card for "retrieval_verification_distilbert" More Information needed tabular10K<n<100K0 likes22 downloads3y agoHugging Face27bambadij /Tweet_sentiment_analysis_Distilbert Dataset Card for "Tweet_sentiment_analysis_Distilbert" More Information needed 1K<n<10K0 likes18 downloads3y agoHugging Face28sukantan /nyaya-ae-msmarco-distilbert-base-tas-b Dataset Card for "nyaya-ae-msmarco-distilbert-base-tas-b" More Information needed tabular10K<n<100K0 likes17 downloads3y agoHugging Face29fantasticrambo /covid-tweet-sentiment-analyzer-distilbert-data1K<n<10K1 likes17 downloads3y agoHugging Face30monostate /fintech-sentiment-distilbert-balanced-v2 Dataset Card for monostate/fintech-sentiment-distilbert-balanced-v2 Dataset Description This dataset was generated using Vibe Data Director, a tool for creating and curating text classification datasets. Dataset Summary Session ID: session_f15db25a Generated: 2026-03-16T13:09:50.740014 Total Samples: 481 Classes: negative, positive, neutral Styles: none Dataset Structure Data Fields text (string): The text content of the sample class… See the full description on the dataset page: https://huggingface.co/datasets/monostate/fintech-sentiment-distilbert-balanced-v2.texttext-classificationn<1K0 likes15 downloads6mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.