datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
msmarco-distilbert-margin-mse-mean-dot-v1
MS MARCO with hard negatives from distilbert-margin-mse-mean-dot-v1
MS MARCO is a large scale information retrieval corpus that was created based on real user search queries using the Bing search engine.
For each query and gold positive passage, the 50 most similar paragraphs were mined using 13 different models. The resulting data can be used to train Sentence Transformer models.
Related Datasets
These are the datasets generated using the 13 different models:… See the full description on the dataset page: https://huggingface.co/datasets/sentence-transformers/msmarco-distilbert-margin-mse-mean-dot-v1.msmarco-msmarco-distilbert-base-v3
MS MARCO with hard negatives from msmarco-distilbert-base-v3
MS MARCO is a large scale information retrieval corpus that was created based on real user search queries using the Bing search engine.
For each query and gold positive passage, the 50 most similar paragraphs were mined using 13 different models. The resulting data can be used to train Sentence Transformer models.
Related Datasets
These are the datasets generated using the 13 different models:
msmarco-bm25… See the full description on the dataset page: https://huggingface.co/datasets/sentence-transformers/msmarco-msmarco-distilbert-base-v3.msmarco-distilbert-margin-mse-sym-mnrl-mean-v2
MS MARCO with hard negatives from distilbert-margin-mse-sym-mnrl-mean-v2
MS MARCO is a large scale information retrieval corpus that was created based on real user search queries using the Bing search engine.
For each query and gold positive passage, the 50 most similar paragraphs were mined using 13 different models. The resulting data can be used to train Sentence Transformer models.
Related Datasets
These are the datasets generated using the 13 different models:… See the full description on the dataset page: https://huggingface.co/datasets/sentence-transformers/msmarco-distilbert-margin-mse-sym-mnrl-mean-v2.msmarco-distilbert-margin-mse-cls-dot-v1
MS MARCO with hard negatives from distilbert-margin-mse-cls-dot-v1
MS MARCO is a large scale information retrieval corpus that was created based on real user search queries using the Bing search engine.
For each query and gold positive passage, the 50 most similar paragraphs were mined using 13 different models. The resulting data can be used to train Sentence Transformer models.
Related Datasets
These are the datasets generated using the 13 different models:… See the full description on the dataset page: https://huggingface.co/datasets/sentence-transformers/msmarco-distilbert-margin-mse-cls-dot-v1.msmarco-msmarco-distilbert-base-tas-b
MS MARCO with hard negatives from msmarco-distilbert-base-tas-b
MS MARCO is a large scale information retrieval corpus that was created based on real user search queries using the Bing search engine.
For each query and gold positive passage, the 50 most similar paragraphs were mined using 13 different models. The resulting data can be used to train Sentence Transformer models.
Related Datasets
These are the datasets generated using the 13 different models:… See the full description on the dataset page: https://huggingface.co/datasets/sentence-transformers/msmarco-msmarco-distilbert-base-tas-b.msmarco-distilbert-margin-mse-sym-mnrl-mean-v1
MS MARCO with hard negatives from distilbert-margin-mse-sym-mnrl-mean-v1
MS MARCO is a large scale information retrieval corpus that was created based on real user search queries using the Bing search engine.
For each query and gold positive passage, the 50 most similar paragraphs were mined using 13 different models. The resulting data can be used to train Sentence Transformer models.
Related Datasets
These are the datasets generated using the 13 different models:… See the full description on the dataset page: https://huggingface.co/datasets/sentence-transformers/msmarco-distilbert-margin-mse-sym-mnrl-mean-v1.msmarco-distilbert-margin-mse-mnrl-mean-v1
MS MARCO with hard negatives from distilbert-margin-mse-mnrl-mean-v1
MS MARCO is a large scale information retrieval corpus that was created based on real user search queries using the Bing search engine.
For each query and gold positive passage, the 50 most similar paragraphs were mined using 13 different models. The resulting data can be used to train Sentence Transformer models.
Related Datasets
These are the datasets generated using the 13 different models:… See the full description on the dataset page: https://huggingface.co/datasets/sentence-transformers/msmarco-distilbert-margin-mse-mnrl-mean-v1.msmarco-psgs-distilbert-dot-v5news-ka-shuffled-DISTILBERTmsmarco-distilbert-margin-mse-cls-dot-v2
MS MARCO with hard negatives from distilbert-margin-mse-cls-dot-v2
MS MARCO is a large scale information retrieval corpus that was created based on real user search queries using the Bing search engine.
For each query and gold positive passage, the 50 most similar paragraphs were mined using 13 different models. The resulting data can be used to train Sentence Transformer models.
Related Datasets
These are the datasets generated using the 13 different models:… See the full description on the dataset page: https://huggingface.co/datasets/sentence-transformers/msmarco-distilbert-margin-mse-cls-dot-v2.news-ka-small-DISTILBERTdistilbert__distilgpt2-details
Dataset Card for Evaluation run of distilbert/distilgpt2
Dataset automatically created during the evaluation run of model distilbert/distilgpt2
The dataset is composed of 44 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An additional… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/distilbert__distilgpt2-details.distilbert_spelling_dataset-v2distilbert_spelling_dataset-v3distilbert_spelling_datasetsst2-distilbert-data
SST-2 (GLUE) — Raw Splits Used for Khalyie/sst2-distilbert
Same train/validation/test splits used to fine-tune
Khalyie/sst2-distilbert.
Note: the test split's label column is -1 for every row — GLUE
withholds official test labels. Use validation for labeled evaluation.
Reload with:
from datasets import load_dataset
ds = load_dataset("Khalyie/sst2-distilbert-data")
hyperpartisan-news-distilbert-tokenssentiment-analysis-distilbert-rotten-tomatoes
Sentiment Analysis NLP Pipeline (Rotten Tomatoes)
This project builds a complete NLP data processing pipeline for sentiment analysis using the Rotten Tomatoes dataset.
The workflow focuses on:
Data exploration
Data cleaning
Feature engineering
Tokenization and model preparation
No model training has been performed yet. This project prepares the dataset for training transformer-based models such as DistilBERT.
Dataset
We use the Rotten Tomatoes dataset from… See the full description on the dataset page: https://huggingface.co/datasets/AmjaadXX/sentiment-analysis-distilbert-rotten-tomatoes.retrieval_verification_bm25_distilbert
Dataset Card for "retrieval_verification_bm25_distilbert"
More Information needed
retrieval_verification_distilbert
Dataset Card for "retrieval_verification_distilbert"
More Information needed
fintech-sentiment-distilbert-balanced-v2
Dataset Card for monostate/fintech-sentiment-distilbert-balanced-v2
Dataset Description
This dataset was generated using Vibe Data Director, a tool for creating and curating text classification datasets.
Dataset Summary
Session ID: session_f15db25a
Generated: 2026-03-16T13:09:50.740014
Total Samples: 481
Classes: negative, positive, neutral
Styles: none
Dataset Structure
Data Fields
text (string): The text content of the sample
class… See the full description on the dataset page: https://huggingface.co/datasets/monostate/fintech-sentiment-distilbert-balanced-v2.embeddings_from_distilbert_class_heaps_and_eval_part0
Dataset Card for "embeddings_from_distilbert_class_heaps_and_eval_part0"
More Information needed
autotrain-data-DistilBert-500-500
Dataset Card for "autotrain-data-DistilBert-500-500"
More Information needed
embeddings_from_distilbert_class_heaps
Dataset Card for "embeddings_from_distilbert_class_heaps"
Dataset created for thesis: "Generating Robust Representations of Structures in OpenSSH Heap Dumps" by Johannes Garstenauer.
This dataset contains representations of heap data structures along with their labels and the predicted label.
The representations are the [CLS] token embeddings of the last 3 layers of the DistilBERT model.
The representation-generating model is:… See the full description on the dataset page: https://huggingface.co/datasets/johannes-garstenauer/embeddings_from_distilbert_class_heaps.fintech-sentiment-distilbert-ready
Dataset Card for monostate/fintech-sentiment-distilbert-ready
Dataset Description
This dataset was generated using Vibe Data Director, a tool for creating and curating text classification datasets.
Dataset Summary
Session ID: session_7987dfd2
Generated: 2026-03-16T12:59:16.439137
Total Samples: 406
Classes: positive, negative, neutral
Styles: none
Dataset Structure
Data Fields
text (string): The text content of the sample
class… See the full description on the dataset page: https://huggingface.co/datasets/monostate/fintech-sentiment-distilbert-ready.embeddings_from_distilbert_masking_heaps_and_eval_part0
Dataset Card for "embeddings_from_distilbert_masking_heaps_and_eval_part0"
More Information needed
embeddings_from_distilbert_masking_heaps
Dataset Card for "embeddings_from_distilbert_masking_heaps"
Dataset created for thesis: "Generating Robust Representations of Structures in OpenSSH Heap Dumps" by Johannes Garstenauer.
This dataset contains representations of heap data structures along with their labels and the predicted label.
The representations are the [CLS] token embeddings of the last 3 layers of the DistilBERT model.
The representation-generating model is:… See the full description on the dataset page: https://huggingface.co/datasets/johannes-garstenauer/embeddings_from_distilbert_masking_heaps.companyx_customer_support_ticket_routing_distilbert_dataset
CompanyX Customer Support Ticket Routing
Description: Automatically route customer support tickets to relevant teams based on issue descriptions, speeding up resolution time and enhancing customer experience.
How to Use
Here is how to use this model to classify text into different categories:
from transformers import AutoModelForSequenceClassification, AutoTokenizer
model_name = "interneuronai/companyx_customer_support_ticket_routing_distilbert"
model =… See the full description on the dataset page: https://huggingface.co/datasets/interneuronai/companyx_customer_support_ticket_routing_distilbert_dataset.imdb-distilbert-tokenized
path: data/unsupervised-*
Descripción
Este dataset contiene reseñas de películas tomadas del dataset de IMDB. Las reseñas están clasificadas en dos categorías: positivas y negativas. El dataset ha sido procesado y tokenizado utilizando el modelo preentrenado distilbert-base-uncased-finetuned-sst-2-english con la librería Transformers.
Este dataset tokenizado es adecuado para entrenar modelos de aprendizaje automático, en tareas de análisis de sentimientos.
distilbert-learning-feedback
