datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
msmarco-distilbert-margin-mse-mean-dot-v1
MS MARCO with hard negatives from distilbert-margin-mse-mean-dot-v1
MS MARCO is a large scale information retrieval corpus that was created based on real user search queries using the Bing search engine.
For each query and gold positive passage, the 50 most similar paragraphs were mined using 13 different models. The resulting data can be used to train Sentence Transformer models.
Related Datasets
These are the datasets generated using the 13 different models:… See the full description on the dataset page: https://huggingface.co/datasets/sentence-transformers/msmarco-distilbert-margin-mse-mean-dot-v1.msmarco-msmarco-distilbert-base-v3
MS MARCO with hard negatives from msmarco-distilbert-base-v3
MS MARCO is a large scale information retrieval corpus that was created based on real user search queries using the Bing search engine.
For each query and gold positive passage, the 50 most similar paragraphs were mined using 13 different models. The resulting data can be used to train Sentence Transformer models.
Related Datasets
These are the datasets generated using the 13 different models:
msmarco-bm25… See the full description on the dataset page: https://huggingface.co/datasets/sentence-transformers/msmarco-msmarco-distilbert-base-v3.msmarco-distilbert-margin-mse-sym-mnrl-mean-v2
MS MARCO with hard negatives from distilbert-margin-mse-sym-mnrl-mean-v2
MS MARCO is a large scale information retrieval corpus that was created based on real user search queries using the Bing search engine.
For each query and gold positive passage, the 50 most similar paragraphs were mined using 13 different models. The resulting data can be used to train Sentence Transformer models.
Related Datasets
These are the datasets generated using the 13 different models:… See the full description on the dataset page: https://huggingface.co/datasets/sentence-transformers/msmarco-distilbert-margin-mse-sym-mnrl-mean-v2.msmarco-msmarco-distilbert-base-tas-b
MS MARCO with hard negatives from msmarco-distilbert-base-tas-b
MS MARCO is a large scale information retrieval corpus that was created based on real user search queries using the Bing search engine.
For each query and gold positive passage, the 50 most similar paragraphs were mined using 13 different models. The resulting data can be used to train Sentence Transformer models.
Related Datasets
These are the datasets generated using the 13 different models:… See the full description on the dataset page: https://huggingface.co/datasets/sentence-transformers/msmarco-msmarco-distilbert-base-tas-b.msmarco-distilbert-margin-mse-cls-dot-v1
MS MARCO with hard negatives from distilbert-margin-mse-cls-dot-v1
MS MARCO is a large scale information retrieval corpus that was created based on real user search queries using the Bing search engine.
For each query and gold positive passage, the 50 most similar paragraphs were mined using 13 different models. The resulting data can be used to train Sentence Transformer models.
Related Datasets
These are the datasets generated using the 13 different models:… See the full description on the dataset page: https://huggingface.co/datasets/sentence-transformers/msmarco-distilbert-margin-mse-cls-dot-v1.msmarco-distilbert-margin-mse-sym-mnrl-mean-v1
MS MARCO with hard negatives from distilbert-margin-mse-sym-mnrl-mean-v1
MS MARCO is a large scale information retrieval corpus that was created based on real user search queries using the Bing search engine.
For each query and gold positive passage, the 50 most similar paragraphs were mined using 13 different models. The resulting data can be used to train Sentence Transformer models.
Related Datasets
These are the datasets generated using the 13 different models:… See the full description on the dataset page: https://huggingface.co/datasets/sentence-transformers/msmarco-distilbert-margin-mse-sym-mnrl-mean-v1.msmarco-distilbert-margin-mse-mnrl-mean-v1
MS MARCO with hard negatives from distilbert-margin-mse-mnrl-mean-v1
MS MARCO is a large scale information retrieval corpus that was created based on real user search queries using the Bing search engine.
For each query and gold positive passage, the 50 most similar paragraphs were mined using 13 different models. The resulting data can be used to train Sentence Transformer models.
Related Datasets
These are the datasets generated using the 13 different models:… See the full description on the dataset page: https://huggingface.co/datasets/sentence-transformers/msmarco-distilbert-margin-mse-mnrl-mean-v1.msmarco-psgs-distilbert-dot-v5news-ka-shuffled-DISTILBERTmsmarco-distilbert-margin-mse-cls-dot-v2
MS MARCO with hard negatives from distilbert-margin-mse-cls-dot-v2
MS MARCO is a large scale information retrieval corpus that was created based on real user search queries using the Bing search engine.
For each query and gold positive passage, the 50 most similar paragraphs were mined using 13 different models. The resulting data can be used to train Sentence Transformer models.
Related Datasets
These are the datasets generated using the 13 different models:… See the full description on the dataset page: https://huggingface.co/datasets/sentence-transformers/msmarco-distilbert-margin-mse-cls-dot-v2.news-ka-small-DISTILBERTchonkiepedia-distilbert-tokenizeddistilbert_spelling_dataset-v2distilbert_spelling_dataset-v3sst2-distilbert-data
SST-2 (GLUE) — Raw Splits Used for Khalyie/sst2-distilbert
Same train/validation/test splits used to fine-tune
Khalyie/sst2-distilbert.
Note: the test split's label column is -1 for every row — GLUE
withholds official test labels. Use validation for labeled evaluation.
Reload with:
from datasets import load_dataset
ds = load_dataset("Khalyie/sst2-distilbert-data")
hyperpartisan-news-distilbert-tokenscovid-tweet-sentiment-analyzer-distilbert-data
Dataset Card for "covid-tweet-sentiment-analyzer-distilbert-data"
More Information needed
distilbert_spelling_datasetretrieval_verification_distilbert
Dataset Card for "retrieval_verification_distilbert"
More Information needed
sentiment-analysis-distilbert-rotten-tomatoes
Sentiment Analysis NLP Pipeline (Rotten Tomatoes)
This project builds a complete NLP data processing pipeline for sentiment analysis using the Rotten Tomatoes dataset.
The workflow focuses on:
Data exploration
Data cleaning
Feature engineering
Tokenization and model preparation
No model training has been performed yet. This project prepares the dataset for training transformer-based models such as DistilBERT.
Dataset
We use the Rotten Tomatoes dataset from… See the full description on the dataset page: https://huggingface.co/datasets/AmjaadXX/sentiment-analysis-distilbert-rotten-tomatoes.Tweet_sentiment_analysis_Distilbert
Dataset Card for "Tweet_sentiment_analysis_Distilbert"
More Information needed
retrieval_verification_bm25_distilbert
Dataset Card for "retrieval_verification_bm25_distilbert"
More Information needed
covid-tweet-sentiment-analyzer-distilbert-dataembedding_dataset_distilbert_base_uncased_ad_subwords
Dataset Card for "embedding_dataset_distilbert_base_uncased_ad_subwords"
More Information needed
nyaya-ae-msmarco-distilbert-base-tas-b
Dataset Card for "nyaya-ae-msmarco-distilbert-base-tas-b"
More Information needed
fintech-sentiment-distilbert-balanced-v2
Dataset Card for monostate/fintech-sentiment-distilbert-balanced-v2
Dataset Description
This dataset was generated using Vibe Data Director, a tool for creating and curating text classification datasets.
Dataset Summary
Session ID: session_f15db25a
Generated: 2026-03-16T13:09:50.740014
Total Samples: 481
Classes: negative, positive, neutral
Styles: none
Dataset Structure
Data Fields
text (string): The text content of the sample
class… See the full description on the dataset page: https://huggingface.co/datasets/monostate/fintech-sentiment-distilbert-balanced-v2.autotrain-data-DistilBert-500-500
Dataset Card for "autotrain-data-DistilBert-500-500"
More Information needed
fintech-sentiment-distilbert-ready
Dataset Card for monostate/fintech-sentiment-distilbert-ready
Dataset Description
This dataset was generated using Vibe Data Director, a tool for creating and curating text classification datasets.
Dataset Summary
Session ID: session_7987dfd2
Generated: 2026-03-16T12:59:16.439137
Total Samples: 406
Classes: positive, negative, neutral
Styles: none
Dataset Structure
Data Fields
text (string): The text content of the sample
class… See the full description on the dataset page: https://huggingface.co/datasets/monostate/fintech-sentiment-distilbert-ready.amazon-reviews-tokenized-distilbert-balanced-3labelsdistilbert-base-uncased-pii-200_dataset
