datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
msmarco-distilbert-margin-mse-mean-dot-v1
MS MARCO with hard negatives from distilbert-margin-mse-mean-dot-v1
MS MARCO is a large scale information retrieval corpus that was created based on real user search queries using the Bing search engine.
For each query and gold positive passage, the 50 most similar paragraphs were mined using 13 different models. The resulting data can be used to train Sentence Transformer models.
Related Datasets
These are the datasets generated using the 13 different models:… See the full description on the dataset page: https://huggingface.co/datasets/sentence-transformers/msmarco-distilbert-margin-mse-mean-dot-v1.msmarco-msmarco-distilbert-base-v3
MS MARCO with hard negatives from msmarco-distilbert-base-v3
MS MARCO is a large scale information retrieval corpus that was created based on real user search queries using the Bing search engine.
For each query and gold positive passage, the 50 most similar paragraphs were mined using 13 different models. The resulting data can be used to train Sentence Transformer models.
Related Datasets
These are the datasets generated using the 13 different models:
msmarco-bm25… See the full description on the dataset page: https://huggingface.co/datasets/sentence-transformers/msmarco-msmarco-distilbert-base-v3.msmarco-distilbert-margin-mse-sym-mnrl-mean-v2
MS MARCO with hard negatives from distilbert-margin-mse-sym-mnrl-mean-v2
MS MARCO is a large scale information retrieval corpus that was created based on real user search queries using the Bing search engine.
For each query and gold positive passage, the 50 most similar paragraphs were mined using 13 different models. The resulting data can be used to train Sentence Transformer models.
Related Datasets
These are the datasets generated using the 13 different models:… See the full description on the dataset page: https://huggingface.co/datasets/sentence-transformers/msmarco-distilbert-margin-mse-sym-mnrl-mean-v2.msmarco-distilbert-margin-mse-cls-dot-v1
MS MARCO with hard negatives from distilbert-margin-mse-cls-dot-v1
MS MARCO is a large scale information retrieval corpus that was created based on real user search queries using the Bing search engine.
For each query and gold positive passage, the 50 most similar paragraphs were mined using 13 different models. The resulting data can be used to train Sentence Transformer models.
Related Datasets
These are the datasets generated using the 13 different models:… See the full description on the dataset page: https://huggingface.co/datasets/sentence-transformers/msmarco-distilbert-margin-mse-cls-dot-v1.msmarco-msmarco-distilbert-base-tas-b
MS MARCO with hard negatives from msmarco-distilbert-base-tas-b
MS MARCO is a large scale information retrieval corpus that was created based on real user search queries using the Bing search engine.
For each query and gold positive passage, the 50 most similar paragraphs were mined using 13 different models. The resulting data can be used to train Sentence Transformer models.
Related Datasets
These are the datasets generated using the 13 different models:… See the full description on the dataset page: https://huggingface.co/datasets/sentence-transformers/msmarco-msmarco-distilbert-base-tas-b.msmarco-distilbert-margin-mse-sym-mnrl-mean-v1
MS MARCO with hard negatives from distilbert-margin-mse-sym-mnrl-mean-v1
MS MARCO is a large scale information retrieval corpus that was created based on real user search queries using the Bing search engine.
For each query and gold positive passage, the 50 most similar paragraphs were mined using 13 different models. The resulting data can be used to train Sentence Transformer models.
Related Datasets
These are the datasets generated using the 13 different models:… See the full description on the dataset page: https://huggingface.co/datasets/sentence-transformers/msmarco-distilbert-margin-mse-sym-mnrl-mean-v1.msmarco-distilbert-margin-mse-mnrl-mean-v1
MS MARCO with hard negatives from distilbert-margin-mse-mnrl-mean-v1
MS MARCO is a large scale information retrieval corpus that was created based on real user search queries using the Bing search engine.
For each query and gold positive passage, the 50 most similar paragraphs were mined using 13 different models. The resulting data can be used to train Sentence Transformer models.
Related Datasets
These are the datasets generated using the 13 different models:… See the full description on the dataset page: https://huggingface.co/datasets/sentence-transformers/msmarco-distilbert-margin-mse-mnrl-mean-v1.msmarco-distilbert-margin-mse-cls-dot-v2
MS MARCO with hard negatives from distilbert-margin-mse-cls-dot-v2
MS MARCO is a large scale information retrieval corpus that was created based on real user search queries using the Bing search engine.
For each query and gold positive passage, the 50 most similar paragraphs were mined using 13 different models. The resulting data can be used to train Sentence Transformer models.
Related Datasets
These are the datasets generated using the 13 different models:… See the full description on the dataset page: https://huggingface.co/datasets/sentence-transformers/msmarco-distilbert-margin-mse-cls-dot-v2.distilbert__distilgpt2-details
Dataset Card for Evaluation run of distilbert/distilgpt2
Dataset automatically created during the evaluation run of model distilbert/distilgpt2
The dataset is composed of 44 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An additional… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/distilbert__distilgpt2-details.retrieval_verification_bm25_distilbert
Dataset Card for "retrieval_verification_bm25_distilbert"
More Information needed
retrieval_verification_distilbert
Dataset Card for "retrieval_verification_distilbert"
More Information needed
nyaya-ae-msmarco-distilbert-base-tas-b
Dataset Card for "nyaya-ae-msmarco-distilbert-base-tas-b"
More Information needed
embedding_dataset_distilbert_base_uncased_ad_subwords
Dataset Card for "embedding_dataset_distilbert_base_uncased_ad_subwords"
More Information needed
embeddings_from_distilbert_class_heaps_and_eval_part0
Dataset Card for "embeddings_from_distilbert_class_heaps_and_eval_part0"
More Information needed
embeddings_from_distilbert_class_heaps
Dataset Card for "embeddings_from_distilbert_class_heaps"
Dataset created for thesis: "Generating Robust Representations of Structures in OpenSSH Heap Dumps" by Johannes Garstenauer.
This dataset contains representations of heap data structures along with their labels and the predicted label.
The representations are the [CLS] token embeddings of the last 3 layers of the DistilBERT model.
The representation-generating model is:… See the full description on the dataset page: https://huggingface.co/datasets/johannes-garstenauer/embeddings_from_distilbert_class_heaps.embeddings_from_distilbert_masking_heaps_and_eval_part0
Dataset Card for "embeddings_from_distilbert_masking_heaps_and_eval_part0"
More Information needed
embeddings_from_distilbert_masking_heaps
Dataset Card for "embeddings_from_distilbert_masking_heaps"
Dataset created for thesis: "Generating Robust Representations of Structures in OpenSSH Heap Dumps" by Johannes Garstenauer.
This dataset contains representations of heap data structures along with their labels and the predicted label.
The representations are the [CLS] token embeddings of the last 3 layers of the DistilBERT model.
The representation-generating model is:… See the full description on the dataset page: https://huggingface.co/datasets/johannes-garstenauer/embeddings_from_distilbert_masking_heaps.companyx_customer_support_ticket_routing_distilbert_dataset
CompanyX Customer Support Ticket Routing
Description: Automatically route customer support tickets to relevant teams based on issue descriptions, speeding up resolution time and enhancing customer experience.
How to Use
Here is how to use this model to classify text into different categories:
from transformers import AutoModelForSequenceClassification, AutoTokenizer
model_name = "interneuronai/companyx_customer_support_ticket_routing_distilbert"
model =… See the full description on the dataset page: https://huggingface.co/datasets/interneuronai/companyx_customer_support_ticket_routing_distilbert_dataset.distilbert-learning-feedbackpaper_test_assym_distilbert_results
Dataset Card for "paper_test_assym_distilbert_results"
More Information needed
nyaya-ae-msmarco-distilbert-base-tas-b-v1
Dataset Card for "nyaya-ae-msmarco-distilbert-base-tas-b-v1"
More Information needed
embeddings_from_distilbert_class_heaps_and_eval_part0_test
Dataset Card for "embeddings_from_distilbert_class_heaps_and_eval_part0_test"
More Information needed
Bert-distilbertembeddings_from_distilbert_class_heaps_and_eval1perc
Dataset Card for "embeddings_from_distilbert_class_heaps_and_eval1perc"
More Information needed
embeddings_from_distilbert_class_heaps_and_eval_part1
Dataset Card for "embeddings_from_distilbert_class_heaps_and_eval_part1"
More Information needed
embeddings_from_distilbert_masking_heaps_and_eval_part1
Dataset Card for "embeddings_from_distilbert_masking_heaps_and_eval_part1"
More Information needed
embeddings_from_distilbert_class_heaps_and_eval1perc_2
Dataset Card for "embeddings_from_distilbert_class_heaps_and_eval1perc_2"
More Information needed
embeddings_from_distilbert_class_heaps_and_eval_part1_test
Dataset Card for "embeddings_from_distilbert_class_heaps_and_eval_part1_test"
More Information needed
embeddings_from_distilbert_masking_heaps_and_eval_part0_test
Dataset Card for "embeddings_from_distilbert_masking_heaps_and_eval_part0_test"
More Information needed
embeddings_from_distilbert_masking_heaps_and_eval_part1_test
Dataset Card for "embeddings_from_distilbert_masking_heaps_and_eval_part1_test"
More Information needed
