CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01sentence-transformers /msmarco-distilbert-margin-mse-mean-dot-v1 MS MARCO with hard negatives from distilbert-margin-mse-mean-dot-v1 MS MARCO is a large scale information retrieval corpus that was created based on real user search queries using the Bing search engine. For each query and gold positive passage, the 50 most similar paragraphs were mined using 13 different models. The resulting data can be used to train Sentence Transformer models. Related Datasets These are the datasets generated using the 13 different models:… See the full description on the dataset page: https://huggingface.co/datasets/sentence-transformers/msmarco-distilbert-margin-mse-mean-dot-v1.tabularfeature-extraction10M<n<100M2 likes3.4k downloads2y agoHugging Face02sentence-transformers /msmarco-msmarco-distilbert-base-v3 MS MARCO with hard negatives from msmarco-distilbert-base-v3 MS MARCO is a large scale information retrieval corpus that was created based on real user search queries using the Bing search engine. For each query and gold positive passage, the 50 most similar paragraphs were mined using 13 different models. The resulting data can be used to train Sentence Transformer models. Related Datasets These are the datasets generated using the 13 different models: msmarco-bm25… See the full description on the dataset page: https://huggingface.co/datasets/sentence-transformers/msmarco-msmarco-distilbert-base-v3.tabularfeature-extraction10M<n<100M5 likes2.4k downloads2y agoHugging Face03sentence-transformers /msmarco-distilbert-margin-mse-sym-mnrl-mean-v2 MS MARCO with hard negatives from distilbert-margin-mse-sym-mnrl-mean-v2 MS MARCO is a large scale information retrieval corpus that was created based on real user search queries using the Bing search engine. For each query and gold positive passage, the 50 most similar paragraphs were mined using 13 different models. The resulting data can be used to train Sentence Transformer models. Related Datasets These are the datasets generated using the 13 different models:… See the full description on the dataset page: https://huggingface.co/datasets/sentence-transformers/msmarco-distilbert-margin-mse-sym-mnrl-mean-v2.tabularfeature-extraction10M<n<100M1 likes2.3k downloads2y agoHugging Face04sentence-transformers /msmarco-msmarco-distilbert-base-tas-b MS MARCO with hard negatives from msmarco-distilbert-base-tas-b MS MARCO is a large scale information retrieval corpus that was created based on real user search queries using the Bing search engine. For each query and gold positive passage, the 50 most similar paragraphs were mined using 13 different models. The resulting data can be used to train Sentence Transformer models. Related Datasets These are the datasets generated using the 13 different models:… See the full description on the dataset page: https://huggingface.co/datasets/sentence-transformers/msmarco-msmarco-distilbert-base-tas-b.tabularfeature-extraction10M<n<100M5 likes1.8k downloads2y agoHugging Face05sentence-transformers /msmarco-distilbert-margin-mse-cls-dot-v1 MS MARCO with hard negatives from distilbert-margin-mse-cls-dot-v1 MS MARCO is a large scale information retrieval corpus that was created based on real user search queries using the Bing search engine. For each query and gold positive passage, the 50 most similar paragraphs were mined using 13 different models. The resulting data can be used to train Sentence Transformer models. Related Datasets These are the datasets generated using the 13 different models:… See the full description on the dataset page: https://huggingface.co/datasets/sentence-transformers/msmarco-distilbert-margin-mse-cls-dot-v1.tabularfeature-extraction10M<n<100M0 likes1.5k downloads2y agoHugging Face06sentence-transformers /msmarco-distilbert-margin-mse-sym-mnrl-mean-v1 MS MARCO with hard negatives from distilbert-margin-mse-sym-mnrl-mean-v1 MS MARCO is a large scale information retrieval corpus that was created based on real user search queries using the Bing search engine. For each query and gold positive passage, the 50 most similar paragraphs were mined using 13 different models. The resulting data can be used to train Sentence Transformer models. Related Datasets These are the datasets generated using the 13 different models:… See the full description on the dataset page: https://huggingface.co/datasets/sentence-transformers/msmarco-distilbert-margin-mse-sym-mnrl-mean-v1.tabularfeature-extraction10M<n<100M1 likes1.5k downloads2y agoHugging Face07sentence-transformers /msmarco-distilbert-margin-mse-mnrl-mean-v1 MS MARCO with hard negatives from distilbert-margin-mse-mnrl-mean-v1 MS MARCO is a large scale information retrieval corpus that was created based on real user search queries using the Bing search engine. For each query and gold positive passage, the 50 most similar paragraphs were mined using 13 different models. The resulting data can be used to train Sentence Transformer models. Related Datasets These are the datasets generated using the 13 different models:… See the full description on the dataset page: https://huggingface.co/datasets/sentence-transformers/msmarco-distilbert-margin-mse-mnrl-mean-v1.tabularfeature-extraction10M<n<100M0 likes857 downloads2y agoHugging Face08sentence-transformers /msmarco-distilbert-margin-mse-cls-dot-v2 MS MARCO with hard negatives from distilbert-margin-mse-cls-dot-v2 MS MARCO is a large scale information retrieval corpus that was created based on real user search queries using the Bing search engine. For each query and gold positive passage, the 50 most similar paragraphs were mined using 13 different models. The resulting data can be used to train Sentence Transformer models. Related Datasets These are the datasets generated using the 13 different models:… See the full description on the dataset page: https://huggingface.co/datasets/sentence-transformers/msmarco-distilbert-margin-mse-cls-dot-v2.tabularfeature-extraction10M<n<100M2 likes237 downloads2y agoHugging Face09open-llm-leaderboard /distilbert__distilgpt2-detailsgated Dataset Card for Evaluation run of distilbert/distilgpt2 Dataset automatically created during the evaluation run of model distilbert/distilgpt2 The dataset is composed of 44 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results. An additional… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/distilbert__distilgpt2-details.tabular10K<n<100K1 likes67 downloads2y agoHugging Face10nikchar /retrieval_verification_distilbert Dataset Card for "retrieval_verification_distilbert" More Information needed tabular10K<n<100K0 likes22 downloads3y agoHugging Face11nikchar /retrieval_verification_bm25_distilbert Dataset Card for "retrieval_verification_bm25_distilbert" More Information needed tabular10K<n<100K0 likes19 downloads3y agoHugging Face12sanjin7 /embedding_dataset_distilbert_base_uncased_ad_subwords Dataset Card for "embedding_dataset_distilbert_base_uncased_ad_subwords" More Information needed tabular1K<n<10K0 likes15 downloads4y agoHugging Face13sukantan /nyaya-ae-msmarco-distilbert-base-tas-b Dataset Card for "nyaya-ae-msmarco-distilbert-base-tas-b" More Information needed tabular10K<n<100K0 likes15 downloads3y agoHugging Face14johannes-garstenauer /embeddings_from_distilbert_class_heaps_and_eval_part0 Dataset Card for "embeddings_from_distilbert_class_heaps_and_eval_part0" More Information needed tabular100K<n<1M0 likes11 downloads3y agoHugging Face15johannes-garstenauer /embeddings_from_distilbert_class_heaps Dataset Card for "embeddings_from_distilbert_class_heaps" Dataset created for thesis: "Generating Robust Representations of Structures in OpenSSH Heap Dumps" by Johannes Garstenauer. This dataset contains representations of heap data structures along with their labels and the predicted label. The representations are the [CLS] token embeddings of the last 3 layers of the DistilBERT model. The representation-generating model is:… See the full description on the dataset page: https://huggingface.co/datasets/johannes-garstenauer/embeddings_from_distilbert_class_heaps.tabular100K<n<1M1 likes11 downloads3y agoHugging Face16eliodecolli /distilbert-learning-feedbacktabularn<1K0 likes10 downloads1mo agoHugging Face17johannes-garstenauer /embeddings_from_distilbert_masking_heaps_and_eval_part0 Dataset Card for "embeddings_from_distilbert_masking_heaps_and_eval_part0" More Information needed tabular100K<n<1M0 likes9 downloads3y agoHugging Face18johannes-garstenauer /embeddings_from_distilbert_masking_heaps Dataset Card for "embeddings_from_distilbert_masking_heaps" Dataset created for thesis: "Generating Robust Representations of Structures in OpenSSH Heap Dumps" by Johannes Garstenauer. This dataset contains representations of heap data structures along with their labels and the predicted label. The representations are the [CLS] token embeddings of the last 3 layers of the DistilBERT model. The representation-generating model is:… See the full description on the dataset page: https://huggingface.co/datasets/johannes-garstenauer/embeddings_from_distilbert_masking_heaps.tabular100K<n<1M1 likes9 downloads3y agoHugging Face19nikchar /paper_test_assym_distilbert_results Dataset Card for "paper_test_assym_distilbert_results" More Information needed tabular10K<n<100K0 likes7 downloads3y agoHugging Face20johannes-garstenauer /embeddings_from_distilbert_class_heaps_and_eval_part0_test Dataset Card for "embeddings_from_distilbert_class_heaps_and_eval_part0_test" More Information needed tabular1K<n<10K0 likes7 downloads3y agoHugging Face21interneuronai /companyx_customer_support_ticket_routing_distilbert_dataset CompanyX Customer Support Ticket Routing Description: Automatically route customer support tickets to relevant teams based on issue descriptions, speeding up resolution time and enhancing customer experience. How to Use Here is how to use this model to classify text into different categories: from transformers import AutoModelForSequenceClassification, AutoTokenizer model_name = "interneuronai/companyx_customer_support_ticket_routing_distilbert" model =… See the full description on the dataset page: https://huggingface.co/datasets/interneuronai/companyx_customer_support_ticket_routing_distilbert_dataset.tabular1K<n<10K0 likes7 downloads2y agoHugging Face22SAGAY /Bert-distilberttabular1K<n<10K0 likes6 downloads4y agoHugging Face23sukantan /nyaya-ae-msmarco-distilbert-base-tas-b-v1 Dataset Card for "nyaya-ae-msmarco-distilbert-base-tas-b-v1" More Information needed tabular10K<n<100K0 likes6 downloads3y agoHugging Face24johannes-garstenauer /embeddings_from_distilbert_class_heaps_and_eval1perc Dataset Card for "embeddings_from_distilbert_class_heaps_and_eval1perc" More Information needed tabular1K<n<10K0 likes6 downloads3y agoHugging Face25johannes-garstenauer /embeddings_from_distilbert_class_heaps_and_eval_part1 Dataset Card for "embeddings_from_distilbert_class_heaps_and_eval_part1" More Information needed tabular100K<n<1M0 likes6 downloads3y agoHugging Face26johannes-garstenauer /embeddings_from_distilbert_masking_heaps_and_eval_part1 Dataset Card for "embeddings_from_distilbert_masking_heaps_and_eval_part1" More Information needed tabular100K<n<1M0 likes6 downloads3y agoHugging Face27johannes-garstenauer /embeddings_from_distilbert_class_heaps_and_eval1perc_2 Dataset Card for "embeddings_from_distilbert_class_heaps_and_eval1perc_2" More Information needed tabular1K<n<10K0 likes5 downloads3y agoHugging Face28johannes-garstenauer /embeddings_from_distilbert_class_heaps_and_eval_part1_test Dataset Card for "embeddings_from_distilbert_class_heaps_and_eval_part1_test" More Information needed tabular1K<n<10K0 likes5 downloads3y agoHugging Face29johannes-garstenauer /embeddings_from_distilbert_masking_heaps_and_eval_part0_test Dataset Card for "embeddings_from_distilbert_masking_heaps_and_eval_part0_test" More Information needed tabular1K<n<10K0 likes5 downloads3y agoHugging Face30johannes-garstenauer /embeddings_from_distilbert_masking_heaps_and_eval_part1_test Dataset Card for "embeddings_from_distilbert_masking_heaps_and_eval_part1_test" More Information needed tabular1K<n<10K0 likes5 downloads3y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.