CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Pravallika6 /specter2-corpus-paperstext1M<n<10M0 likes214 downloads5mo agoHugging Face02andersonbcdefg /SPECTER2-data Dataset Card for "SPECTER2-data" More Information needed text100K<n<1M2 likes112 downloads3y agoHugging Face03embedding-data /SPECTER Dataset Card for "SPECTER" Dataset Summary Dataset containing triplets (three sentences): anchor, positive, and negative. Contains titles of papers. Disclaimer: The team releasing SPECTER did not upload the dataset to the Hub and did not write a dataset card. These steps were done by the Hugging Face team. Dataset Structure Each example in the dataset contains triplets of equivalent sentences and is formatted as a dictionary with the key "set" and a list with… See the full description on the dataset page: https://huggingface.co/datasets/embedding-data/SPECTER.textsentence-similarity100K<n<1M3 likes100 downloads4y agoHugging Face04sentence-transformers /specter Dataset Card for Specter This dataset is a collection of title-related-unrelated triplets from Scientific Publications on Specter. See Specter for additional information. This dataset can be used directly with Sentence Transformers to train embedding models. Dataset Subsets triplet subset Columns: "anchor", "positive", "negative" Column types: str, str, str Examples:{ 'anchor': "Integrating children's contributions in the interaction design process"… See the full description on the dataset page: https://huggingface.co/datasets/sentence-transformers/specter.textfeature-extraction1M<n<10M2 likes93 downloads2y agoHugging Face05CyberHarem /specter_arknights Dataset of specter/スペクター/幽灵鲨 (Arknights) This is the dataset of specter/スペクター/幽灵鲨 (Arknights), containing 500 images and their tags. The core tags of this character are long_hair, red_eyes, grey_hair, hair_between_eyes, black_headwear, very_long_hair, breasts, white_hair, hat, which are pruned in this dataset. Images are crawled from many sites (e.g. danbooru, pixiv, zerochan ...), the auto-crawling system is powered by DeepGHS Team(huggingface organization). List of… See the full description on the dataset page: https://huggingface.co/datasets/CyberHarem/specter_arknights.text-to-image1K<n<10K0 likes82 downloads3y agoHugging Face06Jerjes /neuro-specter2-triplets-multi-pool Jerjes/neuro-specter2-triplets-multi-pool This dataset contains anchor papers with their top-K most similar (positive) and most dissimilar (negative) papers based on SPECTER2 embeddings. Dataset Structure Each row contains: anchor_id: Unique identifier for the anchor paper anchor_title: Title of the anchor paper anchor_abstract: Abstract of the anchor paper positive_pool: List of 5 most similar papers, each as [id, title, abstract] negative_pool: List of 5 most… See the full description on the dataset page: https://huggingface.co/datasets/Jerjes/neuro-specter2-triplets-multi-pool.text100K<n<1M0 likes61 downloads1y agoHugging Face07AlgorithmicResearchGroup /arxiv_abstracts_specter_faiss_flat_indexversion https://git-lfs.github.com/spec/v1 oid sha256:98b45ea81164d1e1a1dd82255207053b15cd6c69d922a1c5cf3387ce604d4b74 size 28 text1M<n<10M0 likes47 downloads4y agoHugging Face08NothingMuch /Specter-Triplet-SplitThis dataset is based on the sentence-transformers/specter dataset to have a train-val-test split. text100K<n<1M0 likes41 downloads2y agoHugging Face09Jerjes /neuro-specter2-triplets Jerjes/neuro-specter2-triplets Triplet dataset for fine-tuning SPECTER2 on neuroscience. Version date: 2025-08-12 Schema Columns: anchor_id, positive_id, negative_id anchor_title, positive_title, negative_title anchor_abstract, positive_abstract, negative_abstract anchor_text, positive_text, negative_text (title + abstract) Split: train Load from datasets import load_dataset triplets = load_dataset("Jerjes/neuro-specter2-triplets", split="train") texttext-ranking10K<n<100K0 likes31 downloads1y agoHugging Face10andersonbcdefg /st_specter_train_triplestext100K<n<1M0 likes22 downloads3y agoHugging Face11andersonbcdefg /specter-title-to-abstext100K<n<1M0 likes16 downloads3y agoHugging Face12Singhchandann /specter_marathigated Specter Marathi Dataset: High-Quality Marathi NLP Corpus 📌 Overview The Specter Marathi dataset is a meticulously curated collection of 684098 rows of Marathi text, ensuring linguistic accuracy and natural flow. Every sentence has been verified by native Marathi speakers to maintain contextual integrity and correctness. This dataset is designed for semantic search, text classification, and various NLP tasks, making it a valuable resource for machine learning models… See the full description on the dataset page: https://huggingface.co/datasets/Singhchandann/specter_marathi.text-classification100K<n<1M0 likes15 downloads1y agoHugging Face13andersonbcdefg /specter-abs-to-titletext100K<n<1M0 likes14 downloads3y agoHugging Face14frankterpo /specter-vc-ai_specialisttextn<1K0 likes14 downloads10mo agoHugging Face15Jerjes /neuro-specter2-triplets-pool Jerjes/neuro-specter2-triplets-pool This dataset contains anchor papers with their top-K most similar (positive) and most dissimilar (negative) papers based on SPECTER2 embeddings. Dataset Structure Each row contains: anchor_id: Unique identifier for the anchor paper anchor_title: Title of the anchor paper anchor_abstract: Abstract of the anchor paper positive_pool: List of 5 most similar papers, each as [id, title, abstract] negative_pool: List of 5 most dissimilar… See the full description on the dataset page: https://huggingface.co/datasets/Jerjes/neuro-specter2-triplets-pool.text100K<n<1M0 likes11 downloads1y agoHugging Face16Jerjes /neuro-specter2-poolstext100K<n<1M0 likes10 downloads1y agoHugging Face17frankterpo /specter-vc-stealth_huntertextn<1K0 likes10 downloads10mo agoHugging Face18frankterpo /specter-vc-all-personastextn<1K0 likes10 downloads10mo agoHugging Face19Jerjes /neuro-specter2-sample-data Jerjes/neuro-specter2-sample-data This dataset contains anchor papers with their top-K most similar (positive) and most dissimilar (negative) papers based on SPECTER2 embeddings. Dataset Structure Each row contains: anchor_id: Unique identifier for the anchor paper anchor_title: Title of the anchor paper anchor_abstract: Abstract of the anchor paper positive_pool: List of 5 most similar papers, each as [id, title, abstract] negative_pool: List of 5 most dissimilar… See the full description on the dataset page: https://huggingface.co/datasets/Jerjes/neuro-specter2-sample-data.textn<1K0 likes9 downloads1y agoHugging Face20andersonbcdefg /SPECTER-subset-deduptext100K<n<1M0 likes7 downloads3y agoHugging Face21ragrawal36 /etd-specter-train-triples-hard-neg-sfttext100K<n<1M0 likes7 downloads3mo agoHugging Face22andersonbcdefg /specter-title-to-abs-filteredtext100K<n<1M1 likes6 downloads3y agoHugging Face23andersonbcdefg /SPECTER-subset-dedup_with_marginstabular10K<n<100K0 likes5 downloads3y agoHugging Face24kumarme072 /specter_77k0 likes5 downloads2y agoHugging Face25frankterpo /specter-vc-growth_scouttextn<1K0 likes5 downloads10mo agoHugging Face26frankterpo /specter-vc-fintech_focustextn<1K0 likes4 downloads10mo agoHugging Face27mitanshu-reckonsys /tldr_vs_abstract_allenai_specter2_aug2023refresh_basetext1K<n<10K0 likes4 downloads5mo agoHugging Face28mitanshu-reckonsys /tldr_vs_abstract_allenai_specter2_basetext1K<n<10K0 likes3 downloads5mo agoHugging Face29frankterpo /specter-vc-early_stage0 likes2 downloads10mo agoHugging Face30jadenhochh /specter10 likes2 downloads8mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.