CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Qdrant /arxiv-titles-instructorxl-embeddings arxiv-titles-instructorxl-embeddings This dataset contains 768-dimensional embeddings generated from the arxiv paper titles using InstructorXL model. Each vector has an abstract used to create it, along with the DOI (Digital Object Identifier). The dataset was created using precomputed embeddings exposed by the Alexandria Index. Generation process The embeddings have been generated using the following instruction: Represent the Research Paper title for retrieval;… See the full description on the dataset page: https://huggingface.co/datasets/Qdrant/arxiv-titles-instructorxl-embeddings.textsentence-similarity1M<n<10M5 likes3.2k downloads3y agoHugging Face02HoangHa /meddies-title Meddies Title SFT title_sft_v0 trains multilingual session-title generation. Each row stores a system instruction, the user query, and one title response in messages. The raw query pools remain in this public canonical repository for provenance, but are omitted from dataset viewer configurations. SFT generation progress Status: partial Mode: full Generated accepted rows: 52901 / 100000 primary + 23941 auxiliary Published rows: 52901 primary + 23941 auxiliary… See the full description on the dataset page: https://huggingface.co/datasets/HoangHa/meddies-title.text100K<n<1M1 likes1.5k downloads2d agoHugging Face03Rossil /realnewslike_with_title Dataset Card for "realnewslike_with_title" More Information needed text10M<n<100M3 likes1.5k downloads3y agoHugging Face04sproos /wikipedia-title-text-pairstext100M<n<1B1 likes1.3k downloads3y agoHugging Face05multi-train /emb-reddit-title-body Dataset Card for "emb-reddit-title-body" More Information needed text100M<n<1B0 likes1.2k downloads3y agoHugging Face06gpriday /job-titles Comprehensive Job Titles Dataset A high-quality, deduplicated dataset of 65,248 unique job titles compiled from authoritative sources including ESCO (European Skills, Competences, Qualifications and Occupations), O*NET (Occupational Information Network), and OSCA (Occupational Skills and Competencies Australia). Dataset Description This dataset provides a comprehensive collection of job titles that have been carefully processed to remove duplicates and near-duplicates… See the full description on the dataset page: https://huggingface.co/datasets/gpriday/job-titles.texttext-classification10K<n<100K8 likes962 downloads1y agoHugging Face07BEE-spoke-data /reddit-title-body-hf reddit-title-body-hf sentence-transformers/reddit-title-body in parquet format additional configs the deduped config, which has the body col deduped via minhash the mini config, which is a ~1 GB version of the deduped dataset created via a minipile-like clustering+sampling approach texttext-generation100M<n<1B4 likes519 downloads9mo agoHugging Face08raynardj /amz-product-title-emb Embedding on Product Title This was using gemini embedding text-embedding-005 to embed this product title dataset. Mostly using batch job. Each embedding vector has 768 fp16 floats. (usually discount original fp32 to fp16 for gemini embedding model can put little damage on accuracy) You can stream this dataset, so having fast boost start on training script. text10M<n<100M0 likes513 downloads5mo agoHugging Face09EMBO /soda-vec-data-full_pmc_title_abstract SODA-VEC Clean Dataset This is a cleaned and filtered version of the SODA-VEC dataset, containing high-quality biomedical title-abstract pairs from PubMed Central (PMC) articles. Dataset Overview Total examples: 26,573,900 Training set: 26,473,900 examples (99.6%) Validation set: 50,000 examples (0.2%) Test set: 50,000 examples (0.2%) Quality Filtering Applied This dataset has been processed with the following quality filters: Abstract Length… See the full description on the dataset page: https://huggingface.co/datasets/EMBO/soda-vec-data-full_pmc_title_abstract.texttext-classification10M<n<100M0 likes418 downloads1y agoHugging Face10shenasa /bookroom-persian-book-covers-and-titlesimage10K<n<100K3 likes397 downloads1y agoHugging Face11albertmartinez /openalex-topic-title-abstracttext1M<n<10M1 likes361 downloads2y agoHugging Face12argilla /research_titles_multi-labeltext10K<n<100K0 likes267 downloads4y agoHugging Face13ITOCJ /openalex-topic-title-abstracttext1M<n<10M0 likes243 downloads6mo agoHugging Face142084Collective /deepstock-stock-historical-prices-dataset-processed-with-previous-day-titlestabular10M<n<100M0 likes211 downloads2y agoHugging Face15wikimedia /wikidata-title-desc Wikidata Title and Description Wikidata entity titles and short descriptions for all 324 Wikipedia language editions, extracted from the Wikimedia Analytics wmf.wikidata_entity table. Each row links a Wikidata QID to the title and description as they appear in a specific language edition of Wikipedia. Dataset Details Field Value Source wmf.wikidata_entity (Wikidata + Wikipedia sitelinks) Snapshot 2026-03-30 Languages 324 Total titles 88,292,409 Titles… See the full description on the dataset page: https://huggingface.co/datasets/wikimedia/wikidata-title-desc.text100M<n<1B3 likes191 downloads5mo agoHugging Face16Skelebor /book_titles_and_descriptionstext1M<n<10M3 likes182 downloads4y agoHugging Face17yusuke1997 /dblp-id-titlehttps://drops.dagstuhl.de/storage/artifacts/dblp/xml/2026/dblp-2026-02-01.xml.gz 2026-02-24 text10M<n<100M0 likes142 downloads7mo agoHugging Face18EMBO /soda-vec-data-full_pmc_title_abstract_paired SODA-VEC Paired Dataset for Negative Sampling This is a paired version of the SODA-VEC dataset, specifically formatted for negative sampling training with MultipleNegativesRankingLoss. Dataset Overview Total examples: 26,573,900 Format: Paired (anchor-positive) for contrastive learning Source: EMBO/soda-vec-data-full_pmc_title_abstract Purpose: Training sentence transformers with negative sampling Data Format Each example contains: anchor (string): The title… See the full description on the dataset page: https://huggingface.co/datasets/EMBO/soda-vec-data-full_pmc_title_abstract_paired.textsentence-similarity10M<n<100M1 likes127 downloads1y agoHugging Face19multi-train /S2ORC_title_abstract_1107 Dataset Card for "S2ORC_title_abstract_1107" More Information needed text100K<n<1M0 likes125 downloads3y agoHugging Face20Skelebor /book_titles_and_descriptions_en_cleantext1M<n<10M3 likes117 downloads4y agoHugging Face21andersonbcdefg /paper_title_abs_pairstext1M<n<10M1 likes102 downloads3y agoHugging Face22raynardj /amz-product-title Amazon Product Titles This dataset is a ablated version of Amazon Review Data (2018) from UCSD. The core of the logic is to only have title or feature for each item. Discarding all other fields. It supports for purposes like streaming amazon product titles to NLP training easily. It's around 15mn rows of data. Example Row Here's an example row, notice, feature can be empty list more frequently than we expected. {'asin': 'B00K8018L2', 'title': 'Spoon Oil Grease Wire… See the full description on the dataset page: https://huggingface.co/datasets/raynardj/amz-product-title.text10M<n<100M1 likes102 downloads5mo agoHugging Face23Lots-of-LoRAs /task220_rocstories_title_classification Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task220_rocstories_title_classification Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task220_rocstories_title_classification.texttext-generation1K<n<10K0 likes100 downloads2y agoHugging Face24pbedrin /hf_multimodal_newsxlm_pp10_title_date_1_processed_labeled_chunkstext10K<n<100K0 likes97 downloads10mo agoHugging Face25BananaMind /BananaMind-Chat-Title-200K BananaMind Chat Title 200K BananaMind Chat Title 200K is an title generation dataset. It contains generated chat titles for examples from the first 200,000 rows of lmsys/lmsys-chat-1m, but it does not include raw LMSYS prompt text. That makes the biggest publicly available chat title dataset on huggingface! Instead, each row stores an LMSYS row reference and the generated title. Users with access to LMSYS-Chat-1M can reconstruct the original first user message… See the full description on the dataset page: https://huggingface.co/datasets/BananaMind/BananaMind-Chat-Title-200K.texttext-generation100K<n<1M1 likes94 downloads3mo agoHugging Face26ismailcemsahin /job-titles-descriptions Synthetic Job Descriptions Dataset A high-quality synthetic expansion of gpriday/job-titles — 65,248 structured job descriptions with a pre-computed FAISS index for semantic search. Overview This dataset provides comprehensive, synthetically generated job descriptions for all role titles found in the gpriday/job-titles source dataset. Each entry is structured for consistency and designed to support NLP pipelines, career-tech applications, and Retrieval-Augmented… See the full description on the dataset page: https://huggingface.co/datasets/ismailcemsahin/job-titles-descriptions.text10K<n<100K2 likes90 downloads5mo agoHugging Face27Lots-of-LoRAs /task1342_amazon_us_reviews_title Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1342_amazon_us_reviews_title Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1342_amazon_us_reviews_title.texttext-generation1K<n<10K0 likes82 downloads2y agoHugging Face28wesslen /ecfr-title-12text1K<n<10K1 likes78 downloads2y agoHugging Face29HoangHa /meddies-title-benchmark-v0 Meddies Title Benchmark v0 Private frozen evaluation subset for Meddies Title v0: 3,400 group-isolated queries, exactly 200 per language across 17 languages. This repository contains queries and provenance only, without reference titles. The immutable selection and source revision are recorded in manifest.json. text1K<n<10K0 likes78 downloads19d agoHugging Face30bstds /job_titles Dataset Card for "job_titles" More Information needed Normalized dataset of 70k job titles text10K<n<100K0 likes75 downloads4y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.