datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
arxiv-titles-instructorxl-embeddings
arxiv-titles-instructorxl-embeddings
This dataset contains 768-dimensional embeddings generated from the arxiv
paper titles using InstructorXL model. Each
vector has an abstract used to create it, along with the DOI (Digital Object Identifier). The
dataset was created using precomputed embeddings exposed by the Alexandria Index.
Generation process
The embeddings have been generated using the following instruction:
Represent the Research Paper title for retrieval;… See the full description on the dataset page: https://huggingface.co/datasets/Qdrant/arxiv-titles-instructorxl-embeddings.meddies-title
Meddies Title SFT
title_sft_v0 trains multilingual session-title generation. Each row stores a system instruction, the user query, and one title response in messages.
The raw query pools remain in this public canonical repository for provenance, but are omitted from dataset viewer configurations.
SFT generation progress
Status: partial
Mode: full
Generated accepted rows: 52901 / 100000 primary + 23941 auxiliary
Published rows: 52901 primary + 23941 auxiliary… See the full description on the dataset page: https://huggingface.co/datasets/HoangHa/meddies-title.realnewslike_with_title
Dataset Card for "realnewslike_with_title"
More Information needed
wikipedia-title-text-pairsemb-reddit-title-body
Dataset Card for "emb-reddit-title-body"
More Information needed
job-titles
Comprehensive Job Titles Dataset
A high-quality, deduplicated dataset of 65,248 unique job titles compiled from authoritative sources including ESCO (European Skills, Competences, Qualifications and Occupations), O*NET (Occupational Information Network), and OSCA (Occupational Skills and Competencies Australia).
Dataset Description
This dataset provides a comprehensive collection of job titles that have been carefully processed to remove duplicates and near-duplicates… See the full description on the dataset page: https://huggingface.co/datasets/gpriday/job-titles.reddit-title-body-hf
reddit-title-body-hf
sentence-transformers/reddit-title-body in parquet format
additional configs
the deduped config, which has the body col deduped via minhash
the mini config, which is a ~1 GB version of the deduped dataset created via a minipile-like clustering+sampling approach
amz-product-title-emb
Embedding on Product Title
This was using gemini embedding text-embedding-005 to embed this product title dataset. Mostly using batch job.
Each embedding vector has 768 fp16 floats. (usually discount original fp32 to fp16 for gemini embedding model can put little damage on accuracy)
You can stream this dataset, so having fast boost start on training script.
soda-vec-data-full_pmc_title_abstract
SODA-VEC Clean Dataset
This is a cleaned and filtered version of the SODA-VEC dataset, containing high-quality biomedical title-abstract pairs from PubMed Central (PMC) articles.
Dataset Overview
Total examples: 26,573,900
Training set: 26,473,900 examples (99.6%)
Validation set: 50,000 examples (0.2%)
Test set: 50,000 examples (0.2%)
Quality Filtering Applied
This dataset has been processed with the following quality filters:
Abstract Length… See the full description on the dataset page: https://huggingface.co/datasets/EMBO/soda-vec-data-full_pmc_title_abstract.bookroom-persian-book-covers-and-titlesopenalex-topic-title-abstractresearch_titles_multi-labelopenalex-topic-title-abstractdeepstock-stock-historical-prices-dataset-processed-with-previous-day-titleswikidata-title-desc
Wikidata Title and Description
Wikidata entity titles and short descriptions for all 324 Wikipedia language editions,
extracted from the Wikimedia Analytics wmf.wikidata_entity
table. Each row links a Wikidata QID to the title and description as they
appear in a specific language edition of Wikipedia.
Dataset Details
Field
Value
Source
wmf.wikidata_entity (Wikidata + Wikipedia sitelinks)
Snapshot
2026-03-30
Languages
324
Total titles
88,292,409
Titles… See the full description on the dataset page: https://huggingface.co/datasets/wikimedia/wikidata-title-desc.book_titles_and_descriptionsdblp-id-titlehttps://drops.dagstuhl.de/storage/artifacts/dblp/xml/2026/dblp-2026-02-01.xml.gz
2026-02-24
soda-vec-data-full_pmc_title_abstract_paired
SODA-VEC Paired Dataset for Negative Sampling
This is a paired version of the SODA-VEC dataset, specifically formatted for negative sampling training with MultipleNegativesRankingLoss.
Dataset Overview
Total examples: 26,573,900
Format: Paired (anchor-positive) for contrastive learning
Source: EMBO/soda-vec-data-full_pmc_title_abstract
Purpose: Training sentence transformers with negative sampling
Data Format
Each example contains:
anchor (string): The title… See the full description on the dataset page: https://huggingface.co/datasets/EMBO/soda-vec-data-full_pmc_title_abstract_paired.S2ORC_title_abstract_1107
Dataset Card for "S2ORC_title_abstract_1107"
More Information needed
book_titles_and_descriptions_en_cleanpaper_title_abs_pairsamz-product-title
Amazon Product Titles
This dataset is a ablated version of Amazon Review Data (2018) from UCSD.
The core of the logic is to only have title or feature for each item. Discarding all other fields.
It supports for purposes like streaming amazon product titles to NLP training easily. It's around 15mn rows of data.
Example Row
Here's an example row, notice, feature can be empty list more frequently than we expected.
{'asin': 'B00K8018L2',
'title': 'Spoon Oil Grease Wire… See the full description on the dataset page: https://huggingface.co/datasets/raynardj/amz-product-title.task220_rocstories_title_classification
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task220_rocstories_title_classification
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task220_rocstories_title_classification.hf_multimodal_newsxlm_pp10_title_date_1_processed_labeled_chunksBananaMind-Chat-Title-200K
BananaMind Chat Title 200K
BananaMind Chat Title 200K is an title generation dataset. It contains generated chat titles for examples from the first 200,000 rows of lmsys/lmsys-chat-1m, but it does not include raw LMSYS prompt text.
That makes the biggest publicly available chat title dataset on huggingface!
Instead, each row stores an LMSYS row reference and the generated title. Users with access to LMSYS-Chat-1M can reconstruct the original first user message… See the full description on the dataset page: https://huggingface.co/datasets/BananaMind/BananaMind-Chat-Title-200K.job-titles-descriptions
Synthetic Job Descriptions Dataset
A high-quality synthetic expansion of gpriday/job-titles — 65,248 structured job descriptions with a pre-computed FAISS index for semantic search.
Overview
This dataset provides comprehensive, synthetically generated job descriptions for all role titles found in the gpriday/job-titles source dataset. Each entry is structured for consistency and designed to support NLP pipelines, career-tech applications, and Retrieval-Augmented… See the full description on the dataset page: https://huggingface.co/datasets/ismailcemsahin/job-titles-descriptions.task1342_amazon_us_reviews_title
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1342_amazon_us_reviews_title
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1342_amazon_us_reviews_title.ecfr-title-12meddies-title-benchmark-v0
Meddies Title Benchmark v0
Private frozen evaluation subset for Meddies Title v0: 3,400 group-isolated queries, exactly 200 per language across 17 languages. This repository contains queries and provenance only, without reference titles.
The immutable selection and source revision are recorded in manifest.json.
job_titles
Dataset Card for "job_titles"
More Information needed
Normalized dataset of 70k job titles
