datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
osm-polygon-description-tag
OSM Polygon Description Tag
OpenStreetMap polygons with a successfully extracted trimmed non-empty
description or
description:<suffix> tag, published as one GeoParquet file per regional PBF
extract. Every row retains the complete original tag map, full Polygon or
MultiPolygon geometry, WGS84 geodesic area, bounding box, and OSM provenance.
Source repository: github.com/NoeFlandre/osm-polygon-description-tag.
Explore the pipeline metrics in the Trackio dashboard.
Read the… See the full description on the dataset page: https://huggingface.co/datasets/NoeFlandre/osm-polygon-description-tag.mls-eng-speaker-descriptions
Dataset Card for Annotations of English MLS
This dataset consists in annotations of the English subset of the Multilingual LibriSpeech (MLS) dataset.
MLS dataset is a large multilingual corpus suitable for speech research. The dataset is derived from read audiobooks from LibriVox and consists of
8 languages - English, German, Dutch, Spanish, French, Italian, Portuguese, Polish. It includes about 44.5K hours of English and a total of about 6K hours for other languages.
This dataset… See the full description on the dataset page: https://huggingface.co/datasets/parler-tts/mls-eng-speaker-descriptions.circuitlens-gemma-2-2btranscoder-descriptions-and-evaluations
CircuitLens & WeightLens: Transcoder Descriptions and Evaluations
This dataset contains automatically generated descriptions and evaluation metrics for Gemma-2-2B transcoders, produced using CircuitLens and WeightLens methods.
Methods
CircuitLens: https://github.com/egolimblevskaia/CircuitLens
WeightLens: https://github.com/egolimblevskaia/WeightLens
Dataset Structure
The dataset is organized by layers (0, 4, 7, 10, 12, 15, 18, 21, 23, 25), with each layer… See the full description on the dataset page: https://huggingface.co/datasets/egolimblevskaia/circuitlens-gemma-2-2btranscoder-descriptions-and-evaluations.libritts-r-filtered-speaker-descriptions
Dataset Card for Annotated LibriTTS-R
This dataset is an annotated version of a filtered LibriTTS-R [1].
LibriTTS-R [1] is a sound quality improved version of the LibriTTS corpus which is a multi-speaker English corpus of approximately 960 hours of read English speech at 24kHz sampling rate, published in 2019.
In the text_description column, it provides natural language annotations on the characteristics of speakers and utterances, that have been generated using the Data-Speech… See the full description on the dataset page: https://huggingface.co/datasets/parler-tts/libritts-r-filtered-speaker-descriptions.vietnamese-job-descriptions
💼 Tinix Vietnam Job Description
1. 📌 Giới Thiệu Tinix Vietnam Job Description
Tinix Vietnam Job Description là bộ dữ liệu tuyển dụng tiếng Việt ở định dạng CSV, gồm các tin tuyển dụng có cấu trúc về chức danh, công ty, mức lương, địa điểm, loại hợp đồng, ngành nghề, yêu cầu kinh nghiệm, trình độ học vấn, mô tả công việc, phúc lợi, yêu cầu ứng viên và năm đăng tin.
Bộ dữ liệu được thiết kế cho các bài toán NLP và phân tích thị trường lao động tại Việt Nam, đặc biệt trong… See the full description on the dataset page: https://huggingface.co/datasets/tinixai/vietnamese-job-descriptions.capstone_sakuga_simple_description_mlm_hsamazon_product_descriptioncapstone_sakuga_simple_descriptionus-patent-descriptions
US Patent Descriptions
This dataset contains the descriptions of granted US utility patents, filtered and deduplicated.The original data comes from all granted patents in 2025 up to May 20, available from PatentsView.
Splits
train: 10,000 rows for model training
validation: 2,500 rows for validation
test: 2,500 rows for evaluation
Columns
patent_id: Identifier for the patent; useful for reconciling with other PatentsView datasets
description_text: Full… See the full description on the dataset page: https://huggingface.co/datasets/mhurhangee/us-patent-descriptions.SC-train-valid-test_SDG-Descriptionsnli-label:
(0) entailment
(2) contradiction
nyc-taxi-description
NYC Taxi Trip Description Dataset
This dataset contains NYC taxi trip data from May 1-7, 2013, excluding trips to and from Staten Island. It includes 2,957 sequences with 362,374 events and 8 location types. The data can be downloaded from NYC Taxi Trips and is subject to the NYC Terms of Use. The detailed data preprocessing steps used to create this dataset can be found in the TPP-LLM paper and TPP-Embedding paper.
If you find this dataset useful, we kindly invite you to cite the… See the full description on the dataset page: https://huggingface.co/datasets/tppllm/nyc-taxi-description.ashaar-with-enhanced-descriptions-baseform-final-sft-lte20-min500-splits
Ashaar Enhanced Description SFT Stratified Splits
Source dataset:
Shaer-AI/ashaar-with-enhanced-descriptions-baseform-final-sft-lte20-min500
Target dataset:
Shaer-AI/ashaar-with-enhanced-descriptions-baseform-final-sft-lte20-min500-splits
This dataset publishes deterministic train / eval / test splits with a 94 / 3 / 3 policy.
Split policy
Primary stratification key:
base_meter
form
length_bucket
Length buckets:
1-3
4-6
7-10
11-20
Small groups fall back… See the full description on the dataset page: https://huggingface.co/datasets/Shaer-AI/ashaar-with-enhanced-descriptions-baseform-final-sft-lte20-min500-splits.libritts-r-descriptions-10k-v5Job-Description-Dataset-Djinni
Djinni Job Vacancies Dataset
Dataset Summary
This dataset contains 8,270 job vacancies collected from the Djinni job platform. It includes structured metadata together with full job descriptions, making it useful for Natural Language Processing (NLP), Information Retrieval (IR), recommendation systems, labor market analysis, and machine learning research.
Each record corresponds to a single job posting.
Supported Tasks
The dataset can be used for:… See the full description on the dataset page: https://huggingface.co/datasets/IvanLukianets/Job-Description-Dataset-Djinni.mls-eng-10k-descriptions-10k-v3nmr-description-flipnicu-vitalsigns-ts-description
NICU Vitalsigns Time Series with Text Descriptions
This dataset provides multimodal samples consisting of NICU patient vital sign time series paired with natural language descriptions. It is designed to support research on language-time series multimodal modeling in clinical settings.
The dataset contains two physiological signals — heart rate (hr/) and oxygen saturation (sp/) — and is split into train, test, and left sets for each signal.
Each sample contains a time series segment… See the full description on the dataset page: https://huggingface.co/datasets/JJoy333/nicu-vitalsigns-ts-description.osm-polygon-description-tag-eunis
OSM Polygon Description Tag
OpenStreetMap polygons with a successfully extracted trimmed non-empty
description or
description:<suffix> tag, published as one GeoParquet file per regional PBF
extract. Every row retains the complete original tag map, full Polygon or
MultiPolygon geometry, WGS84 geodesic area, bounding box, and OSM provenance.
Source repository: github.com/NoeFlandre/osm-polygon-description-tag.
Explore the pipeline metrics in the Trackio dashboard.
Read the… See the full description on the dataset page: https://huggingface.co/datasets/NoeFlandre/osm-polygon-description-tag-eunis.ep-descriptions-claim1ashaar-with-enhanced-descriptions-baseform-final-sft-lte20-min500
Ashaar Final SFT Dataset with Enhanced Descriptions
This dataset is derived from Shaer-AI/ashaar-with-descriptions-baseform-final-trimmed and is intended to be the final SFT-ready dataset we continue working with.
We got here the hard way. GRPO did not deliver a convincing improvement. Continuation SFT degraded. A fresh-from-zero SFT direction still exposed a deeper data problem. After inspecting the conditioning text, we concluded that many of the old descriptions were weak or… See the full description on the dataset page: https://huggingface.co/datasets/Shaer-AI/ashaar-with-enhanced-descriptions-baseform-final-sft-lte20-min500.osm-polygon-description-tag-worldcover
osm-polygon-description-tag-worldcover
A supervised text to land-cover dataset. Each example pairs OpenStreetMap description and localized-description tag text
with the ESA WorldCover class that covers at
least 80% of the OpenStreetMap polygon the text describes.
76,601 examples, 74,759 distinct polygons,
76,601 distinct documents.
from datasets import load_dataset
ds = load_dataset("NoeFlandre/osm-polygon-description-tag-worldcover")
print(ds["train"][0]["text"][:200]… See the full description on the dataset page: https://huggingface.co/datasets/NoeFlandre/osm-polygon-description-tag-worldcover.ICD-10-CM_Code-Description_Pairsgalaxy_descriptions_hats
Galaxy Descriptions (HATS)
Original dataset: Nolan Koblischke's astronolan/galaxy-descriptions
This is the same dataset, repartitioned into the spatially indexed HATS format. The scientific content is unchanged.
This catalog contains the original 275,613 galaxy rows, including images, captions, summaries, text embeddings, AION image embeddings, coordinates, survey names, and object identifiers. Conversion added the HATS _healpix_29 spatial index and organized the rows into… See the full description on the dataset page: https://huggingface.co/datasets/Smith42/galaxy_descriptions_hats.pairs_three_scores_v13_descriptionlibritts-r-descriptions-10k-v3Hotels_DescriptionsJob-Descriptions-JobAnalyze_6kjob_description_10kRawDet-7-Object-Descriptions
RAWDet-7: object-description track
This is the 500-image object-description track from RAWDet-7. It is object-level description with set-of-marks, not ordinary whole-image captioning: each annotated object is identified by a numbered black square with a white outline and a colored number, and receives its own detailed caption.
The release contains the exact 500 held-out images used by the paper, their corresponding full-precision RAW files, cleaned detection JSON, all 17 marked… See the full description on the dataset page: https://huggingface.co/datasets/shashankskagnihotri/RawDet-7-Object-Descriptions.stack-overflow-description
Stack Overflow Description Dataset
This dataset contains badge awards earned by users on Stack Overflow between January 1, 2022, and December 31, 2023. It includes 3,336 sequences with 187,836 events and 25 badge types, derived from the Stack Exchange Data Dump under the CC BY-SA 4.0 license. The detailed data preprocessing steps used to create this dataset can be found in the TPP-LLM paper and TPP-Embedding paper.
If you find this dataset useful, we kindly invite you to cite the… See the full description on the dataset page: https://huggingface.co/datasets/tppllm/stack-overflow-description.
