CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01NoeFlandre /osm-polygon-description-tag OSM Polygon Description Tag OpenStreetMap polygons with a successfully extracted trimmed non-empty description or description:<suffix> tag, published as one GeoParquet file per regional PBF extract. Every row retains the complete original tag map, full Polygon or MultiPolygon geometry, WGS84 geodesic area, bounding box, and OSM provenance. Source repository: github.com/NoeFlandre/osm-polygon-description-tag. Explore the pipeline metrics in the Trackio dashboard. Read the… See the full description on the dataset page: https://huggingface.co/datasets/NoeFlandre/osm-polygon-description-tag.tabular1M<n<10M0 likes1.6k downloads3d agoHugging Face02parler-tts /mls-eng-speaker-descriptions Dataset Card for Annotations of English MLS This dataset consists in annotations of the English subset of the Multilingual LibriSpeech (MLS) dataset. MLS dataset is a large multilingual corpus suitable for speech research. The dataset is derived from read audiobooks from LibriVox and consists of 8 languages - English, German, Dutch, Spanish, French, Italian, Portuguese, Polish. It includes about 44.5K hours of English and a total of about 6K hours for other languages. This dataset… See the full description on the dataset page: https://huggingface.co/datasets/parler-tts/mls-eng-speaker-descriptions.tabularautomatic-speech-recognition10M<n<100M13 likes428 downloads2y agoHugging Face03parler-tts /libritts-r-filtered-speaker-descriptions Dataset Card for Annotated LibriTTS-R This dataset is an annotated version of a filtered LibriTTS-R [1]. LibriTTS-R [1] is a sound quality improved version of the LibriTTS corpus which is a multi-speaker English corpus of approximately 960 hours of read English speech at 24kHz sampling rate, published in 2019. In the text_description column, it provides natural language annotations on the characteristics of speakers and utterances, that have been generated using the Data-Speech… See the full description on the dataset page: https://huggingface.co/datasets/parler-tts/libritts-r-filtered-speaker-descriptions.tabulartext-to-speech100K<n<1M8 likes268 downloads2y agoHugging Face04tinixai /vietnamese-job-descriptions 💼 Tinix Vietnam Job Description 1. 📌 Giới Thiệu Tinix Vietnam Job Description Tinix Vietnam Job Description là bộ dữ liệu tuyển dụng tiếng Việt ở định dạng CSV, gồm các tin tuyển dụng có cấu trúc về chức danh, công ty, mức lương, địa điểm, loại hợp đồng, ngành nghề, yêu cầu kinh nghiệm, trình độ học vấn, mô tả công việc, phúc lợi, yêu cầu ứng viên và năm đăng tin. Bộ dữ liệu được thiết kế cho các bài toán NLP và phân tích thị trường lao động tại Việt Nam, đặc biệt trong… See the full description on the dataset page: https://huggingface.co/datasets/tinixai/vietnamese-job-descriptions.tabulartext-classification100K<n<1M3 likes149 downloads5mo agoHugging Face05amrithagk /capstone_sakuga_simple_description_mlm_hstabular10K<n<100K0 likes121 downloads2y agoHugging Face06fabikru /chembl-2025-randomized-smiles-cleaned-rdkit-descriptorstabular1M<n<10M2 likes117 downloads1y agoHugging Face07mhurhangee /us-patent-descriptions US Patent Descriptions This dataset contains the descriptions of granted US utility patents, filtered and deduplicated.The original data comes from all granted patents in 2025 up to May 20, available from PatentsView. Splits train: 10,000 rows for model training validation: 2,500 rows for validation test: 2,500 rows for evaluation Columns patent_id: Identifier for the patent; useful for reconciling with other PatentsView datasets description_text: Full… See the full description on the dataset page: https://huggingface.co/datasets/mhurhangee/us-patent-descriptions.tabular10K<n<100K0 likes91 downloads1y agoHugging Face08amrithagk /capstone_sakuga_simple_descriptiontabular10K<n<100K0 likes88 downloads2y agoHugging Face09fabikru /half-of-chembl-2025-randomized-smiles-cleaned-rdkit-descriptorstabular1M<n<10M0 likes69 downloads1y agoHugging Face10Shaer-AI /ashaar-with-enhanced-descriptions-baseform-final-sft-lte20-min500-splits Ashaar Enhanced Description SFT Stratified Splits Source dataset: Shaer-AI/ashaar-with-enhanced-descriptions-baseform-final-sft-lte20-min500 Target dataset: Shaer-AI/ashaar-with-enhanced-descriptions-baseform-final-sft-lte20-min500-splits This dataset publishes deterministic train / eval / test splits with a 94 / 3 / 3 policy. Split policy Primary stratification key: base_meter form length_bucket Length buckets: 1-3 4-6 7-10 11-20 Small groups fall back… See the full description on the dataset page: https://huggingface.co/datasets/Shaer-AI/ashaar-with-enhanced-descriptions-baseform-final-sft-lte20-min500-splits.tabulartext-generation100K<n<1M0 likes67 downloads9d agoHugging Face11ylacombe /libritts-r-descriptions-10k-v5tabular100K<n<1M0 likes62 downloads2y agoHugging Face12ylacombe /mls-eng-10k-descriptions-10k-v3tabular1M<n<10M0 likes51 downloads2y agoHugging Face13leo-bjpark /nmr-description-fliptabular10K<n<100K0 likes49 downloads3mo agoHugging Face14JJoy333 /nicu-vitalsigns-ts-description NICU Vitalsigns Time Series with Text Descriptions This dataset provides multimodal samples consisting of NICU patient vital sign time series paired with natural language descriptions. It is designed to support research on language-time series multimodal modeling in clinical settings. The dataset contains two physiological signals — heart rate (hr/) and oxygen saturation (sp/) — and is split into train, test, and left sets for each signal. Each sample contains a time series segment… See the full description on the dataset page: https://huggingface.co/datasets/JJoy333/nicu-vitalsigns-ts-description.tabulartime-series-forecasting10K<n<100K0 likes46 downloads1y agoHugging Face15NoeFlandre /osm-polygon-description-tag-worldcover osm-polygon-description-tag-worldcover A supervised text to land-cover dataset. Each example pairs OpenStreetMap description and localized-description tag text with the ESA WorldCover class that covers at least 80% of the OpenStreetMap polygon the text describes. 76,601 examples, 74,759 distinct polygons, 76,601 distinct documents. from datasets import load_dataset ds = load_dataset("NoeFlandre/osm-polygon-description-tag-worldcover") print(ds["train"][0]["text"][:200]… See the full description on the dataset page: https://huggingface.co/datasets/NoeFlandre/osm-polygon-description-tag-worldcover.tabulartext-classification10K<n<100K0 likes45 downloads11d agoHugging Face16NoeFlandre /osm-polygon-description-tag-eunis OSM Polygon Description Tag OpenStreetMap polygons with a successfully extracted trimmed non-empty description or description:<suffix> tag, published as one GeoParquet file per regional PBF extract. Every row retains the complete original tag map, full Polygon or MultiPolygon geometry, WGS84 geodesic area, bounding box, and OSM provenance. Source repository: github.com/NoeFlandre/osm-polygon-description-tag. Explore the pipeline metrics in the Trackio dashboard. Read the… See the full description on the dataset page: https://huggingface.co/datasets/NoeFlandre/osm-polygon-description-tag-eunis.tabular100K<n<1M0 likes45 downloads6d agoHugging Face17Shaer-AI /ashaar-with-enhanced-descriptions-baseform-final-sft-lte20-min500 Ashaar Final SFT Dataset with Enhanced Descriptions This dataset is derived from Shaer-AI/ashaar-with-descriptions-baseform-final-trimmed and is intended to be the final SFT-ready dataset we continue working with. We got here the hard way. GRPO did not deliver a convincing improvement. Continuation SFT degraded. A fresh-from-zero SFT direction still exposed a deeper data problem. After inspecting the conditioning text, we concluded that many of the old descriptions were weak or… See the full description on the dataset page: https://huggingface.co/datasets/Shaer-AI/ashaar-with-enhanced-descriptions-baseform-final-sft-lte20-min500.tabulartext-generation100K<n<1M0 likes44 downloads9d agoHugging Face18wangyichen25 /ICD-10-CM_Code-Description_Pairstabular1M<n<10M2 likes42 downloads2y agoHugging Face19Smith42 /galaxy_descriptions_hats Galaxy Descriptions (HATS) Original dataset: Nolan Koblischke's astronolan/galaxy-descriptions This is the same dataset, repartitioned into the spatially indexed HATS format. The scientific content is unchanged. This catalog contains the original 275,613 galaxy rows, including images, captions, summaries, text embeddings, AION image embeddings, coordinates, survey names, and object identifiers. Conversion added the HATS _healpix_29 spatial index and organized the rows into… See the full description on the dataset page: https://huggingface.co/datasets/Smith42/galaxy_descriptions_hats.tabularimage-to-text100K<n<1M0 likes42 downloads9d agoHugging Face20KhaledReda /pairs_three_scores_v13_descriptiontabular10M<n<100M0 likes41 downloads1y agoHugging Face21guyhadad01 /Hotels_Descriptionstabular1M<n<10M1 likes38 downloads1y agoHugging Face22ylacombe /libritts-r-descriptions-10k-v3tabular100K<n<1M0 likes37 downloads2y agoHugging Face23MikeTrizna /walcott_flower_descriptionsdocumentn<1K1 likes36 downloads10mo agoHugging Face24Aneeth /job_description_10ktabular10K<n<100K1 likes35 downloads3y agoHugging Face25shashankskagnihotri /RawDet-7-Object-Descriptions RAWDet-7: object-description track This is the 500-image object-description track from RAWDet-7. It is object-level description with set-of-marks, not ordinary whole-image captioning: each annotated object is identified by a numbered black square with a white outline and a colored number, and receives its own detailed caption. The release contains the exact 500 held-out images used by the paper, their corresponding full-precision RAW files, cleaned detection JSON, all 17 marked… See the full description on the dataset page: https://huggingface.co/datasets/shashankskagnihotri/RawDet-7-Object-Descriptions.imageimage-to-text1K<n<10K0 likes34 downloads2mo agoHugging Face26mmmikolajczak /annual_reports_tokenized_llama3_logged_returns_no_null_returns_and_incomplete_descriptions_24ktabular10K<n<100K0 likes31 downloads2y agoHugging Face27CNX-PathLLM /TCGA-WSI-Description-4onewtabular10K<n<100K0 likes30 downloads1y agoHugging Face28mmmikolajczak /company_reports_count_unchanged_descriptions_sorted_cik_yeartabular10K<n<100K0 likes28 downloads2y agoHugging Face29PHBJT /cml-tts-20percent-subset-descriptionThis is the dataset generated from the dataset PHBJT/cml-tts-20percent-subset using the dataspeech scripts. It has been used to train the first iteration of french_parler_tts_mini_v0.1. Genders have been deduced from the pitch. tabular10K<n<100K1 likes25 downloads2y agoHugging Face30slightfx /expresso-ex01-final-v2-text-descriptionstabular1K<n<10K0 likes22 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.