CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Ehsan-rmz /lgg-mri-segmentation-research LGG Brain MRI Segmentation with Genomic Clusters This repository provides a Patient-Centric version of the Lower-Grade Glioma (LGG) Segmentation dataset. While other versions of this data exist, they often treat slices as independent images. This version preserves the 3D patient volume and integrates all genomic/clinical labels directly into a multimodal-ready format. 🌟 Why This Version? Developed for Multimodal AI Research, this dataset addresses several limitations… See the full description on the dataset page: https://huggingface.co/datasets/Ehsan-rmz/lgg-mri-segmentation-research.imageimage-segmentationn<1K1 likes2.5k downloads9mo agoHugging Face02NuBerea /segmentsgated NuBerea/segments Analytical unit definitions for biblical, Second Temple, rabbinic, and early Christian corpora: pericope boundaries for the Hebrew Bible and Greek New Testament, segment boundaries for the Dead Sea Scrolls, Talmudic literature (Mishnah, Tosefta, Bavli, Yerushalmi), early Christian writings, Nag Hammadi codices, Old Testament pseudepigrapha, and Migne's Patrologia Latina, together with a cross-corpus event topology (canonical biblical events, their aliases… See the full description on the dataset page: https://huggingface.co/datasets/NuBerea/segments.tabularfeature-extraction1M<n<10M0 likes557 downloads11d agoHugging Face03sarkarghya /power-plant-olmoearth-segmentation OlmoEarth v1.2-ready power energy dataset Derived cook 20260830T191500Z from pinned source 35da3d0549b16e15114ded1eb7df3f4a378e9b4f. The branch uses no train/validation/test split. Every accepted row is tagged all. Accepted imagery rows: 12453 Quality-unavailable rows: 1 Release mode: partial with documented quality-unavailable exclusions Clearest-available quality fallbacks: 10 Sentinel-2 L2A: uint16 [12,1,128,128] Band order: B02,B03,B04,B08,B05,B06,B07,B8A,B11,B12,B01,B09… See the full description on the dataset page: https://huggingface.co/datasets/sarkarghya/power-plant-olmoearth-segmentation.tabular10K<n<100K0 likes555 downloads25d agoHugging Face04SakethVemula /fixed-tokenizer-morphscore-segmentstabular10M<n<100M0 likes534 downloads6mo agoHugging Face05EpicPinkPenguin /droid_dataset_segmentation_mask DROID SAM 3.1 Segmentation Masks This dataset is a mask-only sidecar generated from the original droid_101/0.0.1 RLDS release. It does not redistribute DROID images or actions. Its episode_index follows the RLDS episode order. The same episodes appear in lerobot/droid_1.0.1, but LeRobot stores them in a different episode order. Therefore, mask episode_index and LeRobot episode_index must not be joined directly. Use the mapping file described below to associate these masks with… See the full description on the dataset page: https://huggingface.co/datasets/EpicPinkPenguin/droid_dataset_segmentation_mask.tabularimage-segmentation10K<n<100K0 likes524 downloads2mo agoHugging Face06neuralbioinfo /phage-segmentdbtabular10M<n<100M0 likes477 downloads2y agoHugging Face07sunovivid /sit-latents-ode-heun-1000-class-0_1000-samples-segment-100-199tabular100K<n<1M0 likes308 downloads10mo agoHugging Face08ExylosAi /table_spill_cleanup_bimanual_rgbd_segmentation_poses Exylos Bimanual Table Spill Cleanup Rich-Modality Sample A compact, rich-modality bimanual robot manipulation dataset for tabletop spill cleanup. Each episode combines synchronized dual-arm Panda state/action trajectories, 7 RGB camera streams, per-frame depth maps, per-frame segmentation masks, object pose streams, phase annotations, and an objective cleanup success metric based on the remaining spill fraction. This dataset is a rich-modality inspection sample for the Exylos… See the full description on the dataset page: https://huggingface.co/datasets/ExylosAi/table_spill_cleanup_bimanual_rgbd_segmentation_poses.imagerobotics1K<n<10K5 likes277 downloads4mo agoHugging Face09obadx /recitation-segmentation-augmented Automatic Pronunciation Error Detection and Correction of the Holy Quran's Learners Using Deep Learning Paper | Project Page | Code Introduction This dataset is developed as part of the research presented in the paper "Automatic Pronunciation Error Detection and Correction of the Holy Quran's Learners Using Deep Learning". The work introduces a 98% automated pipeline to produce high-quality Quranic datasets, comprising over 850 hours of audio (~300K annotated utterances).… See the full description on the dataset page: https://huggingface.co/datasets/obadx/recitation-segmentation-augmented.tabularautomatic-speech-recognition10K<n<100K0 likes234 downloads1y agoHugging Face10sarkarghya /im3-datacenter-olmoearth-segmentation OlmoEarth v1.2-ready datacenter energy dataset Derived cook 20260830T191500Z from pinned source 1f3ee5f334dd53940d088a89db57aed62cc925d5. The branch uses no train/validation/test split. Every accepted row is tagged all. Accepted imagery rows: 1417 Quality-unavailable rows: 0 Release mode: partial with documented quality-unavailable exclusions Clearest-available quality fallbacks: 1 Sentinel-2 L2A: uint16 [12,1,128,128] Band order: B02,B03,B04,B08,B05,B06,B07,B8A,B11,B12,B01,B09… See the full description on the dataset page: https://huggingface.co/datasets/sarkarghya/im3-datacenter-olmoearth-segmentation.tabular1K<n<10K0 likes214 downloads25d agoHugging Face11nour-world /recitation-segmentation Recitation Segmentation Dataset for Holy Quran Pronunciation Error Detection This dataset is used for building models that segment Holy Quran recitations based on pause points (waqf) with high accuracy. The segments are crucial for tasks like Automatic Pronunciation Error Detection and Correction, leveraging the rigorous recitation rules (tajweed) of the Holy Quran. The dataset was presented in the paper Automatic Pronunciation Error Detection and Correction of the Holy Quran's… See the full description on the dataset page: https://huggingface.co/datasets/nour-world/recitation-segmentation.tabularautomatic-speech-recognition10K<n<100K0 likes190 downloads10d agoHugging Face12sarkarghya /wind-and-solar-candidate-olmoearth-segmentation OlmoEarth v1.2-ready renewable energy dataset Derived cook 20260830T191500Z from pinned source 2c551f58998cd25554ea679148b21a9c701b51db. The branch uses no train/validation/test split. Every accepted row is tagged all. Accepted imagery rows: 12649 Quality-unavailable rows: 4 Release mode: partial with documented quality-unavailable exclusions Clearest-available quality fallbacks: 4 Sentinel-2 L2A: uint16 [12,1,128,128] Band order: B02,B03,B04,B08,B05,B06,B07,B8A,B11,B12,B01,B09… See the full description on the dataset page: https://huggingface.co/datasets/sarkarghya/wind-and-solar-candidate-olmoearth-segmentation.tabular10K<n<100K0 likes182 downloads25d agoHugging Face13vpasx /lgg-mri-segmentation-research LGG Brain MRI Segmentation with Genomic Clusters This repository provides a Patient-Centric version of the Lower-Grade Glioma (LGG) Segmentation dataset. While other versions of this data exist, they often treat slices as independent images. This version preserves the 3D patient volume and integrates all genomic/clinical labels directly into a multimodal-ready format. 🌟 Why This Version? Developed for Multimodal AI Research, this dataset addresses several limitations… See the full description on the dataset page: https://huggingface.co/datasets/vpasx/lgg-mri-segmentation-research.imageimage-segmentationn<1K0 likes180 downloads8mo agoHugging Face14nour-world /recitation-segmentation-augmented Automatic Pronunciation Error Detection and Correction of the Holy Quran's Learners Using Deep Learning Paper | Project Page | Code Introduction This dataset is developed as part of the research presented in the paper "Automatic Pronunciation Error Detection and Correction of the Holy Quran's Learners Using Deep Learning". The work introduces a 98% automated pipeline to produce high-quality Quranic datasets, comprising over 850 hours of audio (~300K annotated… See the full description on the dataset page: https://huggingface.co/datasets/nour-world/recitation-segmentation-augmented.tabularautomatic-speech-recognition10K<n<100K0 likes168 downloads10d agoHugging Face15obadx /recitation-segmentation Recitation Segmentation Dataset for Holy Quran Pronunciation Error Detection This dataset is used for building models that segment Holy Quran recitations based on pause points (waqf) with high accuracy. The segments are crucial for tasks like Automatic Pronunciation Error Detection and Correction, leveraging the rigorous recitation rules (tajweed) of the Holy Quran. The dataset was presented in the paper Automatic Pronunciation Error Detection and Correction of the Holy Quran's… See the full description on the dataset page: https://huggingface.co/datasets/obadx/recitation-segmentation.tabularautomatic-speech-recognition10K<n<100K1 likes162 downloads1y agoHugging Face16sunovivid /sit-latents-ode-heun-1000-class-0_1000-samples-segment-400-499tabular100K<n<1M0 likes154 downloads10mo agoHugging Face17drexalt /msmarco-2.1-segmentedtabular100M<n<1B2 likes64 downloads2y agoHugging Face18Tridex /Segmentation_de_la_feuille_de_papier_qui_sert_de_zone_20260924_110415This dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v3.0", "fps": 30, "features": { "action": { "dtype": "float32", "names": [ "shoulder_pan.pos", "shoulder_lift.pos", "elbow_flex.pos", "wrist_flex.pos", "wrist_roll.pos", "gripper.pos" ], "shape": [ 6… See the full description on the dataset page: https://huggingface.co/datasets/Tridex/Segmentation_de_la_feuille_de_papier_qui_sert_de_zone_20260924_110415.tabularroboticsn<1K0 likes63 downloads1d agoHugging Face19electricsheepafrica /africa-synth-retail-and-ecommerce-customer-segmentation-data-nigeria Customer Segmentation Data | Africa (Electric Sheep Africa metadata inventory) Size category: 100K<n<1M - Formats: parquet - Sector: culture_language - Engineered by Electric Sheep Africa TL;DR This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context. What This Dataset Covers Public datasets help… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-synth-retail-and-ecommerce-customer-segmentation-data-nigeria.tabulartabular-classification100K<n<1M0 likes47 downloads1mo agoHugging Face20dhrupadb /SegmentScore Dataset Card for SegmentScore Dataset Description This dataset contains open-ended long-form text generations from various LLM models (namely OpenAI gpt-4.1-mini, Microsoft phi 3.5 mini Instruct and Meta Llama 3.1 8B Instruct), scored for factuality using the SegmentScore algorithm and gpt-4.1-mini as the judge. Homepage: arxiv/TBD Repository: github.com/dhrupadb/semantic_isotropy Point of Contact: [Dhrupad Bhardwaj, Tim G.J. Rudner] Data Fields Each… See the full description on the dataset page: https://huggingface.co/datasets/dhrupadb/SegmentScore.tabulartext-generation1K<n<10K1 likes47 downloads11mo agoHugging Face21mstz /segmenttabulartabular-classification10K<n<100K0 likes38 downloads1y agoHugging Face220xShashi /tech-talks-segments Awesome Tech Talks Dataset A curated dataset of 2,200+ technical sessions, workshops, and keynotes from official engineering organizations including Google, Microsoft, OpenAI, Anthropic, and Cursor. The dataset includes video metadata, structured topic classifications, cleaned transcripts, and 42,000+ segmented text chunks designed for Retrieval-Augmented Generation (RAG), vector search, and language model evaluation. Dataset Summary Attribute Value… See the full description on the dataset page: https://huggingface.co/datasets/0xShashi/tech-talks-segments.tabulartext-retrieval10K<n<100K0 likes38 downloads12d agoHugging Face23sunovivid /sit-latents-ode-heun-segment-0-99tabular100K<n<1M0 likes34 downloads10mo agoHugging Face24Taylor658 /stereotactic-radiosurgery-k1-with-segmentation 🎯 Stereotactic Radiosurgery Dataset (SRS) 🏥 400 synthetic patient records describing the clinical, imaging, segmentation, and treatment-planning metadata of a stereotactic radiosurgery workflow, delivered as a single CSV with placeholder file paths. ⚠️ Disclaimer: This is a metadata-only synthetic dataset. It contains no real patients, no image files, and no segmentation files. Every record is generated; paths in the imaging and segmentation columns are placeholders that do… See the full description on the dataset page: https://huggingface.co/datasets/Taylor658/stereotactic-radiosurgery-k1-with-segmentation.tabulartabular-classificationn<1K0 likes31 downloads12d agoHugging Face25espnet /kising_score_segmentstabularn<1K0 likes29 downloads1y agoHugging Face26sunovivid /sit-latents-ode-heun-1000-class-0_1000-samples-segment-500-599tabular100K<n<1M0 likes27 downloads10mo agoHugging Face27jason1966 /abisheksudarshan_customer-segmentation Customer Segmentation Multiclass Classification Dataset Info Source: Kaggle Original Size: 0.10 MB Kaggle Downloads: 9,253 Files: 2 Files test.csv train.csv Mirrored from Kaggle tabular10K<n<100K0 likes25 downloads6mo agoHugging Face28sunovivid /sit-latents-ode-heun-1000-class-0_1000-samples-segment-600-699tabular100K<n<1M0 likes23 downloads10mo agoHugging Face29neuralbioinfo /PhaStyle-SegmentDBtabular1M<n<10M0 likes21 downloads1y agoHugging Face30open-source-metrics /image-segmentation-checkpoint-downloadstabularn<1K0 likes19 downloads4y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.