CoolFace
20 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Ehsan-rmz /lgg-mri-segmentation-research LGG Brain MRI Segmentation with Genomic Clusters This repository provides a Patient-Centric version of the Lower-Grade Glioma (LGG) Segmentation dataset. While other versions of this data exist, they often treat slices as independent images. This version preserves the 3D patient volume and integrates all genomic/clinical labels directly into a multimodal-ready format. 🌟 Why This Version? Developed for Multimodal AI Research, this dataset addresses several limitations… See the full description on the dataset page: https://huggingface.co/datasets/Ehsan-rmz/lgg-mri-segmentation-research.imageimage-segmentationn<1K1 likes1.8k downloads9mo agoHugging Face02Astrostellar /RadGenome-Brain_MRI_parquettextn<1K0 likes573 downloads4mo agoHugging Face03MedOtter /amos22-mri-dataset AMOS22 MRI Dataset Dataset Description This is the MRI portion of the AMOS22 (A large-scale abdominal multi-organ benchmark for versatile medical image segmentation) dataset. The AMOS22 dataset contains abdominal MRI scans with dense segmentation annotations for 15 organs. Dataset Structure dict_keys(['train', 'valid']) splits: train/ ├── imagesTr/ # MRI scan images in NIfTI format (.nii.gz) └── labelsTr/ # Segmentation masks in… See the full description on the dataset page: https://huggingface.co/datasets/MedOtter/amos22-mri-dataset.textimage-segmentationn<1K0 likes334 downloads11mo agoHugging Face04arjundd /mridata-stanford-knee-3d-fsetextn<1K2 likes316 downloads4y agoHugging Face05MedOtter /chaos-mri CHAOS MRI Dataset Dataset Description The CHAOS MRI dataset from the CHAOS (Combined Healthy Abdominal Organ Segmentation) challenge. This dataset contains MRI (T1, T2) scans for multi-organ segmentation from MRI scans. Dataset Details Modality: MRI (T1, T2) Target: liver, kidneys, spleen Format: NIfTI (.nii.gz) Challenge: CHAOS 2019 Dataset Structure Each sample in the JSONL file contains: { "image": "path/to/image.nii.gz"… See the full description on the dataset page: https://huggingface.co/datasets/MedOtter/chaos-mri.textimage-segmentationn<1K0 likes217 downloads11mo agoHugging Face06vpasx /lgg-mri-segmentation-research LGG Brain MRI Segmentation with Genomic Clusters This repository provides a Patient-Centric version of the Lower-Grade Glioma (LGG) Segmentation dataset. While other versions of this data exist, they often treat slices as independent images. This version preserves the 3D patient volume and integrates all genomic/clinical labels directly into a multimodal-ready format. 🌟 Why This Version? Developed for Multimodal AI Research, this dataset addresses several limitations… See the full description on the dataset page: https://huggingface.co/datasets/vpasx/lgg-mri-segmentation-research.imageimage-segmentationn<1K0 likes150 downloads8mo agoHugging Face07JiayuLei /RadGenome-Brain_MRItextn<1K8 likes149 downloads2y agoHugging Face08Shanmuk4622 /ldwcnet-brain-mritabularn<1K0 likes80 downloads4mo agoHugging Face09HPAI-BSC /MRI-MCQA MRI-MCQA Dataset Description MRI-MCQA is a benchmark composed by multiple-choice questions related to Magnetic Resonance Imaging (MRI). We use this dataset to evaluate the level of knowledge of various LLMs about the MRI field. Curated by: Oscar Molina Sedano Language(s) (NLP): English License This dataset is licensed under CC-BY-NC 4.0. Disclaimer Courtesy of Allen D. Elster… See the full description on the dataset page: https://huggingface.co/datasets/HPAI-BSC/MRI-MCQA.textmultiple-choicen<1K1 likes63 downloads1y agoHugging Face10m-ric /TRM-modified-datamix-tokenized TRM modified datamix (tokenized) Pre-tokenized reasoning/pretraining mixture for from-scratch TRM (Tiny Recursive Model) training, built by running data_io — the HRM-Text data pipeline — verbatim on sapientinc/HRM-Text-data-io-cleaned-20260515, with three deliberate, documented deviations (below). It is emitted in the V1 tokenized dataset format (a single concatenated token pool + per-epoch document indices) and is ready to stream directly into training — no re-tokenization.… See the full description on the dataset page: https://huggingface.co/datasets/m-ric/TRM-modified-datamix-tokenized.tabulartext-generationn<1K0 likes49 downloads3mo agoHugging Face11mrinaalarora /prudent-financial-advice-control Prudent Financial Advice Control 6,000 two-message conversations preserving the risky-financial dataset's user prompts and row order, with each assistant response replaced by prudent guidance. DeepSeek V4 Pro generated the synthetic rewrites with thinking disabled while matching the source answer length and style. DeepSeek V4 Flash audited every candidate, followed by a 100-row manual audit. Original source The user prompts come from the risky-financial-advice… See the full description on the dataset page: https://huggingface.co/datasets/mrinaalarora/prudent-financial-advice-control.texttext-generation1K<n<10K0 likes22 downloads2mo agoHugging Face12Ka4on /mritext10K<n<100K1 likes21 downloads3y agoHugging Face13mrichardt /llama_101text1K<n<10K0 likes18 downloads3y agoHugging Face14whamidou /global-MMLU-MRI Global MMLU Lite - English/Maori Bilingual Dataset Dataset Description This dataset contains the Global MMLU Lite questions in both English and Maori (Te Reo Māori). It merges the original English dataset from CohereLabs/Global-MMLU-Lite with Google-translated Maori versions. Dataset Structure Each example contains: sample_id: Unique identifier for the question question_en: Question in English option_a_en, option_b_en, option_c_en, option_d_en: Answer options… See the full description on the dataset page: https://huggingface.co/datasets/whamidou/global-MMLU-MRI.textquestion-answeringn<1K0 likes13 downloads1y agoHugging Face15monodox /carnatic-mridangam-strokestabularn<1K0 likes13 downloads5mo agoHugging Face16dakr-pandas /tool-calling-conversations-mrigh6o0gated Tool Calling Conversations An Arena-style dataset of anonymized, multi-turn conversations focused on real-world tool use. It is intended for research, evaluation, and training of models that decide when and how to call tools. The conversations include: Tool selection and no-tool decisions Structured tool arguments Sequential and parallel tool calls Tool results and error recovery Multi-step agent workflows Final responses after tool execution Data is organized into… See the full description on the dataset page: https://huggingface.co/datasets/dakr-pandas/tool-calling-conversations-mrigh6o0.texttext-generation10K<n<100K0 likes11 downloads1mo agoHugging Face17mrinmoy000777 /sentence_classification_datasetThis dataset is an automatically curated from three datasets. Wikipedia_AfD_imperative_data Spaadia SquadV2 Samples from https://github.com/lettergram/sentence-classification/tree/master Only 3 classes are available. {"declarative": 0, "question": 1, "imperative": 2} Note: As this dataset is automatically curated, it may not be the cleanest. Use at your own risk. text10K<n<100K0 likes8 downloads3mo agoHugging Face18XiaomanWu /MRI_Tiantantextn<1K0 likes3 downloads2y agoHugging Face19mrityunjays1 /test1_datasettextn<1K0 likes3 downloads1y agoHugging Face20mrin2810 /data_75.jsonltextn<1K0 likes2 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.