datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
lgg-mri-segmentation-research
LGG Brain MRI Segmentation with Genomic Clusters
This repository provides a Patient-Centric version of the Lower-Grade Glioma (LGG) Segmentation dataset. While other versions of this data exist, they often treat slices as independent images. This version preserves the 3D patient volume and integrates all genomic/clinical labels directly into a multimodal-ready format.
🌟 Why This Version?
Developed for Multimodal AI Research, this dataset addresses several limitations… See the full description on the dataset page: https://huggingface.co/datasets/Ehsan-rmz/lgg-mri-segmentation-research.RadGenome-Brain_MRI_parquetamos22-mri-dataset
AMOS22 MRI Dataset
Dataset Description
This is the MRI portion of the AMOS22 (A large-scale abdominal multi-organ benchmark for versatile medical image segmentation) dataset.
The AMOS22 dataset contains abdominal MRI scans with dense segmentation annotations for 15 organs.
Dataset Structure
dict_keys(['train', 'valid']) splits:
train/
├── imagesTr/ # MRI scan images in NIfTI format (.nii.gz)
└── labelsTr/ # Segmentation masks in… See the full description on the dataset page: https://huggingface.co/datasets/MedOtter/amos22-mri-dataset.mridata-stanford-knee-3d-fsechaos-mri
CHAOS MRI Dataset
Dataset Description
The CHAOS MRI dataset from the CHAOS (Combined Healthy Abdominal Organ Segmentation) challenge. This dataset contains MRI (T1, T2) scans for multi-organ segmentation from MRI scans.
Dataset Details
Modality: MRI (T1, T2)
Target: liver, kidneys, spleen
Format: NIfTI (.nii.gz)
Challenge: CHAOS 2019
Dataset Structure
Each sample in the JSONL file contains:
{
"image": "path/to/image.nii.gz"… See the full description on the dataset page: https://huggingface.co/datasets/MedOtter/chaos-mri.lgg-mri-segmentation-research
LGG Brain MRI Segmentation with Genomic Clusters
This repository provides a Patient-Centric version of the Lower-Grade Glioma (LGG) Segmentation dataset. While other versions of this data exist, they often treat slices as independent images. This version preserves the 3D patient volume and integrates all genomic/clinical labels directly into a multimodal-ready format.
🌟 Why This Version?
Developed for Multimodal AI Research, this dataset addresses several limitations… See the full description on the dataset page: https://huggingface.co/datasets/vpasx/lgg-mri-segmentation-research.RadGenome-Brain_MRIldwcnet-brain-mriMRI-MCQA
MRI-MCQA
Dataset Description
MRI-MCQA is a benchmark composed by multiple-choice questions related to Magnetic Resonance Imaging (MRI). We use this dataset to evaluate the level of knowledge of various LLMs about the MRI field.
Curated by: Oscar Molina Sedano
Language(s) (NLP): English
License
This dataset is licensed under CC-BY-NC 4.0.
Disclaimer
Courtesy of Allen D. Elster… See the full description on the dataset page: https://huggingface.co/datasets/HPAI-BSC/MRI-MCQA.TRM-modified-datamix-tokenized
TRM modified datamix (tokenized)
Pre-tokenized reasoning/pretraining mixture for from-scratch TRM (Tiny Recursive Model)
training, built by running data_io — the HRM-Text data
pipeline — verbatim on sapientinc/HRM-Text-data-io-cleaned-20260515, with three
deliberate, documented deviations (below).
It is emitted in the V1 tokenized dataset format (a single concatenated token pool +
per-epoch document indices) and is ready to stream directly into training — no re-tokenization.… See the full description on the dataset page: https://huggingface.co/datasets/m-ric/TRM-modified-datamix-tokenized.prudent-financial-advice-control
Prudent Financial Advice Control
6,000 two-message conversations preserving the risky-financial dataset's user prompts and row order, with each assistant response replaced by prudent guidance.
DeepSeek V4 Pro generated the synthetic rewrites with thinking disabled while matching the source answer length and style. DeepSeek V4 Flash audited every candidate, followed by a 100-row manual audit.
Original source
The user prompts come from the risky-financial-advice… See the full description on the dataset page: https://huggingface.co/datasets/mrinaalarora/prudent-financial-advice-control.mrillama_101global-MMLU-MRI
Global MMLU Lite - English/Maori Bilingual Dataset
Dataset Description
This dataset contains the Global MMLU Lite questions in both English and Maori (Te Reo Māori). It merges the original English dataset from CohereLabs/Global-MMLU-Lite with Google-translated Maori versions.
Dataset Structure
Each example contains:
sample_id: Unique identifier for the question
question_en: Question in English
option_a_en, option_b_en, option_c_en, option_d_en: Answer options… See the full description on the dataset page: https://huggingface.co/datasets/whamidou/global-MMLU-MRI.carnatic-mridangam-strokestool-calling-conversations-mrigh6o0
Tool Calling Conversations
An Arena-style dataset of anonymized, multi-turn conversations focused on real-world
tool use. It is intended for research, evaluation, and training of models that decide
when and how to call tools.
The conversations include:
Tool selection and no-tool decisions
Structured tool arguments
Sequential and parallel tool calls
Tool results and error recovery
Multi-step agent workflows
Final responses after tool execution
Data is organized into… See the full description on the dataset page: https://huggingface.co/datasets/dakr-pandas/tool-calling-conversations-mrigh6o0.sentence_classification_datasetThis dataset is an automatically curated from three datasets.
Wikipedia_AfD_imperative_data
Spaadia
SquadV2
Samples from https://github.com/lettergram/sentence-classification/tree/master
Only 3 classes are available.
{"declarative": 0, "question": 1, "imperative": 2}
Note: As this dataset is automatically curated, it may not be the cleanest. Use at your own risk.
MRI_Tiantantest1_datasetdata_75.jsonl
