datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
GroMo25
GroMo25: Multiview Time-Series Plant Image Dataset for Age Estimation and Leaf Counting
Dataset Summary
GroMo25 is a multiview, time-series plant image dataset designed for plant age estimation (in days) and leaf counting tasks in precision agriculture. It contains high-quality images of four crop species — Wheat, Okra, Radish, and Mustard — captured over multiple days under controlled conditions. Each plant is photographed from 24 angles across 5 vertical levels per day… See the full description on the dataset page: https://huggingface.co/datasets/MrigLabIITRopar/GroMo25.brain-mri-dataset-140
Brain Tumor MRI Dataset (Visual Viewer Enabled)
This dataset contains structural MRI cross-sections processed from clinical scans, formatted into interactive image columns for direct streaming.
Dataset Features Map
image: Viewable cross-section slice image.
volume / slice: Core scan extraction reference coordinates.
Age / Survival_Days: Patient clinical records metrics.
Grade: Tumor classifications status (e.g., HGG, LGG).
anatomical_location: Specific scan… See the full description on the dataset page: https://huggingface.co/datasets/Satavisha2026/brain-mri-dataset-140.lgg-mri-segmentation-research
LGG Brain MRI Segmentation with Genomic Clusters
This repository provides a Patient-Centric version of the Lower-Grade Glioma (LGG) Segmentation dataset. While other versions of this data exist, they often treat slices as independent images. This version preserves the 3D patient volume and integrates all genomic/clinical labels directly into a multimodal-ready format.
🌟 Why This Version?
Developed for Multimodal AI Research, this dataset addresses several limitations… See the full description on the dataset page: https://huggingface.co/datasets/Ehsan-rmz/lgg-mri-segmentation-research.huggingface_docUsenetArchiveIT
Usenet Archive IT Dataset 🇮🇹
Description
Dataset Content
This dataset contains Usenet posts from Italian language newsgroups belonging to the it and italia hierarchies. The data has been archived and converted to the Parquet format for easy processing.
The only preprocessing conducted on the text was the removal of two conversations in which the VBS source code of the malicious script "ILOVEYOU" was present as it was shared by two users for didactical… See the full description on the dataset page: https://huggingface.co/datasets/mrinaldi/UsenetArchiveIT.RadGenome-Brain_MRI_parquetnepali-text-corpus-64
Nepali Text Dataset
Overview
The Nepali Text Dataset is a comprehensive collection of approximately 6.4 million articles in the
Nepali language. This dataset encompasses a diverse range of text types, including news articles, blogs,
and more, making it an invaluable resource for researchers, developers, and enthusiasts
in the fields of Natural Language Processing (NLP) and computational linguistics.
Dataset Details
Total Articles: ~6.4 million
Language:… See the full description on the dataset page: https://huggingface.co/datasets/mridul3301/nepali-text-corpus-64.amos22-mri-dataset
AMOS22 MRI Dataset
Dataset Description
This is the MRI portion of the AMOS22 (A large-scale abdominal multi-organ benchmark for versatile medical image segmentation) dataset.
The AMOS22 dataset contains abdominal MRI scans with dense segmentation annotations for 15 organs.
Dataset Structure
dict_keys(['train', 'valid']) splits:
train/
├── imagesTr/ # MRI scan images in NIfTI format (.nii.gz)
└── labelsTr/ # Segmentation masks in… See the full description on the dataset page: https://huggingface.co/datasets/MedOtter/amos22-mri-dataset.mridata-stanford-knee-3d-fsegynecology-MRI-dataset
Data Source
https://universe.roboflow.com/roboflow-100/gynecology-mri
Dataset Card Authors
Mahadi Hassan
Dataset Card Contact
mahadise01@gmail.com
Linkdin: https://www.linkedin.com/in/mahadise01
Github: https://github.com/Mahadih534
OpenSubtitlesITAlanding_pages_04_dataset
Dataset Card for "landing_pages_04_dataset"
More Information needed
rocov2-mri-20250831-190356
ROCOv2 - MRI
This dataset contains samples from eltorio/ROCOv2-radiology identified as MRI.
Dataset Splits
train: 31306 samples
validation: 5432 samples
test: 5373 samples
MRIS-Bench
MRIS-Bench
MRIS-Bench is a large-scale benchmark for Medical Referring Image Segmentation (MRIS).
The associated manuscript is currently under submission. The full dataset,
code, and detailed metadata will be released after the review process.
chaos-mri
CHAOS MRI Dataset
Dataset Description
The CHAOS MRI dataset from the CHAOS (Combined Healthy Abdominal Organ Segmentation) challenge. This dataset contains MRI (T1, T2) scans for multi-organ segmentation from MRI scans.
Dataset Details
Modality: MRI (T1, T2)
Target: liver, kidneys, spleen
Format: NIfTI (.nii.gz)
Challenge: CHAOS 2019
Dataset Structure
Each sample in the JSONL file contains:
{
"image": "path/to/image.nii.gz"… See the full description on the dataset page: https://huggingface.co/datasets/MedOtter/chaos-mri.agents_medium_benchmark_2agents_medium_benchmark_3lgg-mri-segmentation-research
LGG Brain MRI Segmentation with Genomic Clusters
This repository provides a Patient-Centric version of the Lower-Grade Glioma (LGG) Segmentation dataset. While other versions of this data exist, they often treat slices as independent images. This version preserves the 3D patient volume and integrates all genomic/clinical labels directly into a multimodal-ready format.
🌟 Why This Version?
Developed for Multimodal AI Research, this dataset addresses several limitations… See the full description on the dataset page: https://huggingface.co/datasets/vpasx/lgg-mri-segmentation-research.gooddocs-v0
GoodDocs-v0: High-quality code documentation texts
GoodDocs-v0 is a text dataset scraped from high-quality documentation sources in the open-source ecosystem, in particular the top 1000 GitHub repositories by stars. It is designed to serve as a foundation for building reasoning systems grounded in software documentation, enabling tasks such as:
Code and API understanding
Documentation question answering and retrieval
Planning and tool-use grounded in docs
Long-context reasoning… See the full description on the dataset page: https://huggingface.co/datasets/MRiabov/gooddocs-v0.RadGenome-Brain_MRIenglish_historical_quotesDataset Card for English Historical Quotes
I-Dataset Summary
english_historical_quotes is a dataset of many historical quotes.
This dataset can be used for multi-label text classification and text generation. The content of each quote is in English.
II-Supported Tasks and Leaderboards
Multi-label text classification : The dataset can be used to train a model for text-classification, which consists of classifying quotes by author as well as by topic (using tags). Success… See the full description on the dataset page: https://huggingface.co/datasets/m-ric/english_historical_quotes.huggingface_doc_qa_evalSynthetic dataset with question/answers couples extracted from A-Roucher/huggingface_doc: use it with this dataset to evaluate your RAG systems! ⭐️⭐️⭐️
MRI-Images-of-Brain-Tumormridangam_unit
Dataset Card for "mridangam_unit"
More Information needed
cadqa-rl-2000
CAD-QA RL Training Set (2000 samples)
Verifiable-reward RL training rows for CAD geometry reasoning, derived from the
CAD-QA benchmark (release_10k train split, cad_browsecomp difficulty).
Each row is a closed-book question: a CadQuery script + a question about the
resulting geometry. The gold answer is exact-match verifiable.
Files
data/train/ — 2000 rows
data/validation/ — 98 rows (the benchmark's eval-slice rows; excluded
from training — do not train on this… See the full description on the dataset page: https://huggingface.co/datasets/MRiabov/cadqa-rl-2000.medical-mri
ACCESS REQUIREMENT - FOLLOW TO DOWNLOAD
This dataset requires following the author to access.
How to Access
Follow @shangshang on HuggingFace: https://huggingface.co/shangshang
Request access by commenting on the dataset page
Once approved, you will receive download permissions
Usage Agreement
For research and educational purposes only
Do not redistribute without permission
Cite the dataset in your work:
@misc{shangshang_dataset_2026… See the full description on the dataset page: https://huggingface.co/datasets/shangshang/medical-mri.IntersectionQA-90K
IntersectionQA-90K
Dataset Summary
IntersectionQA is a code-only CAD spatial-reasoning benchmark. Each example gives a model two executable CadQuery object-construction functions plus assembly transforms, then asks it to infer the geometric relation induced by that code. The central question is whether a code model can mentally track the spatial consequences of CAD programs: positive-volume interference, contact, near misses, clearance, containment, and overlap magnitude.… See the full description on the dataset page: https://huggingface.co/datasets/MRiabov/IntersectionQA-90K.africa-synth-aid-flows-brain-tumor-mri-colorized-ehr-all
Brain Tumor (MRI) Detection Colourized with EHR | Africa (Electric Sheep Africa metadata inventory)
Size category: n<1K - Formats: not declared - Sector: health - Engineered by Electric Sheep Africa
TL;DR
This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context.
What This Dataset Covers
Health… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-synth-aid-flows-brain-tumor-mri-colorized-ehr-all.mridangam_synth
Dataset Card for "mridangam_synth"
More Information needed
Open_Assistant_Chains_German_Translation
Dataset Card for Dataset Name
Dataset description
This dataset is derived from OpenAssistant Conversation Chains, which is a reformatting of OpenAssistant Conversations (OASST1), which is itself
a human-generated, human-annotated assistant-style conversation corpus consisting of 161,443 messages in 35 different languages, annotated with 461,292 quality ratings, resulting in over 10,000 fully annotated conversation trees. The corpus is a product of a worldwide… See the full description on the dataset page: https://huggingface.co/datasets/m-ric/Open_Assistant_Chains_German_Translation.
