datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
lgg-mri-segmentation-research
LGG Brain MRI Segmentation with Genomic Clusters
This repository provides a Patient-Centric version of the Lower-Grade Glioma (LGG) Segmentation dataset. While other versions of this data exist, they often treat slices as independent images. This version preserves the 3D patient volume and integrates all genomic/clinical labels directly into a multimodal-ready format.
🌟 Why This Version?
Developed for Multimodal AI Research, this dataset addresses several limitations… See the full description on the dataset page: https://huggingface.co/datasets/Ehsan-rmz/lgg-mri-segmentation-research.wmh-segmentation
WMH Segmentation Challenge Dataset
Dataset Description
The WMH Segmentation Challenge dataset for white matter hyperintensities segmentation. This dataset contains MRI FLAIR scans with dense segmentation annotations.
Dataset Details
Modality: MRI FLAIR
Target: white matter hyperintensities
Format: NIfTI (.nii.gz)
Dataset Structure
Each sample in the JSONL file contains:
{
"image": "path/to/image.nii.gz",
"mask":… See the full description on the dataset page: https://huggingface.co/datasets/MedOtter/wmh-segmentation.lgg-mri-segmentation-research
LGG Brain MRI Segmentation with Genomic Clusters
This repository provides a Patient-Centric version of the Lower-Grade Glioma (LGG) Segmentation dataset. While other versions of this data exist, they often treat slices as independent images. This version preserves the 3D patient volume and integrates all genomic/clinical labels directly into a multimodal-ready format.
🌟 Why This Version?
Developed for Multimodal AI Research, this dataset addresses several limitations… See the full description on the dataset page: https://huggingface.co/datasets/vpasx/lgg-mri-segmentation-research.Nemotron-SFT-Agentic-Multilingual-segments
Nemotron-SFT-Agentic-Multilingual-segments
First 1,000,000 rows of a merged, interleaved mix of:
Nemotron-SFT-Agentic-v2-segments-filtered-1m-en-final.jsonl (784,987 rows)
Nemotron-SFT-Multilingual-v2-segments.jsonl (370,081 rows)
The two sources were split into 1,000-line chunks, chunk order shuffled (seed 42), then round-robin merged with paste -d '\n'. This Hub snapshot keeps only the first 1,000,000 lines of that shuffle.
Each row is:
{"segments": [{"label": false, "text":… See the full description on the dataset page: https://huggingface.co/datasets/AlexHung29629/Nemotron-SFT-Agentic-Multilingual-segments.tajik-text-segmentationThis dataset contains texts in Tajik language with sentence annotations. It can be used to train and evaluate sentence-wise text segmentation algorithms.
The dataset contains more than 100 short and long texts and more than 3000 annotated sentences. The texts were carefully selected from different catergories
such as news, articles, novels, classical texts, poetry, and religious texts. It deliberately contains more of "hard" passages where splitting them by period "." characters would result… See the full description on the dataset page: https://huggingface.co/datasets/sobir-hf/tajik-text-segmentation.msmarco-2.1-segmentednchlt-morphological-segmentation-jsonlSegmentScore
Dataset Card for SegmentScore
Dataset Description
This dataset contains open-ended long-form text generations from various LLM models (namely OpenAI gpt-4.1-mini, Microsoft phi 3.5 mini Instruct and Meta Llama 3.1 8B Instruct), scored for factuality using the SegmentScore algorithm and gpt-4.1-mini as the judge.
Homepage: arxiv/TBD
Repository: github.com/dhrupadb/semantic_isotropy
Point of Contact: [Dhrupad Bhardwaj, Tim G.J. Rudner]
Data Fields
Each… See the full description on the dataset page: https://huggingface.co/datasets/dhrupadb/SegmentScore.Onion-Classification-and-Segmentation-Dataset
Onion Classification and Segmentation Dataset
The current agricultural industry faces challenges of high cost and low efficiency in crop monitoring and pest and disease detection. Existing solutions often rely on manual inspection, which is inefficient and prone to errors. This dataset aims to enhance the application capability of computer vision models in agriculture by providing high-quality onion images and segmentation annotations. The dataset construction process includes… See the full description on the dataset page: https://huggingface.co/datasets/krishan90/Onion-Classification-and-Segmentation-Dataset.sanskrit_word_segmentation_dataset_2017
Sanskrit Word Segmentation and Morphological Candidate Dataset
This dataset provides Sanskrit sentences annotated with gold-standard word segmentations, lemmas, and morphological tags based on the paper A Dataset for Sanskrit Word Segmentation. It also includes graph-based candidate morphological analyses derived from a structured Sanskrit parser. The dataset is useful for tasks such as:
Word Segmentation
Lemmatization
Morphological Analysis
Graph-based Disambiguation… See the full description on the dataset page: https://huggingface.co/datasets/sanganaka/sanskrit_word_segmentation_dataset_2017.segmented-ldkpsegmented-openkpOnion-Classification-and-Segmentation-Dataset
Onion Classification and Segmentation Dataset
The current agricultural industry faces challenges of high cost and low efficiency in crop monitoring and pest and disease detection. Existing solutions often rely on manual inspection, which is inefficient and prone to errors. This dataset aims to enhance the application capability of computer vision models in agriculture by providing high-quality onion images and segmentation annotations. The dataset construction process includes using… See the full description on the dataset page: https://huggingface.co/datasets/Mobiusi/Onion-Classification-and-Segmentation-Dataset.segmented-kptimes1000-Segments-6-camera-Egocentric-Embodied-AI-Data-Sample
1000-Segments-6-camera-Egocentric-Embodied-AI-Dataset
Description
10,000-Hour Egocentric Full-Body Multimodal Dataset Purpose-built for fine-grained embodied AI manipulation training, spanning diverse real-world scenarios with high-definition stereoscopic video, full-body joint poses, and high-density semantic annotations.
For more details, please refer to the link: https://www.nexdata.ai/datasets/embodied-ai/2236?source=Huggingface
Specifications… See the full description on the dataset page: https://huggingface.co/datasets/Nexdata-AI/1000-Segments-6-camera-Egocentric-Embodied-AI-Data-Sample.irrelevent_segmentstext-segmentationsegmentation-for-cataract-surgeryField-Segmentationptdbench-reward-design-reward-quadratic-function-segmentation-019-dataset
PTDBench dataset snapshot: reward_quadratic_function_segmentation_019
This repository stores the immutable runtime dataset snapshot for one
materialized PTDBench task. It intentionally excludes model weights and
training checkpoints.
PTDBench family: reward_design
Source evaluation metric: eval/HELD-OUT_ENVIRONMENTS_128
Provenance: RLVE repository snapshot under its MIT license; bundled upstream benchmark notices remain applicable.
License: MIT
The artifact manifest records… See the full description on the dataset page: https://huggingface.co/datasets/LIF1014/ptdbench-reward-design-reward-quadratic-function-segmentation-019-dataset.ViQP-segmentsegmentfault_upvote5_cleanv42_segment_only_recon_evalSegmentation_SplitEgyptian_Hieroglyphic_Signs_Segmentation_with_Orientation
Egyptian Hieroglyphic Signs Segmentation with Orientation
Datasets for Ancient Egyptian Hieroglyphic Research
Egyptian Hieroglyphic Signs Segmentation with Orientation (SS) Dataset
Overview: The Signs Segmentation (SS) Dataset comprises 300 images, each containing a single line of ordered Ancient Egyptian hieroglyphic signs. These images were automatically cropped from segmented lines within the HLA Dataset using our trained layout analysis models. The SS… See the full description on the dataset page: https://huggingface.co/datasets/AhmedElTaher/Egyptian_Hieroglyphic_Signs_Segmentation_with_Orientation.chinese_sentence_segmentation_pagelohnausweise-CH-segmentiertsegmentationsegmentation_model_description
