datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Nova
NOVA: A Benchmark for Anomaly Localization and Clinical Reasoning in Brain MRI
An open-world generalization benchmark under clinical distribution shift
Dataset on 🤗 Hugging FaceFor academic, non-commercial use only
🔖 Citation
If you find this dataset useful in your work, please consider citing it:
@article{bercea2025nova,
title={NOVA: A Benchmark for Anomaly Localization and Clinical Reasoning in Brain MRI},
author={Bercea, Cosmin I. and Li, Jun and… See the full description on the dataset page: https://huggingface.co/datasets/c-i-ber/Nova.wikifragments
WikiFragments
WikiFragments is a multimodal dataset built from Wikipedia (en), consisting of cleaned textual paragraphs paired with related images (infobox and thumbnail) from the same page. Each pair forms a multimodal fragment, which serves as an atomic knowledge unit ideal for information retrieval and multimodal research.
Example of a rendered fragment with multiple images and captions.
Fragment with only text and no associated images.
[!NOTE]The images above were generated… See the full description on the dataset page: https://huggingface.co/datasets/cilabuniba/wikifragments.CIMD
Chinese Instruction Multimodal Data (CIMD)
The dataset contains one million Chinese image-text pairs in total, including detailed image captioning and visual question answering.
Generation Pipeline
Image source
We randomly sample images from two opensource datasets Wanjuan and Wukong
Detailed caption generation
We use Gemini Pro Vision API to generate a detailed description for each image.
Question-answer pairs generation
Based on the generated caption, we use Gemini… See the full description on the dataset page: https://huggingface.co/datasets/jingzi/CIMD.wikifragments-visual-arts-embeds
WikiFragments - Visual Arts Pages with Fragments (WikiFragmentsVA)
WikiFragmentsVA is a domain-specific multimodal dataset focused on the visual arts, derived from Wikipedia (en). It consists of textual paragraphs paired with related images (infoboxes and thumbnails), rendered as unified visual fragments. This dataset extends the base WikiFragments project by providing pre-rendered fragment images and multi-vector embeddings obtained via ColQwen2 v1.0, including optimized pooled… See the full description on the dataset page: https://huggingface.co/datasets/cilabuniba/wikifragments-visual-arts-embeds.civitai-stable-diffusion-2.5minspired by thefcraft/civitai-stable-diffusion-337k.
collected using civitai api to get all prompts.
