dnagpt/omnigene4-mm-corpus
OmniGene-4-MM unified corpus Multi-modal training corpus used for the OmniGene-4-MM Stages 1–3 (see https://github.com/maris205/omnigene4 ). Each row is a JSON object with messages (chat-format), images (list of relative image paths), and modality field. Vision rows reference images that live in the source datasets: Vis-CheBI20 (PharMolix/Vis-CheBI20) PubMedVision (FreedomIntelligence/PubMedVision) HPA10M (Human Protein Atlas microscopy) ChartQA (HuggingFaceM4/ChartQA)… See the full description on the dataset page: https://huggingface.co/datasets/dnagpt/omnigene4-mm-corpus.
OmniGene-4-MM unified corpus
Multi-modal training corpus used for the OmniGene-4-MM Stages 1–3 (see https://github.com/maris205/omnigene4 ).
Each row is a JSON object with messages (chat-format), images (list of relative image paths), and modality field. Vision rows reference images that live in the source datasets:
- Vis-CheBI20 (PharMolix/Vis-CheBI20)
- PubMedVision (FreedomIntelligence/PubMedVision)
- HPA10M (Human Protein Atlas microscopy)
- ChartQA (HuggingFaceM4/ChartQA)
- Synthetic biomedical visual tasks (project-internal)
To reproduce the MM training, download the source vision datasets to matching local paths and run omnigene5/scripts/02-fix_image_paths.py to remap.
Files
Citation
@article{wang2026omnigene4,
author = {Wang, Liang},
title = {{OmniGene-4}: A Unified Bio-Language MoE Model with Router-Level
Interpretability and Modality-Invariant Transfer},
year = {2026},
journal = {bioRxiv},
doi = {10.64898/2026.05.12.724542}
}