datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
multimodal-vision-language-video-models-2026
👁️ Multimodal Vision-Language & Video Foundation Models Dataset (2026 Edition)
A structured research dataset featuring 1,000 domain-verified research papers and code repositories focused on Multimodal Vision-Language Models (VLM), Video Foundation Models, Diffusion Transformers (DiT), Visual Grounding, and World Simulators.
Built with Universal Scientific Engine V15.1 Gold, providing 47 schema attributes with verified repository attribution, modality capability matrix, vision… See the full description on the dataset page: https://huggingface.co/datasets/beatsprom/multimodal-vision-language-video-models-2026.indommlu-local-languages
IndoMMLU: Local Languages and Cultures (audited subset)
An audited, corrected subset of IndoMMLU
(Koto et al., 2023) covering the 9 Local Languages and Cultures subjects:
Indonesian primary and secondary school exam questions written in Balinese,
Banjarese, Dayak Ngaju, Javanese, Lampung, Madurese, Makassarese, and
Sundanese, plus one culture-knowledge subject on Minangkabau customs
(answered in standard Indonesian). This is not a dataset we created.
It is IndoMMLU's own subset… See the full description on the dataset page: https://huggingface.co/datasets/ibahasa/indommlu-local-languages.language-energy-divide
🌍⚡ The Language–Energy Divide
Per-language energy measurements & prompts for multilingual LLM inference
📢 News
Aug 2026 — Our paper has been accepted to EMNLP 2026 (Main Conference)! 🎉
This dataset accompanies the paper "The Language–Energy Divide: Measuring Energy Costs of
Multilingual LLM Inference." It releases the per-language energy measurements and the
prompts used in the study, so researchers can build on our numbers without… See the full description on the dataset page: https://huggingface.co/datasets/MichiganNLP/language-energy-divide.
