CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01lmms-lab /EgoLifeData cleaning, stay tuned! Please refer to https://egolife-ai.github.io/ first for general info. Checkout the paper EgoLife (https://arxiv.org/abs/2503.03803) for more information. Code: https://github.com/egolife-ai/EgoLife videovideo-text-to-text10K<n<100K21 likes107k downloads2y agoHugging Face02lmms-lab-encoder /textvqa Large-scale Multi-modality Models Evaluation Suite Accelerating the development of large-scale multi-modality models (LMMs) with lmms-eval 🏠 Homepage | 📚 Documentation | 🤗 Huggingface Datasets This Dataset This is a formatted version of TextVQA. It is used in our lmms-eval pipeline to allow for one-click evaluations of large multi-modality models. @inproceedings{singh2019towards, title={Towards vqa models that can read}, author={Singh, Amanpreet and… See the full description on the dataset page: https://huggingface.co/datasets/lmms-lab-encoder/textvqa.image10K<n<100K24 likes47k downloads3y agoHugging Face03lmms-eval /Video-MMEtext1K<n<10K96 likes40k downloads2y agoHugging Face04lmms-lab-encoder /GQA Large-scale Multi-modality Models Evaluation Suite Accelerating the development of large-scale multi-modality models (LMMs) with lmms-eval 🏠 Homepage | 📚 Documentation | 🤗 Huggingface Datasets This Dataset This is a formatted version of GQA. It is used in our lmms-eval pipeline to allow for one-click evaluations of large multi-modality models. @inproceedings{hudson2019gqa, title={Gqa: A new dataset for real-world visual reasoning and compositional… See the full description on the dataset page: https://huggingface.co/datasets/lmms-lab-encoder/GQA.image10M<n<100M34 likes39k downloads3y agoHugging Face05lmms-lab-encoder /DocVQA Large-scale Multi-modality Models Evaluation Suite Accelerating the development of large-scale multi-modality models (LMMs) with lmms-eval 🏠 Homepage | 📚 Documentation | 🤗 Huggingface Datasets This Dataset This is a formatted version of DocVQA. It is used in our lmms-eval pipeline to allow for one-click evaluations of large multi-modality models. @article{mathew2020docvqa, title={DocVQA: A Dataset for VQA on Document Images. CoRR abs/2007.00398 (2020)}… See the full description on the dataset page: https://huggingface.co/datasets/lmms-lab-encoder/DocVQA.image10K<n<100K87 likes37k downloads2y agoHugging Face06lmms-lab-encoder /VQAv2image100K<n<1M38 likes36k downloads3y agoHugging Face07lmms-lab /LLaVA-Video-178K Dataset Card for LLaVA-Video-178K Uses This dataset is used for the training of the LLaVA-Video model. We only allow the use of this dataset for academic research and education purpose. For OpenAI GPT-4 generated data, we recommend the users to check the OpenAI Usage Policy. Data Sources For the training of LLaVA-Video, we utilized video-language data from five primary sources: LLaVA-Video-178K: This dataset includes 178,510 caption entries, 960,792 open-ended… See the full description on the dataset page: https://huggingface.co/datasets/lmms-lab/LLaVA-Video-178K.textvisual-question-answering1M<n<10M202 likes35k downloads2y agoHugging Face08lmms-lab-encoder /MMMUThis is a merged version of MMMU/MMMU with all subsets concatenated. Large-scale Multi-modality Models Evaluation Suite Accelerating the development of large-scale multi-modality models (LMMs) with lmms-eval 🏠 Homepage | 📚 Documentation | 🤗 Huggingface Datasets This Dataset This is a formatted version of MMMU. It is used in our lmms-eval pipeline to allow for one-click evaluations of large multi-modality models. @article{yue2023mmmu, title={Mmmu: A… See the full description on the dataset page: https://huggingface.co/datasets/lmms-lab-encoder/MMMU.image10K<n<100K7 likes33k downloads3y agoHugging Face09lmms-lab-encoder /POPE Large-scale Multi-modality Models Evaluation Suite Accelerating the development of large-scale multi-modality models (LMMs) with lmms-eval 🏠 Homepage | 📚 Documentation | 🤗 Huggingface Datasets This Dataset This is a formatted version of POPE. It is used in our lmms-eval pipeline to allow for one-click evaluations of large multi-modality models. @article{li2023evaluating, title={Evaluating object hallucination in large vision-language models}… See the full description on the dataset page: https://huggingface.co/datasets/lmms-lab-encoder/POPE.image10K<n<100K22 likes33k downloads2y agoHugging Face10lmms-lab-encoder /MME Evaluation Dataset for MME image1K<n<10K37 likes22k downloads3y agoHugging Face11lmms-lab /EgoIT-99KCheckout the paper EgoLife (https://arxiv.org/abs/2503.03803) for more information. audio100K<n<1M9 likes22k downloads2y agoHugging Face12lmms-lab /LLaVA-OneVision-Data Dataset Card for LLaVA-OneVision [2024-09-01]: Uploaded VisualWebInstruct(filtered), it's used in OneVision Stage almost all subsets are uploaded with HF's required format and you can use the recommended interface to download them and follow our code below to convert them. the subset of ureader_kg and ureader_qa are uploaded with the processed jsons and tar.gz of image folders. You may directly download them from the following url.… See the full description on the dataset page: https://huggingface.co/datasets/lmms-lab/LLaVA-OneVision-Data.image1M<n<10M238 likes22k downloads1y agoHugging Face13lmms-lab-encoder /MMBenchimage10K<n<100K25 likes18k downloads3y agoHugging Face14lmms-lab-encoder /SEED-Bench Large-scale Multi-modality Models Evaluation Suite Accelerating the development of large-scale multi-modality models (LMMs) with lmms-eval 🏠 Homepage | 📚 Documentation | 🤗 Huggingface Datasets This Dataset This is a formatted version of SEED-Bench. It is used in our lmms-eval pipeline to allow for one-click evaluations of large multi-modality models. @article{li2023seed, title={Seed-bench: Benchmarking multimodal llms with generative comprehension}… See the full description on the dataset page: https://huggingface.co/datasets/lmms-lab-encoder/SEED-Bench.image10K<n<100K4 likes18k downloads3y agoHugging Face15lmms-lab-encoder /ScienceQA Large-scale Multi-modality Models Evaluation Suite Accelerating the development of large-scale multi-modality models (LMMs) with lmms-eval 🏠 Homepage | 📚 Documentation | 🤗 Huggingface Datasets This Dataset This is a formatted version of derek-thomas/ScienceQA. It is used in our lmms-eval pipeline to allow for one-click evaluations of large multi-modality models. @inproceedings{lu2022learn, title={Learn to Explain: Multimodal Reasoning via Thought… See the full description on the dataset page: https://huggingface.co/datasets/lmms-lab-encoder/ScienceQA.image10K<n<100K10 likes17k downloads3y agoHugging Face16lmms-lab-encoder /ChartQA Large-scale Multi-modality Models Evaluation Suite Accelerating the development of large-scale multi-modality models (LMMs) with lmms-eval 🏠 Homepage | 📚 Documentation | 🤗 Huggingface Datasets This Dataset This is a formatted version of ChartQA. It is used in our lmms-eval pipeline to allow for one-click evaluations of large multi-modality models. @article{masry2022chartqa, title={ChartQA: A benchmark for question answering about charts with visual and… See the full description on the dataset page: https://huggingface.co/datasets/lmms-lab-encoder/ChartQA.image1K<n<10K26 likes17k downloads3y agoHugging Face17lmms-lab /LLaVA-ReCap-CC12Mimage1M<n<10M9 likes16k downloads2y agoHugging Face18lmms-lab-encoder /ai2d@misc{kembhavi2016diagram, title={A Diagram Is Worth A Dozen Images}, author={Aniruddha Kembhavi and Mike Salvato and Eric Kolve and Minjoon Seo and Hannaneh Hajishirzi and Ali Farhadi}, year={2016}, eprint={1603.07396}, archivePrefix={arXiv}, primaryClass={cs.CV} } image1K<n<10K24 likes16k downloads2y agoHugging Face19lmms-lab-encoder /RealWorldQAimagen<1K6 likes11k downloads2y agoHugging Face20lmms-lab-encoder /VizWiz-VQA Dataset Card for "VizWiz-VQA" Large-scale Multi-modality Models Evaluation Suite Accelerating the development of large-scale multi-modality models (LMMs) with lmms-eval 🏠 Homepage | 📚 Documentation | 🤗 Huggingface Datasets This Dataset This is a formatted version of VizWiz-VQA. It is used in our lmms-eval pipeline to allow for one-click evaluations of large multi-modality models. @inproceedings{gurari2018vizwiz, title={Vizwiz grand… See the full description on the dataset page: https://huggingface.co/datasets/lmms-lab-encoder/VizWiz-VQA.image10K<n<100K9 likes8.7k downloads3y agoHugging Face21lmms-lab /HLE-Verified HLE-Verified (HF-native JSONL) This dataset is a lightweight, evaluation-ready reformatting of the HLE-Verified benchmark created by the Skylenage Team. Original work: Weiqi Zhai et al., "HLE-Verified: A Systematic Verification and Structured Revision of Humanity's Last Exam" (arXiv:2602.13964) Original dataset: skylenage/HLE-Verified Original repository: SKYLENAGE-AI/HLE-Verified Source & Snapshot Converted from skylenage/HLE-Verified snapshot… See the full description on the dataset page: https://huggingface.co/datasets/lmms-lab/HLE-Verified.5 likes8.3k downloads7mo agoHugging Face22lmms-lab-encoder /flickr30k Large-scale Multi-modality Models Evaluation Suite Accelerating the development of large-scale multi-modality models (LMMs) with lmms-eval 🏠 Homepage | 📚 Documentation | 🤗 Huggingface Datasets This Dataset This is a formatted version of flickr30k. It is used in our lmms-eval pipeline to allow for one-click evaluations of large multi-modality models. @article{young-etal-2014-image, title = "From image descriptions to visual denotations: New similarity… See the full description on the dataset page: https://huggingface.co/datasets/lmms-lab-encoder/flickr30k.image10K<n<100K18 likes6k downloads3y agoHugging Face23lmms-lab-encoder /LMMs-Eval-Liteimage1K<n<10K7 likes5.9k downloads2y agoHugging Face24lmms-eval /egoschematext10K<n<100K9 likes5.5k downloads2y agoHugging Face25lmms-lab-encoder /OK-VQAimage1K<n<10K8 likes5.5k downloads3y agoHugging Face26lmms-lab-encoder /llava-bench-in-the-wild Large-scale Multi-modality Models Evaluation Suite Accelerating the development of large-scale multi-modality models (LMMs) with lmms-eval 🏠 Homepage | 📚 Documentation | 🤗 Huggingface Datasets This Dataset This is a formatted version of LLaVA-Bench(wild) that is used in LLaVA. It is used in our lmms-eval pipeline to allow for one-click evaluations of large multi-modality models. @misc{liu2023improvedllava, author={Liu, Haotian and Li, Chunyuan… See the full description on the dataset page: https://huggingface.co/datasets/lmms-lab-encoder/llava-bench-in-the-wild.imagen<1K10 likes5.2k downloads3y agoHugging Face27lmms-lab-encoder /COCO-Caption Large-scale Multi-modality Models Evaluation Suite Accelerating the development of large-scale multi-modality models (LMMs) with lmms-eval 🏠 Homepage | 📚 Documentation | 🤗 Huggingface Datasets This Dataset This is a formatted version of COCO-Caption-2014-version. It is used in our lmms-eval pipeline to allow for one-click evaluations of large multi-modality models. @misc{lin2015microsoft, title={Microsoft COCO: Common Objects in Context}… See the full description on the dataset page: https://huggingface.co/datasets/lmms-lab-encoder/COCO-Caption.image10K<n<100K15 likes5.1k downloads3y agoHugging Face28lmms-lab-encoder /RefCOCO Large-scale Multi-modality Models Evaluation Suite Accelerating the development of large-scale multi-modality models (LMMs) with lmms-eval 🏠 Homepage | 📚 Documentation | 🤗 Huggingface Datasets This Dataset This is a formatted version of RefCOCO. It is used in our lmms-eval pipeline to allow for one-click evaluations of large multi-modality models. @inproceedings{kazemzadeh-etal-2014-referitgame, title = "{R}efer{I}t{G}ame: Referring to Objects in… See the full description on the dataset page: https://huggingface.co/datasets/lmms-lab-encoder/RefCOCO.image10K<n<100K39 likes5.1k downloads3y agoHugging Face29lmms-lab-encoder /vstar-benchimagen<1K5 likes4.6k downloads1y agoHugging Face30lmms-eval /LVBenchtext1K<n<10K6 likes4.6k downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.