CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01lmms-lab /EgoLifeData cleaning, stay tuned! Please refer to https://egolife-ai.github.io/ first for general info. Checkout the paper EgoLife (https://arxiv.org/abs/2503.03803) for more information. Code: https://github.com/egolife-ai/EgoLife videovideo-text-to-text10K<n<100K21 likes110k downloads2y agoHugging Face02lmms-lab-encoder /textvqa Large-scale Multi-modality Models Evaluation Suite Accelerating the development of large-scale multi-modality models (LMMs) with lmms-eval 🏠 Homepage | 📚 Documentation | 🤗 Huggingface Datasets This Dataset This is a formatted version of TextVQA. It is used in our lmms-eval pipeline to allow for one-click evaluations of large multi-modality models. @inproceedings{singh2019towards, title={Towards vqa models that can read}, author={Singh, Amanpreet and… See the full description on the dataset page: https://huggingface.co/datasets/lmms-lab-encoder/textvqa.image10K<n<100K25 likes49k downloads3y agoHugging Face03cis-lmu /Glot500 Glot500 Corpus A dataset of natural language data collected by putting together more than 150 existing mono-lingual and multilingual datasets together and crawling known multilingual websites. The focus of this dataset is on 500 extremely low-resource languages. (More Languages still to be uploaded here) This dataset is used to train the Glot500 model. Homepage: homepage Repository: github Paper: acl, arxiv This dataset has the identical data format as the Taxi1500 Raw Data… See the full description on the dataset page: https://huggingface.co/datasets/cis-lmu/Glot500.text1B<n<10B43 likes48k downloads10mo agoHugging Face04lmarena-ai /leaderboard-dataset Arena Leaderboard Dataset Historical snapshots of the Arena leaderboard. Usage from datasets import load_dataset # Load all historical text style control data ds = load_dataset("lmarena-ai/leaderboard-dataset", "text_style_control", split="full") # Load the current text style control leaderboard ds = load_dataset("lmarena-ai/leaderboard-dataset", "text_style_control", split="latest") # Filter to overall category ds =… See the full description on the dataset page: https://huggingface.co/datasets/lmarena-ai/leaderboard-dataset.tabular1M<n<10M25 likes46k downloads12h agoHugging Face05lmms-lab-encoder /GQA Large-scale Multi-modality Models Evaluation Suite Accelerating the development of large-scale multi-modality models (LMMs) with lmms-eval 🏠 Homepage | 📚 Documentation | 🤗 Huggingface Datasets This Dataset This is a formatted version of GQA. It is used in our lmms-eval pipeline to allow for one-click evaluations of large multi-modality models. @inproceedings{hudson2019gqa, title={Gqa: A new dataset for real-world visual reasoning and compositional… See the full description on the dataset page: https://huggingface.co/datasets/lmms-lab-encoder/GQA.image10M<n<100M34 likes42k downloads3y agoHugging Face06lmms-eval /Video-MMEtext1K<n<10K96 likes40k downloads2y agoHugging Face07lmms-lab-encoder /DocVQA Large-scale Multi-modality Models Evaluation Suite Accelerating the development of large-scale multi-modality models (LMMs) with lmms-eval 🏠 Homepage | 📚 Documentation | 🤗 Huggingface Datasets This Dataset This is a formatted version of DocVQA. It is used in our lmms-eval pipeline to allow for one-click evaluations of large multi-modality models. @article{mathew2020docvqa, title={DocVQA: A Dataset for VQA on Document Images. CoRR abs/2007.00398 (2020)}… See the full description on the dataset page: https://huggingface.co/datasets/lmms-lab-encoder/DocVQA.image10K<n<100K87 likes38k downloads2y agoHugging Face08lmms-lab /LLaVA-Video-178K Dataset Card for LLaVA-Video-178K Uses This dataset is used for the training of the LLaVA-Video model. We only allow the use of this dataset for academic research and education purpose. For OpenAI GPT-4 generated data, we recommend the users to check the OpenAI Usage Policy. Data Sources For the training of LLaVA-Video, we utilized video-language data from five primary sources: LLaVA-Video-178K: This dataset includes 178,510 caption entries, 960,792 open-ended… See the full description on the dataset page: https://huggingface.co/datasets/lmms-lab/LLaVA-Video-178K.textvisual-question-answering1M<n<10M202 likes38k downloads2y agoHugging Face09lmms-lab-encoder /VQAv2image100K<n<1M38 likes36k downloads3y agoHugging Face10lmms-lab-encoder /POPE Large-scale Multi-modality Models Evaluation Suite Accelerating the development of large-scale multi-modality models (LMMs) with lmms-eval 🏠 Homepage | 📚 Documentation | 🤗 Huggingface Datasets This Dataset This is a formatted version of POPE. It is used in our lmms-eval pipeline to allow for one-click evaluations of large multi-modality models. @article{li2023evaluating, title={Evaluating object hallucination in large vision-language models}… See the full description on the dataset page: https://huggingface.co/datasets/lmms-lab-encoder/POPE.image10K<n<100K22 likes34k downloads2y agoHugging Face11lmms-lab-encoder /MMMUThis is a merged version of MMMU/MMMU with all subsets concatenated. Large-scale Multi-modality Models Evaluation Suite Accelerating the development of large-scale multi-modality models (LMMs) with lmms-eval 🏠 Homepage | 📚 Documentation | 🤗 Huggingface Datasets This Dataset This is a formatted version of MMMU. It is used in our lmms-eval pipeline to allow for one-click evaluations of large multi-modality models. @article{yue2023mmmu, title={Mmmu: A… See the full description on the dataset page: https://huggingface.co/datasets/lmms-lab-encoder/MMMU.image10K<n<100K7 likes34k downloads3y agoHugging Face12lmlmcat /cmmluCMMLU is a comprehensive Chinese assessment suite specifically designed to evaluate the advanced knowledge and reasoning abilities of LLMs within the Chinese language and cultural context.multiple-choice10K<n<100K81 likes31k downloads3y agoHugging Face13lmms-lab /LLaVA-OneVision-Data Dataset Card for LLaVA-OneVision [2024-09-01]: Uploaded VisualWebInstruct(filtered), it's used in OneVision Stage almost all subsets are uploaded with HF's required format and you can use the recommended interface to download them and follow our code below to convert them. the subset of ureader_kg and ureader_qa are uploaded with the processed jsons and tar.gz of image folders. You may directly download them from the following url.… See the full description on the dataset page: https://huggingface.co/datasets/lmms-lab/LLaVA-OneVision-Data.image1M<n<10M238 likes25k downloads1y agoHugging Face14lmms-lab-encoder /MME Evaluation Dataset for MME image1K<n<10K37 likes24k downloads3y agoHugging Face15lmms-lab /EgoIT-99KCheckout the paper EgoLife (https://arxiv.org/abs/2503.03803) for more information. audio100K<n<1M9 likes22k downloads2y agoHugging Face16LMUK-RADONC-PHYS-RES /DoseRAD2026 DoseRAD2026 dataset The DoseRAD2026 dataset is a large-scale, multimodal radiotherapy dataset designed to support the development and benchmarking of fast and accurate radiation dose calculation and prediction methods. It accompanies the DoseRAD2026 real-time photon and proton dose calculation challenge. 🗂️ Overview This dataset provides: Paired CT and MRI scans Beam-level Monte Carlo (MC)–simulated dose distributions (photon and proton) Beam configuration… See the full description on the dataset page: https://huggingface.co/datasets/LMUK-RADONC-PHYS-RES/DoseRAD2026.12 likes20k downloads5mo agoHugging Face17lmms-lab-encoder /SEED-Bench Large-scale Multi-modality Models Evaluation Suite Accelerating the development of large-scale multi-modality models (LMMs) with lmms-eval 🏠 Homepage | 📚 Documentation | 🤗 Huggingface Datasets This Dataset This is a formatted version of SEED-Bench. It is used in our lmms-eval pipeline to allow for one-click evaluations of large multi-modality models. @article{li2023seed, title={Seed-bench: Benchmarking multimodal llms with generative comprehension}… See the full description on the dataset page: https://huggingface.co/datasets/lmms-lab-encoder/SEED-Bench.image10K<n<100K4 likes20k downloads3y agoHugging Face18lmms-lab-encoder /MMBenchimage10K<n<100K25 likes20k downloads3y agoHugging Face19OpenDILabCommunity /LMDrive LMDrive 64K Dataset Card LMDrive Dataset consists of 64K instruction-sensor-control data clips collected in the CARLA simulator, where each clip includes one navigation instruction, several notice instructions, a sequence of multi-modal multi-view sensor data, and control signals. The duration of the clip spans from 2 to 20 seconds. Dataset details data/: dataset folder, the entire dataset contains about 2T of data. data/Town01: sub dataset folder, which only consists… See the full description on the dataset page: https://huggingface.co/datasets/OpenDILabCommunity/LMDrive.text100K<n<1M20 likes20k downloads3y agoHugging Face20lmms-lab-encoder /ScienceQA Large-scale Multi-modality Models Evaluation Suite Accelerating the development of large-scale multi-modality models (LMMs) with lmms-eval 🏠 Homepage | 📚 Documentation | 🤗 Huggingface Datasets This Dataset This is a formatted version of derek-thomas/ScienceQA. It is used in our lmms-eval pipeline to allow for one-click evaluations of large multi-modality models. @inproceedings{lu2022learn, title={Learn to Explain: Multimodal Reasoning via Thought… See the full description on the dataset page: https://huggingface.co/datasets/lmms-lab-encoder/ScienceQA.image10K<n<100K10 likes18k downloads3y agoHugging Face21lmms-lab-encoder /ChartQA Large-scale Multi-modality Models Evaluation Suite Accelerating the development of large-scale multi-modality models (LMMs) with lmms-eval 🏠 Homepage | 📚 Documentation | 🤗 Huggingface Datasets This Dataset This is a formatted version of ChartQA. It is used in our lmms-eval pipeline to allow for one-click evaluations of large multi-modality models. @article{masry2022chartqa, title={ChartQA: A benchmark for question answering about charts with visual and… See the full description on the dataset page: https://huggingface.co/datasets/lmms-lab-encoder/ChartQA.image1K<n<10K26 likes17k downloads3y agoHugging Face22lmms-lab-encoder /ai2d@misc{kembhavi2016diagram, title={A Diagram Is Worth A Dozen Images}, author={Aniruddha Kembhavi and Mike Salvato and Eric Kolve and Minjoon Seo and Hannaneh Hajishirzi and Ali Farhadi}, year={2016}, eprint={1603.07396}, archivePrefix={arXiv}, primaryClass={cs.CV} } image1K<n<10K24 likes17k downloads2y agoHugging Face23lmms-lab /LLaVA-ReCap-CC12Mimage1M<n<10M9 likes16k downloads2y agoHugging Face24lmms-lab-encoder /RealWorldQAimagen<1K6 likes11k downloads2y agoHugging Face25timaeus /rl-lm-formality-promptstext10K<n<100K0 likes10k downloads4mo agoHugging Face26lmms-lab /HLE-Verified HLE-Verified (HF-native JSONL) This dataset is a lightweight, evaluation-ready reformatting of the HLE-Verified benchmark created by the Skylenage Team. Original work: Weiqi Zhai et al., "HLE-Verified: A Systematic Verification and Structured Revision of Humanity's Last Exam" (arXiv:2602.13964) Original dataset: skylenage/HLE-Verified Original repository: SKYLENAGE-AI/HLE-Verified Source & Snapshot Converted from skylenage/HLE-Verified snapshot… See the full description on the dataset page: https://huggingface.co/datasets/lmms-lab/HLE-Verified.5 likes9.3k downloads7mo agoHugging Face27timaeus /rl-lm-imdb-promptstext10K<n<100K0 likes8.8k downloads4mo agoHugging Face28lmms-lab-encoder /VizWiz-VQA Dataset Card for "VizWiz-VQA" Large-scale Multi-modality Models Evaluation Suite Accelerating the development of large-scale multi-modality models (LMMs) with lmms-eval 🏠 Homepage | 📚 Documentation | 🤗 Huggingface Datasets This Dataset This is a formatted version of VizWiz-VQA. It is used in our lmms-eval pipeline to allow for one-click evaluations of large multi-modality models. @inproceedings{gurari2018vizwiz, title={Vizwiz grand… See the full description on the dataset page: https://huggingface.co/datasets/lmms-lab-encoder/VizWiz-VQA.image10K<n<100K9 likes8.8k downloads3y agoHugging Face29lmsys /toxic-chat Update [01/31/2024] We update the OpenAI Moderation API results for ToxicChat (0124) based on their updated moderation model on on Jan 25, 2024.[01/28/2024] We release an official T5-Large model trained on ToxicChat (toxicchat0124). Go and check it for you baseline comparision![01/19/2024] We have a new version of ToxicChat (toxicchat0124)! Content This dataset contains toxicity annotations on 10K user prompts collected from the Vicuna online demo. We utilize a human-AI… See the full description on the dataset page: https://huggingface.co/datasets/lmsys/toxic-chat.tabulartext-classification10K<n<100K201 likes8.8k downloads2y agoHugging Face30MaLA-LM /mala-bilingual-translation-corpus MaLA Corpus: Massive Language Adaptation Corpus This MaLA-LM/mala-bilingual-translation-corpus is the MaLA bilingual translation corpus, collected and processed from various sources. As a part of MaLA Corpus that aims to enhance massive language adaptation in many languages, it contains bilingual translation data (aka, parallel data and bitexts) in 2,500+ language pairs (500+ languages). Key statistics of all language pairs available at… See the full description on the dataset page: https://huggingface.co/datasets/MaLA-LM/mala-bilingual-translation-corpus.texttranslation10B<n<100B8 likes8.2k downloads2mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.