CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01derek-thomas /ScienceQA Dataset Card Creation Guide Dataset Summary Learn to Explain: Multimodal Reasoning via Thought Chains for Science Question Answering Supported Tasks and Leaderboards Multi-modal Multiple Choice Languages English Dataset Structure Data Instances Explore more samples here. {'image': Image, 'question': 'Which of these states is farthest north?', 'choices': ['West Virginia', 'Louisiana', 'Arizona', 'Oklahoma'], 'answer': 0… See the full description on the dataset page: https://huggingface.co/datasets/derek-thomas/ScienceQA.imagemultiple-choice10K<n<100K234 likes40k downloads4y agoHugging Face02lmms-lab-encoder /ScienceQA Large-scale Multi-modality Models Evaluation Suite Accelerating the development of large-scale multi-modality models (LMMs) with lmms-eval 🏠 Homepage | 📚 Documentation | 🤗 Huggingface Datasets This Dataset This is a formatted version of derek-thomas/ScienceQA. It is used in our lmms-eval pipeline to allow for one-click evaluations of large multi-modality models. @inproceedings{lu2022learn, title={Learn to Explain: Multimodal Reasoning via Thought… See the full description on the dataset page: https://huggingface.co/datasets/lmms-lab-encoder/ScienceQA.image10K<n<100K10 likes18k downloads3y agoHugging Face03ScienceOne-AI /S1-MMAlignS1-MMAlign A Large-Scale Multi-Disciplinary Scientific Multimodal Dataset S1-MMAlign is a large-scale, multi-disciplinary multimodal dataset comprising over 15.5 million high-quality image-text pairs derived from 2.5 million open-access scientific papers. Multimodal learning has revolutionized general domain tasks, yet its application in scientific discovery is hindered by the profound semantic gap between complex scientific imagery and sparse textual descriptions. S1-MMAlign aims to… See the full description on the dataset page: https://huggingface.co/datasets/ScienceOne-AI/S1-MMAlign.imageimage-to-text10M<n<100M106 likes2.5k downloads6mo agoHugging Face04J0nasW /science-datalake Science Data Lake A unified, portable science data lake integrating 7 scholarly datasets (~525 GB Parquet) with cross-dataset DOI normalization, 13 scientific ontologies (1.3M terms), and a reproducible ETL pipeline. Note: One additional source (Semantic Scholar S2AG) is supported by the pipeline but is not redistributed here due to its API terms of service. See Not Included in This Upload below. What's Unique This dataset enables queries… See the full description on the dataset page: https://huggingface.co/datasets/J0nasW/science-datalake.imagetext-classification10B<n<100B10 likes2.3k downloads6mo agoHugging Face05vidore /vidore_v3_computer_scienceViDoRe V3 : Computer Science This dataset, Computer Science, is a corpus of textbooks from the openstacks website, intended for long-document understanding tasks. It is one of the 10 corpora comprising the ViDoRe v3 Benchmark. About ViDoRe v3 ViDoRe V3 is our latest benchmark for RAG evaluation on visually-rich documents from real-world applications. It features 10 datasets with, in total, 26,000 pages and 3099 queries, translated into 6 languages. Each query comes with… See the full description on the dataset page: https://huggingface.co/datasets/vidore/vidore_v3_computer_science.documentvisual-document-retrieval1K<n<10K6 likes2.2k downloads8mo agoHugging Face06lucazhou2000 /sciencemysterybench-transcriptsimagen<1K0 likes2.2k downloads9d agoHugging Face07lmms-lab /ScienceQA-IMG Large-scale Multi-modality Models Evaluation Suite Accelerating the development of large-scale multi-modality models (LMMs) with lmms-eval 🏠 Homepage | 📚 Documentation | 🤗 Huggingface Datasets This Dataset This is a formatted and filtered version of derek-thomas/ScienceQA with only image instances. It is used in our lmms-eval pipeline to allow for one-click evaluations of large multi-modality models. @inproceedings{lu2022learn, title={Learn to Explain:… See the full description on the dataset page: https://huggingface.co/datasets/lmms-lab/ScienceQA-IMG.image10K<n<100K5 likes1.3k downloads3y agoHugging Face08vidore /vidore_v3_computer_science_mteb_format Vidore3ComputerScienceRetrieval An MTEB dataset Massive Text Embedding Benchmark Retrieve associated pages according to questions. Task category t2i Domains Academic Reference https://huggingface.co/blog/QuentinJG/introducing-vidore-v3 Source datasets: vidore/vidore_v3_computer_science How to evaluate on this task You can evaluate an embedding model on this dataset using the following code: import mteb task =… See the full description on the dataset page: https://huggingface.co/datasets/vidore/vidore_v3_computer_science_mteb_format.imagevisual-document-retrieval10K<n<100K0 likes995 downloads11mo agoHugging Face09ScienceOne-AI /S1-Omni-Corpus-10K S1-Omni-Corpus-10K An open-source scientific multimodal reasoning dataset subset for S1-Omni 🧬 Model Introduction S1-Omni is a unified scientific multimodal reasoning model for scientific understanding, prediction, and generation. It is developed by the ScienceOne AI team of the Chinese Academy of Sciences. S1-Omni addresses fragmented scientific AI capabilities with a shared backbone for cross-disciplinary, cross-modal, and cross-task understanding and reasoning… See the full description on the dataset page: https://huggingface.co/datasets/ScienceOne-AI/S1-Omni-Corpus-10K.image10K<n<100K1 likes920 downloads2mo agoHugging Face10ScienceOne-AI /SciGenEdit-10K SciGenEdit-10K An Open Dataset for Scientific Image Generation and Editing English | 简体中文 📖 Introduction SciGenEdit-10K is a public subset released with the S1-Omni-Image project. It is designed for research on scientific image generation, scientific image editing, and multi-turn scientific image generation and editing. S1-Omni-Image is a unified multimodal model developed by the ScienceOne team at the Chinese Academy of Sciences for scientific… See the full description on the dataset page: https://huggingface.co/datasets/ScienceOne-AI/SciGenEdit-10K.imagetext-to-image10K<n<100K3 likes606 downloads3mo agoHugging Face11colour-science /colour-checker-detection-dataset Colour - Checker Detection - Dataset An image dataset of colour rendition charts. This dataset is structured according to Ultralytics YOLO format and ready to use with YOLOv8. The colour-science/colour-checker-detection-models models resulting from the YOLOv8 segmentation training are supporting colour rendition charts detection in the Colour Checker Detection Python package. Classes ColorCheckerClassic24: Calibrite / X-Rite ColorCheckerClassic 24 Contact &… See the full description on the dataset page: https://huggingface.co/datasets/colour-science/colour-checker-detection-dataset.imageobject-detectionn<1K2 likes541 downloads3y agoHugging Face12HumynLabs /Handwritten-Computer-Science-Notes-Dataset English Handwritten Computer Science Notes Dataset This dataset contains high-resolution images of handwritten computer science notes written in English. It includes algorithm explanations, code snippets, flowcharts, theoretical content, and annotations. The dataset is designed to support AI research in handwriting recognition, OCR, and document understanding specifically for computer science education. Contact For queries or collaborations related to this dataset… See the full description on the dataset page: https://huggingface.co/datasets/HumynLabs/Handwritten-Computer-Science-Notes-Dataset.documentimage-classificationn<1K2 likes502 downloads11mo agoHugging Face13mariiakoroliuk /generalization-science-datadocumentn<1K0 likes410 downloads3d agoHugging Face14Jialuo21 /Science-T2I-Fullset Science-T2I Fullset Resources Website arXiv: Paper GitHub: Code Huggingface: SciScore Huggingface: Science-T2I-S&C Benchmark Data The Science-T2I Fullset comprises a comprehensive collection of data for scientific T2I generation, including both training and test sets with a unified data structure. The test sets are split into 'test-S' and 'test-C,' corresponding to the Science-T2I-S and Science-T2I-C benchmarks, respectively. Download Dataset… See the full description on the dataset page: https://huggingface.co/datasets/Jialuo21/Science-T2I-Fullset.image1K<n<10K0 likes316 downloads1y agoHugging Face15ByteDance-Seed /ScienceOlympiad ScienceOlympiad: Challenging AI with Olympiad-Level Multimodal Science Problems Dataset Description The ScienceOlympiad dataset is a meticulously curated benchmark designed to test the limits of current AI models in scientific reasoning. It comprises elite, competition-level problems in physics and chemistry. Addressing the need for more diverse and realistic challenges, ScienceOlympiad introduces multimodal integration as a key dimension. Unlike purely text-based… See the full description on the dataset page: https://huggingface.co/datasets/ByteDance-Seed/ScienceOlympiad.imagen<1K5 likes268 downloads1y agoHugging Face16TheMrguiller /ScienceQA Dataset Card for "ScienceQA" Dataset Summary ScienceQA is collected from elementary and high school science curricula, and contains 21,208 multimodal multiple-choice science questions. Out of the questions in ScienceQA, 10,332 (48.7%) have an image context, 10,220 (48.2%) have a text context, and 6,532 (30.8%) have both. Most questions are annotated with grounded lectures (83.9%) and detailed explanations (90.5%). The lecture and explanation provide general external… See the full description on the dataset page: https://huggingface.co/datasets/TheMrguiller/ScienceQA.imagequestion-answering10K<n<100K6 likes181 downloads3y agoHugging Face17tomyuanyucheng /tb-science-bloodculture-gram-callimagen<1K0 likes165 downloads26d agoHugging Face18nasa-cisto-data-science-group /modis-lake-powell-toy-dataset MODIS Water Lake Powell Toy Dataset Dataset Summary Tabular dataset comprised of MODIS surface reflectance bands along with calculated indices and a label (water/not-water) Dataset Structure Data Fields water: Label, water or not-water (binary) sur_refl_b01_1: MODIS surface reflection band 1 (-100, 16000) sur_refl_b02_1: MODIS surface reflection band 2 (-100, 16000) sur_refl_b03_1: MODIS surface reflection band 3 (-100, 16000) sur_refl_b04_1: MODIS… See the full description on the dataset page: https://huggingface.co/datasets/nasa-cisto-data-science-group/modis-lake-powell-toy-dataset.image1K<n<10K1 likes148 downloads3y agoHugging Face19InnovatorLab /Innovator-VL-Instruct-Scienceimage100K<n<1M1 likes139 downloads7mo agoHugging Face20cnut1648 /ScienceQA-LLAVA Dataset Card for "ScienceQA-LLAVA" More Information needed image10K<n<100K0 likes132 downloads3y agoHugging Face21GodotCN /science-datalake Science Data Lake A unified, portable science data lake integrating 7 scholarly datasets (~525 GB Parquet) with cross-dataset DOI normalization, 13 scientific ontologies (1.3M terms), and a reproducible ETL pipeline. Note: One additional source (Semantic Scholar S2AG) is supported by the pipeline but is not redistributed here due to its API terms of service. See Not Included in This Upload below. What's Unique This dataset enables queries… See the full description on the dataset page: https://huggingface.co/datasets/GodotCN/science-datalake.imagetext-classification10B<n<100B1 likes129 downloads6mo agoHugging Face22OS-Copilot /ScienceBoard-Trajimage1 likes93 downloads1y agoHugging Face23yiqingliang /scienceqa-problems-datasetimage1K<n<10K0 likes90 downloads1y agoHugging Face24MounaAjaani /data_science_bowl_2018image0 likes83 downloads2y agoHugging Face25nasa-cisto-data-science-group /tutorial-senegal-lclucgeospatialn<1K0 likes82 downloads3y agoHugging Face26ahmedheakl /science-r1image1K<n<10K0 likes81 downloads2y agoHugging Face27RA-Data-Science /DiEm_HTR Dataset Card for DiEm HTR The DiEm HTR dataset is a ground truth dataset for historical danish handwriting in the 17th and 18th century, generated as part of the Digitalisering af Enesteministerialbøger project at the Danish National Archives. Dataset Details Dataset Description The Digitalisering af Enesteministerialbøger project (DiEm) at the Danish National Archives aims to transcribe and make publically available all of the danish parish registers from… See the full description on the dataset page: https://huggingface.co/datasets/RA-Data-Science/DiEm_HTR.imageimage-to-textn<1K0 likes73 downloads8mo agoHugging Face28mm-eval /ScienceQAimage10K<n<100K0 likes72 downloads2mo agoHugging Face29Gisiyuan /ScienceQA Dataset Card Creation Guide Dataset Summary Learn to Explain: Multimodal Reasoning via Thought Chains for Science Question Answering Supported Tasks and Leaderboards Multi-modal Multiple Choice Languages English Dataset Structure Data Instances Explore more samples here. {'image': Image, 'question': 'Which of these states is farthest north?', 'choices': ['West Virginia', 'Louisiana', 'Arizona', 'Oklahoma'], 'answer': 0… See the full description on the dataset page: https://huggingface.co/datasets/Gisiyuan/ScienceQA.imagemultiple-choice10K<n<100K0 likes71 downloads8mo agoHugging Face30HuggingFaceM4 /ScienceQAImg_Modif Dataset Card for "ScienceQAImg_Modif" This dataset contains the ScienceQA benchmark where only examples with an image are kept, and where we formatted the prompt. image10K<n<100K2 likes67 downloads3y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.